Test-Driven Prompting

Flip the order: write the test first, then prompt the model to make it pass, so correctness is defined before code exists.

TL;DR

  1. Write the test first; it is an executable spec the model's code must satisfy.
  2. The failing test gives the model an exact target and catches hallucinated APIs instantly.
  3. Loop: red (failing test) to green (passing code) to refactor, with the suite as the gate.

Why Test First

    Executable Spec

    A test states the required behavior in a form the model cannot misread.

    expect(slug('A B')).toBe('a-b')
    // the spec, not a paragraph
    Instant Hallucination Check

    Calls to functions that do not exist fail the moment you run the test.

    Fake API -> import error -> red
    (no silent plausible code)
    Unambiguous Done

    Green means done against the spec; there is nothing to debate.

    Red -> not done. Green -> done.

The Loop

    Red

    Write a failing test that captures the next piece of behavior.

    it('throws on empty input', () => {
      expect(() => parse('')).toThrow();
    });
    Green

    Prompt the model to write the minimal code to pass it.

    "Make this test pass. Do not change
    the test. Minimal implementation."
    Refactor

    With tests green, prompt a clean-up that keeps them green.

    "Refactor for readability; tests must
    stay green."

Keep The Test Honest

    You Own It

    Review the test yourself; it is the definition of correct.

    Wrong test -> correctly-wrong code.
    Lock It

    Forbid edits to the test so 'passing' cannot mean weakening it.

    "Change only the implementation,
    not the test file."
    Mean What You Test

    Ensure the assertions capture real intent, not a trivial shortcut.

    Test the contract, not a value the
    code could fake.

Scale The Loop

    One Behavior At A Time

    Add one failing test, make it green, then add the next.

    Small red -> green steps beat one
    big test dump.
    Let Agents Iterate

    An agent can run the suite and keep going until green, within scope.

    "Run the tests; iterate until the
    new test passes."
    Guard Regressions

    Every new green test stays as a guard for future changes.

    The suite grows into a safety net.

Tips

  1. Write or review the test yourself so the spec is correct; then let the model implement against it.
  2. Give the model the test and say 'make this pass without changing the test'.

Warnings

  1. If the model may edit the test, it can 'pass' by weakening the test; lock the test as fixed.
  2. A green test only proves the stated behavior; write tests that actually capture what you mean by correct.

In Practice

FAQ