Test-Driven Prompting
Flip the order: write the test first, then prompt the model to make it pass, so correctness is defined before code exists.
TL;DR
- Write the test first; it is an executable spec the model's code must satisfy.
- The failing test gives the model an exact target and catches hallucinated APIs instantly.
- Loop: red (failing test) to green (passing code) to refactor, with the suite as the gate.
Why Test First
Executable SpecA test states the required behavior in a form the model cannot misread.
expect(slug('A B')).toBe('a-b')
// the spec, not a paragraphInstant Hallucination CheckCalls to functions that do not exist fail the moment you run the test.
Fake API -> import error -> red
(no silent plausible code)Unambiguous DoneGreen means done against the spec; there is nothing to debate.
Red -> not done. Green -> done.The Loop
RedWrite a failing test that captures the next piece of behavior.
it('throws on empty input', () => {
expect(() => parse('')).toThrow();
});GreenPrompt the model to write the minimal code to pass it.
"Make this test pass. Do not change
the test. Minimal implementation."RefactorWith tests green, prompt a clean-up that keeps them green.
"Refactor for readability; tests must
stay green."Keep The Test Honest
You Own ItReview the test yourself; it is the definition of correct.
Wrong test -> correctly-wrong code.Lock ItForbid edits to the test so 'passing' cannot mean weakening it.
"Change only the implementation,
not the test file."Mean What You TestEnsure the assertions capture real intent, not a trivial shortcut.
Test the contract, not a value the
code could fake.Scale The Loop
One Behavior At A TimeAdd one failing test, make it green, then add the next.
Small red -> green steps beat one
big test dump.Let Agents IterateAn agent can run the suite and keep going until green, within scope.
"Run the tests; iterate until the
new test passes."Guard RegressionsEvery new green test stays as a guard for future changes.
The suite grows into a safety net.Tips
- Write or review the test yourself so the spec is correct; then let the model implement against it.
- Give the model the test and say 'make this pass without changing the test'.
Warnings
- If the model may edit the test, it can 'pass' by weakening the test; lock the test as fixed.
- A green test only proves the stated behavior; write tests that actually capture what you mean by correct.
In Practice
You write the failing test that specifies the behavior, then hand it to the model as a fixed target. The model writes code to pass it, the suite is the gate, and the test is never edited.
- You author the test, so the specification is correct and owned by you.
- The prompt hands over the test and forbids changing it.
- The model writes the minimal implementation to turn the test green.
- A refactor step follows, with the same green-suite gate protecting behavior.
# 1. YOU write the failing test (the spec)
import { parseDuration } from './time';
it('parses "1h30m" to 5400 seconds', () => {
expect(parseDuration('1h30m')).toBe(5400);
});
it('throws on an empty string', () => {
expect(() => parseDuration('')).toThrow();
});
# 2. PROMPT the model
Implement `parseDuration(input: string): number` so
the test above passes. Change only the implementation,
not the test. Keep it minimal. Then run the suite.
# 3. After green, prompt a refactor
"Refactor parseDuration for readability.
The tests must stay green. Return a diff."FAQ
The test is an unambiguous, executable specification. It gives the model a precise target, and it fails immediately if the model invents a non-existent function or gets behavior wrong. There is no ambiguity to argue about: the code either turns the test green or it does not.
You own the test, since it defines correctness. You can draft it with the model's help, but review it carefully, because a wrong test will drive the model to write wrong code that passes. Then hand the model the finished test as a fixed target.
State it plainly: 'make the test pass by changing only the implementation; do not modify the test'. In agentic setups, keep the test file out of the editable scope, or review the diff to confirm the test was untouched.
Yes, with the model as the implementer. The red-green-refactor loop is unchanged; what changes is that you spend your effort on the spec (the test) and the review, while the model handles turning red into green.