Verifying & Validating AI Output
Build a habit of verifying AI output: read it, run it, test it, and watch for the anti-patterns models reliably produce.
TL;DR
- Never merge AI code you have not read, run, and tested; confidence is not correctness.
- Let tools verify: type-checker, linter, tests, and the diff view do the objective checking.
- Learn the anti-patterns models repeat so you can spot them fast.
The Verification Ladder
Read ItUnderstand the code; if you cannot explain it, do not ship it.
Can you explain every line? If no,
stop.Run ItExecute it, or at least type-check, before trusting it.
tsc --noEmit && run the real pathTest The EdgesExercise empty, null, boundary, and error cases, not just happy path.
Empty? Null? Max? Error? All
covered?Let Tools Check
Types & LintObjective checks catch fake APIs and style drift instantly.
tsc, eslint -> run every timeMutation SanityBreak the code on purpose; a real test suite goes red.
Introduce a bug -> tests fail?
If not, tests are weak.Diff ReviewRead the diff for scope creep and unintended changes.
git diff # only what you asked for?Spot The Anti-Patterns
Noise & NarrationComments that restate the code and needless over-specification.
// increment i <- delete this
noiseSwallowed ErrorsEmpty catch blocks and ignored failures that hide problems.
catch (e) {} // silent failure,
reject thisDéjà-Vu BugsReintroduced bugs and duplicated logic instead of reuse.
Is this the bug we fixed, back
again?Scale To Risk
Low StakesA throwaway script needs a read and a run.
Script -> read + runHigh StakesData, money, auth, or prod code needs tests and human review.
Prod/auth/$ -> tests + review
+ second pair of eyesOwn ItYou are accountable for merged code, whoever (or what) wrote it.
Your name is on the commit.Tips
- Read the code before you run it; understand what it does rather than trusting it compiles.
- Review the diff for scope creep, changes you did not ask for are a red flag, not a bonus.
Warnings
- A fluent, confident explanation is not evidence; verify the behavior, not the prose around it.
- Tests the model wrote for its own code can pass against its own bugs; sanity-check the assertions.
In Practice
A short, repeatable checklist for AI output, scaled to risk. Running it every time turns 'it looks right' into 'it is verified', and catches the anti-patterns models reliably produce.
- Reading for understanding is the gate: unexplainable code does not merge.
- Tools, type-check, lint, tests, do the objective verification.
- A mutation check confirms the tests actually have teeth.
- Scaling scrutiny to risk keeps effort proportional to consequences.
# Run this on every AI-generated change:
[ ] I read it and can explain every line
[ ] It type-checks and lints clean (tsc, eslint)
[ ] Edge cases tested: empty, null, boundary, error
[ ] I broke the code on purpose -> a test went red
[ ] Diff review: no scope creep, no changes I didn't ask for
[ ] No anti-patterns: narration comments, empty catch,
duplicated logic, a bug we already fixed
[ ] No secrets, no disabled security checks
[ ] Risk check: data/money/auth/prod? -> tests + human review
# If any box fails, fix it or send it back with:
"This has <issue> at <line>. Fix only that; don't
change anything else."FAQ
Read it and make sure you understand it, run it or type-check it, and test the edge cases, not just the happy path. Review the diff for anything you did not ask for. If you cannot explain what a line does, do not merge it.
Because a model can write tests that assert whatever its own (possibly buggy) code does, so they pass against the bug. Sanity-check that the expected values are actually correct, and run a mutation check: break the code and confirm a test goes red.
Over-commented code that narrates every line, swallowed errors (empty catch blocks), edge-case over-specification that adds noise, duplicated logic instead of reuse, re-introducing bugs you already fixed, and 'by-the-book' fixes that miss your actual context. Knowing these speeds up review.
Scale it to risk. A throwaway script needs a read and a run. Code touching data, money, auth, or production needs tests, review, and often a second human. The more consequential the code, the more verification it earns.