Verifying & Validating AI Output

Build a habit of verifying AI output: read it, run it, test it, and watch for the anti-patterns models reliably produce.

TL;DR

  1. Never merge AI code you have not read, run, and tested; confidence is not correctness.
  2. Let tools verify: type-checker, linter, tests, and the diff view do the objective checking.
  3. Learn the anti-patterns models repeat so you can spot them fast.

The Verification Ladder

    Read It

    Understand the code; if you cannot explain it, do not ship it.

    Can you explain every line? If no,
    stop.
    Run It

    Execute it, or at least type-check, before trusting it.

    tsc --noEmit && run the real path
    Test The Edges

    Exercise empty, null, boundary, and error cases, not just happy path.

    Empty? Null? Max? Error? All
    covered?

Let Tools Check

    Types & Lint

    Objective checks catch fake APIs and style drift instantly.

    tsc, eslint -> run every time
    Mutation Sanity

    Break the code on purpose; a real test suite goes red.

    Introduce a bug -> tests fail?
    If not, tests are weak.
    Diff Review

    Read the diff for scope creep and unintended changes.

    git diff  # only what you asked for?

Spot The Anti-Patterns

    Noise & Narration

    Comments that restate the code and needless over-specification.

    // increment i  <- delete this
    noise
    Swallowed Errors

    Empty catch blocks and ignored failures that hide problems.

    catch (e) {}  // silent failure,
    reject this
    Déjà-Vu Bugs

    Reintroduced bugs and duplicated logic instead of reuse.

    Is this the bug we fixed, back
    again?

Scale To Risk

    Low Stakes

    A throwaway script needs a read and a run.

    Script -> read + run
    High Stakes

    Data, money, auth, or prod code needs tests and human review.

    Prod/auth/$ -> tests + review
    + second pair of eyes
    Own It

    You are accountable for merged code, whoever (or what) wrote it.

    Your name is on the commit.

Tips

  1. Read the code before you run it; understand what it does rather than trusting it compiles.
  2. Review the diff for scope creep, changes you did not ask for are a red flag, not a bonus.

Warnings

  1. A fluent, confident explanation is not evidence; verify the behavior, not the prose around it.
  2. Tests the model wrote for its own code can pass against its own bugs; sanity-check the assertions.

In Practice

FAQ