Evaluation & Benchmarks

Build automated test suites, implement LLM-as-a-judge rubrics, and track accuracy benchmarks for AI pipelines.

TL;DR

  1. Build a golden dataset of typed EvalSample records with known answers.
  2. Score open-ended answers with an LLM-as-a-judge and a written rubric.
  3. Run evals in CI so quality regressions fail the build.

LLM-as-a-Judge Rubrics

    Structured Scoring Rubric Prompt

    Prompt frontier model to evaluate answer quality on 1-5 scale.

    function buildJudgePrompt(
      query: string, answer: string, expected: string
    ) {
      return 'Score answer against expected (1 to 5).\n' +
        `Q: ${query}\nExp: ${expected}\nAns: ${answer}\n` +
        'Return JSON: { score, reasoning }';
    }
    Judge Score Parser & Zod Validator

    Validate that judge model outputs typed numeric score.

    import { z } from 'zod';
    const JudgeSchema = z.object({
      score: z.number().min(1).max(5),
      reasoning: z.string().min(10),
    });
    const json = JSON.parse(judgeOutput);
    const result = JudgeSchema.parse(json);
    Pairwise Comparison with Bias Swap

    Evaluate two model candidates swapping order to negate bias.

    async function comparePairwise(
      q: string, a1: string, a2: string
    ) {
      const [s1, s2] = await Promise.all([
        judgePair(q, a1, a2),
        judgePair(q, a2, a1),
      ]);
      const pickA = s1.winner === 'A' && s2.winner === 'B';
      return pickA ? a1 : a2;
    }

Deterministic Evaluation Metrics

    Exact Match & Normalized Equivalence

    Fast zero-cost accuracy check for factual outputs.

    function exactMatch(actual: string, expected: string) {
      const norm = (s: string) => {
        return s.trim().toLowerCase().replace(/\s+/g, ' ');
      };
      return norm(actual) === norm(expected);
    }
    Token F1 Score Calculation

    Calculate precision, recall, and harmonic F1 across tokens.

    function tokenF1(actual: string, expected: string) {
      const a = actual.toLowerCase().split(/\s+/);
      const e = expected.toLowerCase().split(/\s+/);
      const aTokens = new Set(a);
      const eTokens = new Set(e);
      const list = [...aTokens];
      const match = (t: string) => eTokens.has(t);
      const overlap = list.filter(match).length;
      if (overlap === 0) return 0;
      const p = overlap / aTokens.size;
      const r = overlap / eTokens.size;
      return (2 * p * r) / (p + r);
    }
    Regex & Schema Assertions

    Assert output satisfies expected format constraints.

    function assertFormat(val: string, re: RegExp) {
      if (!re.test(val)) {
        throw new Error(`Failed pattern: ${re}`);
      }
    }

Golden Dataset & Test Suite Runner

    Eval Sample Interface Contract

    Structure test case records with inputs and expectations.

    type EvalSample = {
      id: string;
      query: string;
      expected: string;
      category: 'rag' | 'math' | 'formatting';
      minScore: number;
    };
    Continuous Evaluation Suite Runner

    Iterate across golden test dataset and aggregate metrics.

    async function runTestSuite(
      samples: EvalSample[],
      pipelineFn: (q: string) => Promise<string>
    ) {
      const results = [];
      for (const s of samples) {
        const answer = await pipelineFn(s.query);
        const pass = exactMatch(answer, s.expected);
        results.push({ id: s.id, pass });
      }
      const passed = results.filter(r => r.pass).length;
      const passRate = passed / samples.length;
      return { passRate, results };
    }
    CI/CD Regression Threshold Check

    Fail pull request build if test pass rate regresses.

    const MIN_PASS_RATE = 0.95;
    if (report.passRate < MIN_PASS_RATE) {
      const r = report.passRate;
      console.error(`Rate ${r} < ${MIN_PASS_RATE}`);
      process.exit(1);
    }

Synthetic Dataset Generation

    Q&A Pair Generation from Documents

    Generate synthetic test queries from knowledge documents.

    async function makeQA(doc: string, client: any) {
      const p = 'Read document and output 3 queries ' +
        'and answers as JSON:\n' + doc;
      const res = await client.chat.completions.create({
        model: 'gpt-6.1-sol',
        messages: [{ role: 'user', content: p }],
        response_format: { type: 'json_object' },
      });
      return JSON.parse(res.choices[0].message.content!);
    }
    Adversarial Edge Case Generator

    Synthesize edge cases to test system resilience.

    const edgePrompt = 'Generate 5 tricky edge cases ' +
      'testing ambiguous dates and null fields.';
    Dataset Deduplication & Cleaning

    Remove duplicate synthetic queries using embedding distance.

    function dedupeSamples(samples: EvalSample[]) {
      const seen = new Set<string>();
      return samples.filter(s => {
        const k = s.query.trim().toLowerCase();
        if (seen.has(k)) return false;
        seen.add(k);
        return true;
      });
    }

Tips

  1. Swap candidate answer positions during pairwise llmJudge scoring passes to eliminate model order-bias preferences completely across comparative runs.
  2. Measure fast deterministic metrics like exact match, regex assertions, and zod Schema validity before executing expensive LLM judge evaluations.

Warnings

  1. Never ship a prompt change or model swap without checking the passRate across your whole golden dataset first.
  2. Avoid using weak lightweight models as judges because grading nuanced reasoning requires frontier-class gpt-6.1-sol or Claude evaluation.

In Practice

FAQ