Model Selection

Choose a model tier by weighing task difficulty, latency, context size, and cost, then route requests between tiers.

TL;DR

  1. Pick a model tier, not a brand, for each task.
  2. Start with the best model, then step down while evals pass.
  3. Keep every ID in one MODELS config for one-line swaps.

Model Tier Classification

    Frontier Tier

    Highest capability for hard reasoning, coding, and long agent runs.

    const MODELS = {
      frontier: 'claude-opus-5-5',
      // ...other tiers below
    };
    // Best for architecture, hard bugs, agents
    Mid Tier

    Fast and affordable for everyday coding, chat, and tool use.

    const MODELS = {
      frontier: 'claude-opus-5-5',
      balanced: 'claude-sonnet-5-5',
    };
    // Default for most production traffic
    Fast Tier

    Lowest latency and cost for tagging, routing, and extraction.

    const MODELS = {
      frontier: 'claude-opus-5-5',
      balanced: 'claude-sonnet-5-5',
      fast: 'claude-haiku-4-5',
    };
    // Sub-second, built for volume

Decision Matrix Parameters

    Cost Per Million Tokens

    Compare input and output prices across tiers before you commit.

    // USD per million tokens (input / output)
    const PRICE = {
      frontier: [4, 20], // Opus 5.5
      balanced: [2, 10], // Sonnet 5.5
      fast: [1, 5], // Haiku 4.5
    };
    Reasoning Effort

    Tune depth of thinking per request instead of swapping models.

    const res = await client.messages.create({
      model: MODELS.frontier,
      max_tokens: 16000,
      output_config: { effort: 'low' },
      messages,
    });
    // low | medium | high | xhigh | max
    Context Window Check

    Read a model's real limits from the API instead of guessing.

    const info = await client.models.retrieve(
      MODELS.balanced
    );
    console.log(info.max_input_tokens);
    console.log(info.max_tokens);

Cascading Routing Logic

    Task Difficulty Classifier

    Inspect the request before choosing a tier.

    function routeTask(prompt: string) {
      const hard = /refactor|proof|audit/i;
      return hard.test(prompt)
        ? MODELS.frontier
        : MODELS.fast;
    }
    Fallback On Failure

    Escalate when the cheap model fails validation.

    const first = await run(MODELS.fast, input);
    if (schema.safeParse(first).success) return first;
    return run(MODELS.frontier, input);
    Cost Budget Enforcer

    Drop to a cheaper tier when monthly spend passes the limit.

    const active = monthlySpend > limit
      ? MODELS.fast
      : MODELS.balanced;

Provider Feature Checklist

    Prompt Caching

    Check whether repeated context is discounted on your provider.

    // Anthropic: explicit cache_control, reads ~90% off
    // OpenAI: automatic on long shared prefixes
    // See the Prompt Caching sheet
    Capability Flags

    Ask the API what a model supports before relying on it.

    const m = await client.models.retrieve(id);
    console.log(m.capabilities);
    // vision, thinking, structured outputs...
    Fallback Provider

    Keep a second provider configured for outages and rate limits.

    const PROVIDERS = [
      { name: 'anthropic', model: MODELS.balanced },
      { name: 'openai', model: OPENAI_MODEL },
    ];
    // Try in order, log which one served

Tips

  1. Lower output_config.effort before switching to a weaker model, since it often keeps quality while cutting cost and latency.
  2. Ask the provider for a model's limits at runtime, for example client.models.retrieve(id), instead of hard-coding context sizes.

Warnings

  1. Do not pick a model from a leaderboard alone: run your own eval on your own prompts before you commit.
  2. Model IDs and prices change often, so read them from one MODELS config and recheck provider docs before each release.

In Practice

FAQ