Model Selection
Choose a model tier by weighing task difficulty, latency, context size, and cost, then route requests between tiers.
TL;DR
- Pick a model
tier, not a brand, for each task. - Start with the best model, then step down while
evalspass. - Keep every ID in one
MODELSconfig for one-line swaps.
Model Tier Classification
Frontier TierHighest capability for hard reasoning, coding, and long agent runs.
const MODELS = {
frontier: 'claude-opus-5-5',
// ...other tiers below
};
// Best for architecture, hard bugs, agentsMid TierFast and affordable for everyday coding, chat, and tool use.
const MODELS = {
frontier: 'claude-opus-5-5',
balanced: 'claude-sonnet-5-5',
};
// Default for most production trafficFast TierLowest latency and cost for tagging, routing, and extraction.
const MODELS = {
frontier: 'claude-opus-5-5',
balanced: 'claude-sonnet-5-5',
fast: 'claude-haiku-4-5',
};
// Sub-second, built for volumeDecision Matrix Parameters
Cost Per Million TokensCompare input and output prices across tiers before you commit.
// USD per million tokens (input / output)
const PRICE = {
frontier: [4, 20], // Opus 5.5
balanced: [2, 10], // Sonnet 5.5
fast: [1, 5], // Haiku 4.5
};Reasoning EffortTune depth of thinking per request instead of swapping models.
const res = await client.messages.create({
model: MODELS.frontier,
max_tokens: 16000,
output_config: { effort: 'low' },
messages,
});
// low | medium | high | xhigh | maxContext Window CheckRead a model's real limits from the API instead of guessing.
const info = await client.models.retrieve(
MODELS.balanced
);
console.log(info.max_input_tokens);
console.log(info.max_tokens);Cascading Routing Logic
Task Difficulty ClassifierInspect the request before choosing a tier.
function routeTask(prompt: string) {
const hard = /refactor|proof|audit/i;
return hard.test(prompt)
? MODELS.frontier
: MODELS.fast;
}Fallback On FailureEscalate when the cheap model fails validation.
const first = await run(MODELS.fast, input);
if (schema.safeParse(first).success) return first;
return run(MODELS.frontier, input);Cost Budget EnforcerDrop to a cheaper tier when monthly spend passes the limit.
const active = monthlySpend > limit
? MODELS.fast
: MODELS.balanced;Provider Feature Checklist
Prompt CachingCheck whether repeated context is discounted on your provider.
// Anthropic: explicit cache_control, reads ~90% off
// OpenAI: automatic on long shared prefixes
// See the Prompt Caching sheetCapability FlagsAsk the API what a model supports before relying on it.
const m = await client.models.retrieve(id);
console.log(m.capabilities);
// vision, thinking, structured outputs...Fallback ProviderKeep a second provider configured for outages and rate limits.
const PROVIDERS = [
{ name: 'anthropic', model: MODELS.balanced },
{ name: 'openai', model: OPENAI_MODEL },
];
// Try in order, log which one servedTips
- Lower
output_config.effortbefore switching to a weaker model, since it often keeps quality while cutting cost and latency. - Ask the provider for a model's limits at runtime, for example
client.models.retrieve(id), instead of hard-coding context sizes.
Warnings
- Do not pick a model from a leaderboard alone: run your own
evalon your own prompts before you commit. - Model IDs and prices change often, so read them from one
MODELSconfig and recheck provider docs before each release.
In Practice
Routes each prompt to a fast or frontier tier, then escalates if the cheap answer fails validation.
- Keep every model ID in one MODELS config.
- Send easy prompts to the fast tier first.
- Validate the answer, and escalate to the frontier tier on failure.
- Return the final text.
import Anthropic from '@anthropic-ai/sdk';
const client = new Anthropic();
const MODELS = {
fast: 'claude-haiku-4-5',
frontier: 'claude-opus-5-5',
};
async function run(model: string, query: string) {
const res = await client.messages.create({
model,
max_tokens: 500,
messages: [{ role: 'user', content: query }],
});
const block = res.content.find(b => b.type === 'text');
return block?.type === 'text' ? block.text : '';
}
async function dispatch(query: string) {
const first = await run(MODELS.fast, query);
return first.trim() ? first : run(MODELS.frontier, query);
}
console.log(await dispatch('Format this JSON string'));FAQ
Use higher reasoning effort for multi-step coding, math, and planning. For chat, extraction, and classification, low effort is faster and cheaper with little quality loss. Most current models expose this as an effort setting rather than a separate model.
A cascade sends each request to a cheap model first. If validation fails or confidence is low, it escalates to a more capable model. You pay frontier prices only for the hard cases.
Current frontier models offer around one million tokens, and fast tiers may offer less (Claude Haiku 4.5 has 200K). Most single-turn tasks use a small fraction. Long documents are usually cheaper with retrieval than with a giant prompt.