Claude API Fundamentals

Connect to the Anthropic Claude API, configure system prompts, select model tiers, and inspect response stop reasons.

TL;DR

  1. Set ANTHROPIC_API_KEY and let the SDK read it automatically.
  2. Call client.messages.create with a model, max_tokens, and messages.
  3. Check stop_reason before trusting text: it reveals truncated answers.

Client Configuration

    Anthropic Client Setup

    Instantiate SDK client reading credentials from environment variables.

    import Anthropic from '@anthropic-ai/sdk';
    
    const client = new Anthropic();
    // Reads ANTHROPIC_API_KEY automatically
    Model Tier Constants

    Declare type-safe constants for verified production model IDs.

    const OPUS = 'claude-opus-5-5';
    const SONNET = 'claude-sonnet-5-5';
    const HAIKU = 'claude-haiku-4-5';
    // IDs are complete as-is: no date suffix
    Retries And Timeouts

    Tune automatic retries and the request timeout (milliseconds) once per client.

    const client = new Anthropic({
      maxRetries: 3, // default is 2
      timeout: 60_000, // ms
    });

The Messages API

    Top-Level System Prompt

    Pass system instructions as dedicated string parameter outside messages.

    const res = await client.messages.create({
      model: SONNET,
      max_tokens: 500,
      system: 'Respond strictly in bullet points.',
      messages: [{ role: 'user', content: 'Explain DNS' }],
    });
    Message Content Blocks

    Extract text from response content block array safely.

    const block = res.content[0];
    if (block.type === 'text') {
      console.log(block.text);
    }
    Multi-Turn Dialogue

    Preserve prior assistant turns to maintain contextual continuity.

    const history = [
      { role: 'user', content: 'Hi' },
      { role: 'assistant', content: 'Hello!' },
      { role: 'user', content: 'Summarize chat' },
    ];

Stop Reasons And Flow Control

    End Turn Detection

    Confirm message completed naturally without hitting token boundaries.

    if (res.stop_reason === 'end_turn') {
      // Normal completion finished
    }
    Truncation Handling

    Detect when output reached max tokens and trigger continuation.

    if (res.stop_reason === 'max_tokens') {
      console.warn('Response truncated at token limit');
    }
    Custom Stop Sequences

    Halt generation immediately when encountering a sentinel string.

    const res = await client.messages.create({
      model: SONNET,
      max_tokens: 200,
      stop_sequences: ['===END==='],
      messages: [{ role: 'user', content: 'Generate' }],
    });

Token Usage And Latency

    Usage Breakdown

    Inspect input and output token counts for cost accounting.

    const inTok = res.usage.input_tokens;
    const outTok = res.usage.output_tokens;
    console.log(`Tokens: ${inTok} in / ${outTok} out`);
    Cache Read Metrics

    Monitor prompt cache hits and creation write tokens.

    const u = res.usage;
    const read = u.cache_read_input_tokens ?? 0;
    const write = u.cache_creation_input_tokens ?? 0;
    // Cache reads cost 90% less
    Model Switching Logic

    Route requests dynamically based on prompt complexity.

    const model = prompt.length > 2000 ? SONNET : HAIKU;
    const out = await client.messages.create(
      { model, ...opts }
    );

Tips

  1. Use the exact model ID string, such as claude-sonnet-5-5, and keep it in one constant. The IDs are complete as-is, so never add a date suffix.
  2. Supply system instructions via the top-level system property rather than nesting them inside the message role array.

Warnings

  1. Always define an explicit max_tokens parameter because Anthropic requests fail validation if this property is omitted.
  2. Check for max_tokens stop reasons to identify when model responses have been cut off mid-sentence.

In Practice

FAQ