Reasoning & Extended Thinking

Turn on adaptive thinking so the model works through hard problems, and tune how hard it thinks with effort.

TL;DR

  1. Enable reasoning with thinking: { type: 'adaptive' }.
  2. Set depth and cost with output_config.effort, from low to max.
  3. Newer models hide thinking text unless you set display: 'summarized'.

Adaptive Thinking

    Turn It On

    Add the thinking option and the model decides when and how long to think.

    const response = await client.messages.create({
      model: 'claude-opus-5-5',
      max_tokens: 16000,
      thinking: { type: 'adaptive' },
      messages: [{ role: 'user', content: question }],
    });
    Read The Result

    The reply holds thinking blocks and text blocks. Check each block's type.

    for (const block of response.content) {
      if (block.type === 'thinking') {
        console.log(block.thinking);
      }
      if (block.type === 'text') {
        console.log(block.text);
      }
    }
    Show A Summary

    Newer models return empty thinking text unless you ask for a summary.

    thinking: {
      type: 'adaptive',
      display: 'summarized',
    },

Effort Levels

    Set Effort

    Effort lives in output_config and controls thinking depth and total token use.

    output_config: { effort: 'high' },
    // low | medium | high | xhigh | max
    Lower For Simple Work

    Chat, labeling, and quick edits rarely need deep thinking.

    output_config: { effort: 'low' },
    // faster and cheaper for short answers
    Raise For Hard Work

    Use the top levels when correctness matters more than cost.

    output_config: { effort: 'max' },
    // xhigh suits most coding and agent work

When Reasoning Helps

    Good Fits

    Multi-step math, planning, debugging, and tricky logic gain the most.

    Hard: schedule 5 tasks with dependencies
    -> thinking pays for itself.
    Poor Fits

    Lookups, reformatting, and simple extraction do not need it.

    Easy: convert this date to ISO format
    -> use low effort.
    Prompt A Reasoner

    State the goal and constraints. Skip the step-by-step script.

    Weak: "First list options, then score them..."
    Better: "Pick the cheapest plan that
    meets all three requirements."

Keep It Working

    Echo Thinking Blocks

    In a multi-turn chat, send back the full reply, not just the text.

    messages.push({
      role: 'assistant',
      content: response.content,
    });
    Stream Long Runs

    High effort and large max_tokens can run long. Stream to avoid timeouts.

    const stream = client.messages.stream({ ... });
    const response = await stream.finalMessage();
    Older Models

    Haiku 4.5 still uses a fixed budget. It must be under max_tokens and at least 1024.

    model: 'claude-haiku-4-5',
    thinking: { type: 'enabled', budget_tokens: 2000 },

Tips

  1. Set effort explicitly instead of relying on the default, because the default is different on different models.
  2. Give a reasoning model the goal and the rules in your prompt, then let it work out the steps.

Warnings

  1. Thinking tokens are billed as output tokens even when the text is hidden, so high effort on easy requests wastes money.
  2. Do not send budget_tokens to current models. The newest ones reject it with a 400 error, so use adaptive and effort instead.

In Practice

FAQ