Reasoning & Extended Thinking
Turn on adaptive thinking so the model works through hard problems, and tune how hard it thinks with effort.
TL;DR
- Enable reasoning with
thinking: { type: 'adaptive' }. - Set depth and cost with
output_config.effort, fromlowtomax. - Newer models hide thinking text unless you set
display: 'summarized'.
Adaptive Thinking
Turn It OnAdd the thinking option and the model decides when and how long to think.
const response = await client.messages.create({
model: 'claude-opus-5-5',
max_tokens: 16000,
thinking: { type: 'adaptive' },
messages: [{ role: 'user', content: question }],
});Read The ResultThe reply holds thinking blocks and text blocks. Check each block's type.
for (const block of response.content) {
if (block.type === 'thinking') {
console.log(block.thinking);
}
if (block.type === 'text') {
console.log(block.text);
}
}Show A SummaryNewer models return empty thinking text unless you ask for a summary.
thinking: {
type: 'adaptive',
display: 'summarized',
},Effort Levels
Set EffortEffort lives in output_config and controls thinking depth and total token use.
output_config: { effort: 'high' },
// low | medium | high | xhigh | maxLower For Simple WorkChat, labeling, and quick edits rarely need deep thinking.
output_config: { effort: 'low' },
// faster and cheaper for short answersRaise For Hard WorkUse the top levels when correctness matters more than cost.
output_config: { effort: 'max' },
// xhigh suits most coding and agent workWhen Reasoning Helps
Good FitsMulti-step math, planning, debugging, and tricky logic gain the most.
Hard: schedule 5 tasks with dependencies
-> thinking pays for itself.Poor FitsLookups, reformatting, and simple extraction do not need it.
Easy: convert this date to ISO format
-> use low effort.Prompt A ReasonerState the goal and constraints. Skip the step-by-step script.
Weak: "First list options, then score them..."
Better: "Pick the cheapest plan that
meets all three requirements."Keep It Working
Echo Thinking BlocksIn a multi-turn chat, send back the full reply, not just the text.
messages.push({
role: 'assistant',
content: response.content,
});Stream Long RunsHigh effort and large max_tokens can run long. Stream to avoid timeouts.
const stream = client.messages.stream({ ... });
const response = await stream.finalMessage();Older ModelsHaiku 4.5 still uses a fixed budget. It must be under max_tokens and at least 1024.
model: 'claude-haiku-4-5',
thinking: { type: 'enabled', budget_tokens: 2000 },Tips
- Set
effortexplicitly instead of relying on the default, because the default is different on different models. - Give a reasoning model the goal and the rules in your
prompt, then let it work out the steps.
Warnings
- Thinking tokens are billed as output tokens even when the text is hidden, so high
efforton easy requests wastes money. - Do not send
budget_tokensto current models. The newest ones reject it with a 400 error, so useadaptiveandeffortinstead.
In Practice
Ask a multi-step question with high effort, stream the response, and print the summary of the reasoning next to the final answer.
- Turn on adaptive thinking and ask for a summary of the reasoning.
- Set effort to high because the puzzle has several linked steps.
- Stream the request so a long run does not time out.
- Print the thinking summary and the answer separately.
import Anthropic from '@anthropic-ai/sdk';
const client = new Anthropic();
const question =
'3 eggs make 12 muffins. Eggs come in cartons of 6. ' +
'Fewest cartons for 50 muffins?';
const stream = client.messages.stream({
model: 'claude-opus-5-5',
max_tokens: 16000,
thinking: { type: 'adaptive', display: 'summarized' },
output_config: { effort: 'high' },
messages: [{ role: 'user', content: question }],
});
const response = await stream.finalMessage();
for (const block of response.content) {
if (block.type === 'thinking') console.log(block.thinking);
else if (block.type === 'text') console.log(block.text);
}FAQ
It still works, but reasoning models already work through a problem internally before they answer. Asking for the answer plus a short explanation usually beats forcing a fixed list of steps, and it avoids paying for wasted tokens.
You can get a summary, not the raw chain of thought. Set display: 'summarized' on the thinking option. On newer models the default returns empty thinking text. Either way you are billed for the thinking that happened.
Start at high for anything that needs real reasoning, and drop to low for chat, classification, and short answers where thinking adds little. Use xhigh or max only when a wrong answer is costly. Test on real requests before you change a default.
Current models reject a fixed thinking budget, and some reject turning thinking off entirely. Switch to adaptive, lower the effort to save cost, and check the model's notes before copying older code.