Batch Processing & Cost Reduction
Cut AI API costs in half for work nobody is waiting on by submitting it as an asynchronous batch.
TL;DR
- Batches cost half price for work nobody is waiting on, using
batches.create. - Give every request a
custom_id, because results come back in any order. - Poll until
processing_statusisended, then read the results and retry only the failures.
When To Batch
Good FitsBulk work that nobody is waiting on is the ideal case.
Label 10,000 old reviews
Score an eval set overnight
Summarize a document archivePoor FitsAnything a person is waiting for needs a normal call.
Chat reply -> normal call
Live search -> normal callThe TradeYou wait longer and pay half.
50% off all tokens
Most done in under 1 hour
Longest wait: 24 hoursSubmit A Batch
Build The RequestsEach request has a custom_id you choose and the usual message parameters.
const ask = (t: string) =>
`Label positive/negative/neutral: ${t}`;
const requests = reviews.map((text, i) => ({
custom_id: `review-${i}`,
params: {
model: 'claude-haiku-4-5',
max_tokens: 16,
messages: [{ role: 'user', content: ask(text) }],
},
}));Create The BatchOne call queues every request and returns an id to track.
const batch = await client.messages.batches.create({
requests,
});
console.log(batch.id);Mind The LimitsOne batch holds up to 100,000 requests or 256 MB, whichever comes first.
// bigger job? split it into
// several batches of requestsCollect Results
Poll For CompletionCheck the status on a slow timer until the batch has ended.
const check = () =>
client.messages.batches.retrieve(batch.id);
let current = await check();
while (current.processing_status !== 'ended') {
await new Promise(r => setTimeout(r, 60_000));
current = await check();
}Match By custom_idStream the results and file each one under its custom_id.
const labels = new Map<string, string>();
const stream = await client.messages.batches
.results(batch.id);
for await (const item of stream) {
if (item.result.type !== 'succeeded') continue;
const msg = item.result.message;
const first = msg.content[0];
if (first.type === 'text') {
labels.set(item.custom_id, first.text);
}
}Retry FailuresFix errored requests that were invalid and resubmit expired ones.
if (item.result.type === 'expired') {
retryIds.push(item.custom_id);
}
if (item.result.type === 'errored') {
// invalid_request: fix the input first
}Cut Cost Further
Cache Shared TextMark instructions that every request shares. Hits in a batch are not guaranteed.
cache_control: { type: 'ephemeral' }
// on the shared system promptRight-Size The ModelUse a small, fast model for simple labels and keep larger ones for hard cases.
claude-haiku-4-5 for labels
bigger model only for hard casesTrim The OutputCap max_tokens and ask for short answers, since output tokens cost more.
max_tokens: 16
// a one-word label is enoughTips
- Try one or two requests with
messages.createfirst, so you catch mistakes before queuing thousands. - Split big jobs across several
batches.createcalls, since one batch holds up to 100,000 requests or 256 MB.
Warnings
- Results are kept for only 29 days, so save the
results()output to your own storage once the batch ends. - Never rely on result order, because batches finish in any order. Match each result to its input with
custom_id.
In Practice
Queue a list of customer reviews as one batch, wait for it to end, and collect a label for each review by its custom_id.
- Turn each review into a request with its own custom_id.
- Create the batch with one call and keep the batch id.
- Poll until processing_status is ended.
- Read each result and match it back by custom_id.
import Anthropic from '@anthropic-ai/sdk';
const client = new Anthropic();
const reviews = ['Loved it!', 'Arrived broken.'];
const { id } = await client.messages.batches.create({
requests: reviews.map((text, i) => ({
custom_id: `review-${i}`,
params: {
model: 'claude-haiku-4-5',
max_tokens: 16,
messages: [{ role: 'user', content: `Label: ${text}` }],
},
})),
});
// poll retrieve(id) until processing_status === 'ended'
for await (const r of await client.messages.batches.results(id)) {
console.log(r.custom_id, r.result.type);
}FAQ
Batch requests are billed at 50% of the normal price for all token usage. The trade is time: you give up an instant reply in exchange for the discount.
Most batches finish within an hour, and the maximum is 24 hours. A request that cannot finish in that window comes back as expired, so resubmit it.
Yes. Each request accepts the same parameters as a normal Messages API call, so tools, images, and prompt caching all work.
Yes. Call client.messages.batches.cancel(id) and the batch moves to canceling. Requests that already finished keep their results.