Batch Processing & Cost Reduction

Cut AI API costs in half for work nobody is waiting on by submitting it as an asynchronous batch.

TL;DR

  1. Batches cost half price for work nobody is waiting on, using batches.create.
  2. Give every request a custom_id, because results come back in any order.
  3. Poll until processing_status is ended, then read the results and retry only the failures.

When To Batch

    Good Fits

    Bulk work that nobody is waiting on is the ideal case.

    Label 10,000 old reviews
    Score an eval set overnight
    Summarize a document archive
    Poor Fits

    Anything a person is waiting for needs a normal call.

    Chat reply -> normal call
    Live search -> normal call
    The Trade

    You wait longer and pay half.

    50% off all tokens
    Most done in under 1 hour
    Longest wait: 24 hours

Submit A Batch

    Build The Requests

    Each request has a custom_id you choose and the usual message parameters.

    const ask = (t: string) =>
      `Label positive/negative/neutral: ${t}`;
    
    const requests = reviews.map((text, i) => ({
      custom_id: `review-${i}`,
      params: {
        model: 'claude-haiku-4-5',
        max_tokens: 16,
        messages: [{ role: 'user', content: ask(text) }],
      },
    }));
    Create The Batch

    One call queues every request and returns an id to track.

    const batch = await client.messages.batches.create({
      requests,
    });
    console.log(batch.id);
    Mind The Limits

    One batch holds up to 100,000 requests or 256 MB, whichever comes first.

    // bigger job? split it into
    // several batches of requests

Collect Results

    Poll For Completion

    Check the status on a slow timer until the batch has ended.

    const check = () =>
      client.messages.batches.retrieve(batch.id);
    
    let current = await check();
    while (current.processing_status !== 'ended') {
      await new Promise(r => setTimeout(r, 60_000));
      current = await check();
    }
    Match By custom_id

    Stream the results and file each one under its custom_id.

    const labels = new Map<string, string>();
    const stream = await client.messages.batches
      .results(batch.id);
    
    for await (const item of stream) {
      if (item.result.type !== 'succeeded') continue;
      const msg = item.result.message;
      const first = msg.content[0];
      if (first.type === 'text') {
        labels.set(item.custom_id, first.text);
      }
    }
    Retry Failures

    Fix errored requests that were invalid and resubmit expired ones.

    if (item.result.type === 'expired') {
      retryIds.push(item.custom_id);
    }
    if (item.result.type === 'errored') {
      // invalid_request: fix the input first
    }

Cut Cost Further

    Cache Shared Text

    Mark instructions that every request shares. Hits in a batch are not guaranteed.

    cache_control: { type: 'ephemeral' }
    // on the shared system prompt
    Right-Size The Model

    Use a small, fast model for simple labels and keep larger ones for hard cases.

    claude-haiku-4-5 for labels
    bigger model only for hard cases
    Trim The Output

    Cap max_tokens and ask for short answers, since output tokens cost more.

    max_tokens: 16
    // a one-word label is enough

Tips

  1. Try one or two requests with messages.create first, so you catch mistakes before queuing thousands.
  2. Split big jobs across several batches.create calls, since one batch holds up to 100,000 requests or 256 MB.

Warnings

  1. Results are kept for only 29 days, so save the results() output to your own storage once the batch ends.
  2. Never rely on result order, because batches finish in any order. Match each result to its input with custom_id.

In Practice

FAQ