Cut Your Claude API Bill by 90% with Prompt Caching and Batches
Measure real token costs, then stack prompt caching and the Batches API to slash them.
What you'll build
You'll take a working Claude API integration, measure exactly what its input tokens cost, then cut that bill twice: prompt caching bills repeated context at 10% of the normal input rate, and the Message Batches API takes 50% off everything that doesn't need an immediate response. Stacked, a cached token inside a batch costs 5% of the list price.
Prerequisites
- Python 3.9+ (tested on 3.13)
anthropicPython SDK 0.120.2 (current as of August 2026)- An API key from the Claude Console
- Model:
claude-opus-5— $5 per million input tokens, $25 per million output tokens. Cache pricing verified against the official docs: writes cost 1.25× base input (2× for the 1-hour TTL), reads cost 0.1×.
mkdir claude-costs && cd claude-costs
python3 -m venv .venv && source .venv/bin/activate
pip install anthropic==0.120.2
export ANTHROPIC_API_KEY="sk-ant-..."
Step 1: Capture a cost baseline
Most real integrations re-send a big static block on every request — product docs, a style guide, few-shot examples. That's the money leak. Save this as baseline.py:
import anthropic
client = anthropic.Anthropic()
# Stand-in for the static context your app re-sends on every request
DOC_LINE = (
"Acme Widget API reference, section {}: the flux endpoint accepts JSON "
"payloads with a `power` field measured in gigawatts and returns a signed "
"delivery manifest within 30 seconds.\n"
)
SYSTEM = "Answer questions using only the reference below.\n\n" + "".join(
DOC_LINE.format(i) for i in range(200)
)
resp = client.messages.create(
model="claude-opus-5",
max_tokens=1024,
system=SYSTEM,
messages=[{"role": "user", "content": "What unit is the `power` field measured in?"}],
)
u = resp.usage
print(f"input={u.input_tokens} cache_write={u.cache_creation_input_tokens} cache_read={u.cache_read_input_tokens}")
Run python baseline.py. The system prompt is roughly 8,000 tokens, and input_tokens confirms you're paying full price for all of them — about $0.04 per request, $40 per 1,000 requests, before output tokens.
Step 2: Add a cache breakpoint
Caching is a prefix match: the API renders tools → system → messages and caches everything up to a cache_control marker. Any byte change earlier in the prefix invalidates it, so keep stable content first and volatile content (the user's question) after the breakpoint. Save as cached.py:
import anthropic
client = anthropic.Anthropic()
DOC_LINE = (
"Acme Widget API reference, section {}: the flux endpoint accepts JSON "
"payloads with a `power` field measured in gigawatts and returns a signed "
"delivery manifest within 30 seconds.\n"
)
SYSTEM = "Answer questions using only the reference below.\n\n" + "".join(
DOC_LINE.format(i) for i in range(200)
)
def ask(question: str) -> None:
resp = client.messages.create(
model="claude-opus-5",
max_tokens=1024,
system=[
{
"type": "text",
"text": SYSTEM,
"cache_control": {"type": "ephemeral"}, # 5-minute TTL
}
],
messages=[{"role": "user", "content": question}],
)
u = resp.usage
print(f"input={u.input_tokens} cache_write={u.cache_creation_input_tokens} cache_read={u.cache_read_input_tokens}")
ask("What unit is the `power` field measured in?")
ask("How fast is the delivery manifest returned?")
The first call writes the cache at 1.25× input price; every identical-prefix call within the 5-minute TTL (each read refreshes it) bills those tokens at 0.1× — the 90% cut. You get up to 4 breakpoints per request, and on claude-opus-5 a prefix must be at least 512 tokens to cache at all.
Step 3: Batch the non-urgent work
Evaluations, backfills, and nightly summarization don't need answers in seconds. The Batches API takes any list of Messages requests (up to 100,000 requests or 256 MB per batch), processes them asynchronously — most batches finish within an hour — and bills all usage at 50% of standard prices. Caching works inside batches too; use the 1-hour TTL since processing usually outlives the 5-minute one. Save as batch.py:
import time
import anthropic
from anthropic.types.message_create_params import MessageCreateParamsNonStreaming
from anthropic.types.messages.batch_create_params import Request
client = anthropic.Anthropic()
DOC_LINE = (
"Acme Widget API reference, section {}: the flux endpoint accepts JSON "
"payloads with a `power` field measured in gigawatts and returns a signed "
"delivery manifest within 30 seconds.\n"
)
SYSTEM = "Answer questions using only the reference below.\n\n" + "".join(
DOC_LINE.format(i) for i in range(200)
)
shared_system = [
{
"type": "text",
"text": SYSTEM,
"cache_control": {"type": "ephemeral", "ttl": "1h"},
}
]
questions = [
"What unit is the `power` field measured in?",
"What does the flux endpoint return?",
"Summarize the reference in one sentence.",
]
batch = client.messages.batches.create(
requests=[
Request(
custom_id=f"q-{i}",
params=MessageCreateParamsNonStreaming(
model="claude-opus-5",
max_tokens=1024,
system=shared_system,
messages=[{"role": "user", "content": q}],
),
)
for i, q in enumerate(questions)
]
)
print(f"Created {batch.id}, status: {batch.processing_status}")
while batch.processing_status != "ended":
time.sleep(30)
batch = client.messages.batches.retrieve(batch.id)
print(f"status: {batch.processing_status}")
for result in client.messages.batches.results(batch.id):
if result.result.type == "succeeded":
text = next(b.text for b in result.result.message.content if b.type == "text")
print(f"{result.custom_id}: {text}")
else:
print(f"{result.custom_id}: {result.result.type}")
Results arrive in any order, so always key on custom_id, never on position. Results stay downloadable for 29 days.
Verify it works
python cached.py should print (token counts vary by a few):
input=15 cache_write=8021 cache_read=0
input=14 cache_write=0 cache_read=8021
The second line is the win: 8,021 tokens billed at $0.50/MTok instead of $5. python batch.py should end with:
Created msgbatch_013Zva2CMHLNnXjNJJKqJ2EF, status: in_progress
status: ended
q-0: The `power` field is measured in gigawatts.
q-1: The flux endpoint returns a signed delivery manifest.
q-2: ...
At 1,000 requests against this 8,000-token prompt, input costs drop from ~$40 uncached to ~$4 with caching, and ~$2 cached inside a batch — with output tokens also half price in the batch.
Troubleshooting
cache_read_input_tokens is 0 on every request. Caching fails silently, never with an error. Two usual causes: the prefix is under the 512-token minimum for claude-opus-5 (other models require 1,024–4,096), or a byte differs between requests — a timestamp interpolated into the system prompt, an unsorted json.dumps, a tool list that changes order. Diff the rendered prompts of two consecutive requests.
anthropic.AuthenticationError: Error code: 401 - {'type': 'error', 'error': {'type': 'authentication_error', 'message': 'invalid x-api-key'}} — the key isn't reaching the client. Re-run export ANTHROPIC_API_KEY=... in the same shell as the script, and check for a stale key in .env or your shell profile overriding it.
AttributeError: 'MessageBatchErroredResult' object has no attribute 'message' — you read result.result.message without checking the result type. Only succeeded results carry .message; branch on result.result.type first (errored results carry .error instead).
Batch results come back expired. The batch hit the 24-hour processing ceiling before those requests ran — common under heavy demand or when a batch trips rate limits. Expired requests aren't billed; collect their custom_ids and resubmit just those in a new batch.
Next steps
Point cached.py at your real system prompt and tool definitions — a breakpoint on the last system block caches tools and system together. Audit for silent cache invalidators (timestamps, UUIDs, non-deterministic serialization) before trusting your hit rate. Then look at the token counting endpoint to estimate costs pre-flight, and the rate limits page — batches have their own limits separate from the Messages API.
Sources & further reading
- Prompt caching — platform.claude.com
- Batch processing — platform.claude.com
- Pricing — platform.claude.com
- Anthropic Python SDK — github.com
Rachel has been embedded in the developer tooling ecosystem for nearly eight years, covering everything from IDE wars and package-manager drama to the quiet rise of AI-assisted coding. She has a soft spot for open-source maintainers and an unhealthy number of terminal emulators installed on a single laptop.
Discussion 0
No comments yet
Be the first to weigh in.