Advanced Reasoning & Coding Rate Card
Anthropic Claude API Pricing Sheet & Prompt Caching Rates 2026
Anthropic offers state-of-the-art developer models with industry-leading coding benchmark performance and an aggressive Prompt Caching system delivering up to 90% cost reduction on cached context reads.
Anthropic Claude 3.7 & 3.5 Family Token Pricing
| Model Family | Input Rate / 1M | Cache Write / 1M | Cache Read (90% Off) | Output Rate / 1M | Context Limit |
|---|---|---|---|---|---|
| Claude 3.7 Sonnet (Hybrid) | $3.00 | $3.75 (25% extra) | $0.30 (90% off) | $15.00 | 200,000 tokens |
| Claude 3.5 Sonnet (Coding) | $3.00 | $3.75 (25% extra) | $0.30 (90% off) | $15.00 | 200,000 tokens |
| Claude 3.5 Haiku (Fast) | $0.80 | $1.00 (25% extra) | $0.08 (90% off) | $4.00 | 200,000 tokens |
| Claude 3 Opus (High-IQ) | $15.00 | $18.75 (25% extra) | $1.50 (90% off) | $75.00 | 200,000 tokens |
| Claude 3 Haiku (Legacy) | $0.25 | $0.30 (25% extra) | $0.025 (90% off) | $1.25 | 200,000 tokens |
Official Developer SDK / API Integration Example
Python SDK
# Anthropic Claude 3.7 Sonnet Python SDK with Ephemeral Prompt Caching
import anthropic
client = anthropic.Anthropic(api_key="sk-ant-...")
response = client.messages.create(
model="claude-3-7-sonnet-20250219",
max_tokens=4096,
system=[{
"type": "text",
"text": "Enterprise documentation context (10,000+ tokens)...",
"cache_control": {"type": "ephemeral"} # 90% read discount ($0.30/M)
}],
messages=[{"role": "user", "content": "Summarize key endpoints."}]
)
print(response.content[0].text)
Frequently Asked Billing & Pricing Questions
How does Anthropic Prompt Caching achieve 90% savings?
When you add the cache_control breakpoint header to your prompt, Anthropic stores the prompt segment in memory. Writing the cache costs 25% more ($3.75/M on Sonnet), but every subsequent cache read costs only $0.30/M (90% discount).
What is the 5-minute TTL on Anthropic Cache?
Anthropic's prompt cache expires after 5 minutes of inactivity. Every time a cached request arrives, the 5-minute timer resets. For applications with continuous traffic, the cache stays hot indefinitely.
How does Claude 3.7 hybrid thinking affect output tokens?
Claude 3.7 allows developers to configure a specific reasoning token budget. Thinking tokens are billed at the standard output rate ($15.00/M), giving you complete control over spend.