Claude API

Claude API pricing charges per million tokens, and it splits them five ways.

Pricing Model:

Pricing Model:

Pure usage

Pure usage

Pure usage

Packaging Model:

Packaging Model:

Good / Better / Best (GBB)

Good / Better / Best (GBB)

Good / Better / Best (GBB)

Credit Model:

Credit Model:

Prepaid. Managed Agents sessions accept a hard spend budget

Prepaid. Managed Agents sessions accept a hard spend budget

Prepaid. Managed Agents sessions accept a hard spend budget

Updated on:

Claude API pricing: the rate card and the modifiers that stack

Claude API pricing charges per million tokens, and it splits them five ways. Base input, five-minute cache writes, one-hour cache writes, cache hits and output each carry their own rate. OpenAI splits the same bill four ways, so the difference is the second cache duration rather than caching itself. Anthropic then layers modifiers on top for speed, data residency and batching, and several of them stack. The rate card is unusually complete, which makes the stacking rules the part worth reading.

Key takeaways

  • Base input runs $1 per million on Haiku 4.5 up to $10 on Fable 5.1, with output at 5x input across the range.

  • Anthropic prices two cache durations separately: a one-hour cache write costs 2x base input, a five-minute write 1.25x.

  • The Batch API discounts input and output 50%, but cannot be combined with Fast mode.

  • Claude Managed Agents bill $0.08 per session-hour on top of tokens, metered to the millisecond and only while running.

Claude API pricing in 2026

Model

Base input /1M

5m cache write

1h cache write

Cache hit

Output /1M

 

Claude Fable 5.1

$10

$12.50

$20

$0.25

$50

Claude Opus 5

$5

$6.25

$10

$0.50

$25

Claude Sonnet 5

$2

$2.50

$4

$0.20

$10

Claude Sonnet 4.6

$3

$3.75

$6

$0.30

$15

Claude Haiku 4.5

$1

$1.25

$2

$0.10

$5

Source: platform.claude.com/docs/en/about-claude/pricing, read 22 September 2026.

What Claude actually meters

Claude meters tokens, and the five-column rate card is the thing that separates it from most model APIs. Writing to a cache costs more than plain input, and Anthropic charges differently depending on how long you want the cache to live: a five-minute write runs 1.25x base input, a one-hour write runs 2x. Reading from cache is where the saving lands, usually at 0.1x base input.

Fable 5.1 goes further, pricing cache reads at $0.25 per million, which Anthropic documents as 0.025x base input rather than the usual 0.1x. On that model, a cache read costs a quarter of what it costs elsewhere in the range.

Two other meters exist beyond tokens. Web search inside a session costs $10 per 1,000 searches. Claude Managed Agents add a session runtime charge of $0.08 per session-hour, measured to the millisecond and accruing only while a session's status is running. Idle time spent waiting for your next message or a tool confirmation does not count, and runtime replaces container-hour billing rather than adding to it.

How the modifiers stack

This is where Claude API pricing gets genuinely complicated, and Anthropic is unusually clear about which combinations are legal.

  • Batch API: 50% off both input and output. Not available with Fast mode, and not available to Managed Agents sessions, which are stateful.

  • Fast mode: research preview on Opus 5 and Opus 4.8, priced at $10 input and $50 output, double the standard rate. First-party API only, not on AWS or partner clouds.

  • Data residency: pinning inference to the US bills at 1.1x standard rates.

  • Prompt caching multipliers: apply on top of Fast mode and on top of data residency.

So a US-pinned Fast mode request with cache writes carries three multipliers at once. Anthropic publishes a worked example rather than leaving you to compose them, which is more than most providers do.

What happens when you hit the limit

There is no included allowance, so there is nothing to exceed. Billing starts with the first call and spend control sits in account settings, not in the plan you chose.

The exception is Claude Managed Agents. Since August 2026 you can set a hard budget on a session, priced at public list rates, and the session stops when it reaches the cap. That is a real spend ceiling inside the product rather than an alert after the fact.

How Claude API pricing has changed

Date

Milestone

Source

 

1 Sep 2026

Fable 5.1 and Mythos 5.1 launch with cache reads at $0.25 per million, 0.025x base input rather than the usual 0.1x

platform.claude.com/docs/en/release-notes/overview

10 Aug 2026

Sonnet 5's introductory $2 and $10 rates become permanent. A previously scheduled increase to $3 and $15 is cancelled

platform.claude.com/docs/en/release-notes/overview

7 Aug 2026

Session budgets added to Managed Agents as a hard spend cap, alongside per-agent inference geography control

platform.claude.com/docs/en/release-notes/overview

Source: platform.claude.com/docs/en/release-notes/overview, read 22 September 2026.

Flexprice’s Take

Claude publishes the most complete pricing documentation of any model provider in this index, and it needs to, because the modifiers multiply.

Pricing two cache TTLs separately is the right call. A one-hour write costs 2x base input and a five-minute write 1.25x, and charging the same for both would subsidise one of them.

Anthropic also tells you which modifiers stack and which combinations don't exist. Batch takes 50% off input and output but won't run with Fast mode, and finding that out on an invoice would be worse.

The complexity is still real. Three stacking multipliers, a $0.08 per session-hour runtime SKU and $10 per 1,000 web searches mean the effective rate for any one call takes work to derive. That's defensible for a platform this configurable, and still more than a finance team can hold in their head.

Cancelling the scheduled Sonnet 5 increase deserves credit.

Best For

Teams that will actually use caching and batching, where the discounts are large.

Watch Out For

Anyone budgeting a Fast mode workload without checking that Batch is excluded.

Manish Choudhary

CEO & Co-founder, Flexprice

Reselling model access and need per-model, per-customer cost tracking?

Flexprice meters it and reports margin by account.

Flexprice’s Take

Claude publishes the most complete pricing documentation of any model provider in this index, and it needs to, because the modifiers multiply.

Pricing two cache TTLs separately is the right call. A one-hour write costs 2x base input and a five-minute write 1.25x, and charging the same for both would subsidise one of them.

Anthropic also tells you which modifiers stack and which combinations don't exist. Batch takes 50% off input and output but won't run with Fast mode, and finding that out on an invoice would be worse.

The complexity is still real. Three stacking multipliers, a $0.08 per session-hour runtime SKU and $10 per 1,000 web searches mean the effective rate for any one call takes work to derive. That's defensible for a platform this configurable, and still more than a finance team can hold in their head.

Cancelling the scheduled Sonnet 5 increase deserves credit.

Best For

Teams that will actually use caching and batching, where the discounts are large.

Watch Out For

Anyone budgeting a Fast mode workload without checking that Batch is excluded.

Manish Choudhary

CEO & Co-founder, Flexprice

Reselling model access and need per-model, per-customer cost tracking?

Flexprice meters it and reports margin by account.

Customer
Sentiment Highlights

the initial cache costs 25% more than usual context.

Tiberium, on manual cache breakpoints, Hacker News, October 2025

"Meaning I'm eating $20k worth of tokens for a $200 sub."

Hacker News, August 2026

Frequently Asked Questions

Frequently Asked Questions

How much does the Claude API cost?

What is the Claude Batch API discount?

How much does Claude prompt caching save?

Does Claude charge for agent session time?

Launch usage-based billing this week, not next quarter

Launch usage-based billing this week, not next quarter

Get Instant Feedback on Your Pricing | Join the Flexprice Community with 400+ Builders on Slack

Join the Flexprice Community on Slack