Fireworks AI

Fireworks AI pricing runs two meters: per token on serverless, and per GPU-second on on-demand deployments.

Pricing Model:

Pricing Model:

Usage, on two meters, drawn from a prepaid balance

Usage, on two meters, drawn from a prepaid balance

Usage, on two meters, drawn from a prepaid balance

Packaging Model:

Packaging Model:

Freemium, Good / Better / Best (GBB)

Freemium, Good / Better / Best (GBB)

Freemium, Good / Better / Best (GBB)

Credit Model:

Credit Model:

Prepaid, no published expiry or rollover rule

Prepaid, no published expiry or rollover rule

Prepaid, no published expiry or rollover rule

Updated on:

Fireworks AI pricing: the token rate held while the GPU rate doubled

Fireworks AI pricing runs two meters: per token on serverless, and per GPU-second on on-demand deployments. Both draw down a prepaid balance rather than billing in arrears. The token rates sit where the market sits. The GPU rates don't. An H100 hour costs $8.00 today against $4.00 in January 2026, so the meter has moved further than any model you could pick.

Key takeaways

  • GPT-OSS 120B costs $0.15 per million input tokens and $0.60 output, matching Groq, Together AI, Bedrock and Nebius, but cached input runs $0.015, a 90% cut against Groq's 50%.

  • An H100 hour costs $8.00 on-demand, up from $4.00 in January 2026. That's 46% above Together AI's $5.49 dedicated rate and twice its $3.99 cluster rate.

  • Pinning work to a geography costs a flat 1.5x on both meters, on region-restricted deployments and US-only serverless variants alike.

  • Billing went prepaid on 1 July 2026, so there's no overage. A zero balance pauses serverless, deployments and training at once.

Fireworks AI pricing in 2026

Mode

Unit

Rate

Notes

Serverless

Per 1M tokens

GPT-OSS 120B $0.15 / $0.015 cached / $0.60

Three dimensions

Serverless priority

Per 1M tokens

1.2x to 1.5x standard, by model

US-only variants 1.5x

Batch

Per 1M tokens

50% of serverless, input and output


On-demand

Per GPU-hour, by the second

H100 $8.00, B200 $13.00, GB300 $20.00

1.5x if pinned

Managed training

Per 1M training tokens

$0.50 to $40.00 by size and method


Reserved

Custom

Lower GPU-hour prices, about a 1 year term

Sales only

What Fireworks AI actually meters

Fireworks meters tokens on one product and GPU time on the other, and the two never share a line. Serverless splits every request three ways: input, cached input from the prompt cache, and output. Embeddings bill on input alone, $0.008 to $0.10 per million. Models without a named row fall back to a size-based rate covering input and output alike: $0.10 under 4B, $0.20 to 16B, $0.90 above.

On-demand deployments switch the meter to hardware. Charges start when a replica begins accepting requests and run whether or not calls arrive, and each extra GPU adds its full rate. Deployments scale to zero by default, so an idle one costs nothing, though Fireworks says plainly that turning that off means continuous charges.

Two things deserve care. The rate card prints an hourly and a per-minute figure, and the per-minute one rounds up at the third decimal: $0.134 times 60 is $8.04, not the $8.00 printed beside it, so budget from the hourly figure. And geography costs 1.5x on both meters, putting a region-restricted H100 at $12.00 an hour and GLM 5.3 (US) at $2.10 per million input against $1.40.

How credits work

Credits are the only way to pay, and they behave as a balance rather than a wallet with rules. Fireworks moved self-serve accounts to prepaid credits on 1 July 2026, and serverless, deployments and training all draw on the same balance.

Auto Reload buys more when the balance drops to a minimum you choose. With it off, a zero balance pauses usage until you top up. Fireworks publishes no expiry date, no rollover rule and no deduction order, which is defensible with one credit type but leaves a gap for anyone modelling a year out.

What happens when you hit the limit

Fireworks pauses rather than charging you, at three gates. A zero balance with Auto Reload off stops all usage. So does the monthly spend limit, and reaching 100% of it pauses serverless, deployments and training together. Adding credits doesn't raise that limit, and Fireworks devotes a whole FAQ to accounts suspended with credit still on them.

Rate limits bite earlier. Accounts with no payment method or no credits get 10 requests a minute. With both, the account-wide ceiling is a fixed 6,000 RPM that no spend tier raises, though your Tier 1 to Tier 4 spend tier does cap how far the serverless token ceilings climb. Those adaptive limits grow 25% when usage passes half the current limit and reset after 72 quiet hours. On-demand deployments sit outside those ceilings, so a 429 there means busy GPUs.

How Fireworks AI's pricing has changed

Date

Milestone

Source

1 Oct 2026

DeepSeek V4.1 Flash output goes $0.66 to $1.20 per 1M, input $0.22 to $0.30

Vendor

1 Sep 2026

H100 and H200 go $7.00 to $8.00 an hour, B200 $10.00 to $13.00

Vendor

1 Jul 2026

Self-serve accounts move from post-paid invoicing to prepaid credits

Vendor

1 May 2026

H100 and H200 go $6.00 to $7.00, B200 $9.00 to $10.00. A100 leaves the card

Vendor

By 1 Jan 2026

H100 cut $5.80 to $4.00, B200 $11.99 to $9.00. AMD MI300X drops off

Vendor

8 Dec 2025

Model pages start showing cached and uncached input rates separately

Vendor

Flexprice’s Take

Fireworks prices tokens at the market and hardware well above it, so the serverless-to-dedicated move that its own docs recommend has quietly become the expensive one.

The token side is competitive. GPT-OSS 120B matches the $0.15 and $0.60 that Groq, Together AI, Bedrock and Nebius charge, and the $0.015 cached rate beats all of them at 90% off against Groq's 50%. Batch halves input and output alike, with no "selected models" caveat.

Hardware is where it turns. Fireworks' own 2024 post said graduating to on-demand GPUs made sense around 100,000 tokens a minute, when an H100 cost $5.80. At $8.00 and today's rates, GPT-OSS 120B on a 3:1 input to output mix needs roughly 508,000 tokens a minute before the GPU wins, five times that guidance. Together AI's dedicated H100 crosses over near 349,000. The billing mechanics stay clean though: one balance, no overage, spend limits that stop everything at once.

Best For

Serverless token workloads, especially cache-heavy agents on GPT-OSS.

Watch Out For

Pricing a deployment off a GPU rate that has moved four times in a year.

Manish Choudhary

CEO & Co-founder, Flexprice

Charging customers for inference on rates that move this often?

Flexprice tracks cost and margin per model per account, so a vendor repricing doesn't quietly become your loss.

Flexprice’s Take

Fireworks prices tokens at the market and hardware well above it, so the serverless-to-dedicated move that its own docs recommend has quietly become the expensive one.

The token side is competitive. GPT-OSS 120B matches the $0.15 and $0.60 that Groq, Together AI, Bedrock and Nebius charge, and the $0.015 cached rate beats all of them at 90% off against Groq's 50%. Batch halves input and output alike, with no "selected models" caveat.

Hardware is where it turns. Fireworks' own 2024 post said graduating to on-demand GPUs made sense around 100,000 tokens a minute, when an H100 cost $5.80. At $8.00 and today's rates, GPT-OSS 120B on a 3:1 input to output mix needs roughly 508,000 tokens a minute before the GPU wins, five times that guidance. Together AI's dedicated H100 crosses over near 349,000. The billing mechanics stay clean though: one balance, no overage, spend limits that stop everything at once.

Best For

Serverless token workloads, especially cache-heavy agents on GPT-OSS.

Watch Out For

Pricing a deployment off a GPU rate that has moved four times in a year.

Manish Choudhary

CEO & Co-founder, Flexprice

Charging customers for inference on rates that move this often?

Flexprice tracks cost and margin per model per account, so a vendor repricing doesn't quietly become your loss.

Customer
Sentiment Highlights

"I pay as I go on FireWorks.ai for open model inferencing and no matter how much I use this service my monthly bill is between $10 and $40 and much faster than any reasonable home rig."

Pay-as-you-go Fireworks customer, Hacker News, August 2026

"They force you into making a top up (cause otherwise you are by default suspended) then tell you that almost no model is available on serverless and you literally have to rent a GPU (pay more) to use the simplest most basic models and obviously pay them a gazilion dollars."

Mahmoud Hosseinypour, Trustpilot review of Fireworks, July 2026

Frequently Asked Questions

Frequently Asked Questions

How much does Fireworks AI cost?

Does Fireworks AI have a free tier?

Is Fireworks AI cheaper than Together AI?

Do Fireworks AI credits expire?

Launch usage-based billing this week, not next quarter

Launch usage-based billing this week, not next quarter

Get Instant Feedback on Your Pricing | Join the Flexprice Community with 500+ Builders on Slack

Join the Flexprice Community on Slack