Baseten

Baseten pricing meters a minute of instance time, and almost every minute a container is alive counts: image builds, cold starts and idle warm replicas all bill.

Pricing Model:

Pricing Model:

Usage, two meters, invoiced in arrears

Usage, two meters, invoiced in arrears

Usage, two meters, invoiced in arrears

Packaging Model:

Packaging Model:

Good / Better / Best (GBB)

Good / Better / Best (GBB)

Good / Better / Best (GBB)

Credit Model:

Credit Model:

Signup grant, applied before the card. No expiry rule

Signup grant, applied before the card. No expiry rule

Signup grant, applied before the card. No expiry rule

Updated on:

Baseten pricing: billed by the minute, cold start included

Baseten pricing meters a minute of instance time, and almost every minute a container is alive counts: image builds, cold starts and idle warm replicas all bill. A second meter charges per token on Model APIs, where GPT-OSS 120B costs $0.10 per million input and $0.50 output, the lowest rate for that model in this index. Those cheap tokens push the GPU decision further out.

Key takeaways

  • GPT-OSS 120B costs $0.10 per million input tokens and $0.50 output, undercutting the $0.15 and $0.60 Groq, Together AI, Fireworks and Bedrock all charge by 33% and 17%.

  • An H100 costs $0.10833 a minute, or $6.50 an hour, between Together AI's $5.49 dedicated rate and Fireworks' $8.00. Multi-GPU nodes scale linearly, with no volume break.

  • Image builds, cold starts and idle warm replicas all bill, and partial minutes round up. A replica that never boots is free, and so is one scaled to zero.

  • The billing API returns a "model distribution surcharge" beside compute cost, and neither the pricing page nor the docs publish its rate.

Baseten pricing in 2026

Meter

Unit

Rate

Notes

Model APIs

Per 1M tokens

GPT-OSS $0.10/$0.50, GLM-5.3 $1.40/$0.14/$4.40

GPT-OSS: no cache

Dedicated GPU

Per minute

H100 $0.10833, A100 $0.06667, B200 $0.16633

$6.50/$4.00/$9.98 hr

Fractional GPU

Per minute

H100 MIG 40 GiB $0.0625

$3.75 an hour

CPU instances

Per minute

$0.00058 to $0.01382

$0.03 to $0.83 hourly

Basic plan

$0 a month

Everything, plus SOC 2 Type II and HIPAA

Pay as you go

Pro and Enterprise

Custom

Priority GPUs, higher limits, self-hosting

Volume discounts

What Baseten actually meters

Baseten meters the minute an instance spends running on a node, and the published per-minute figure is the hourly rate divided by 60, not the reverse: $0.10833 times 60 is $6.4998 against the $6.50 Baseten publishes. Price from the hourly number.

What counts as a running minute is where this gets specific, and Baseten documents it better than anyone else in this index. Image builds bill, because the build is its own workload. Cold starts and model loading bill, because the replica is already up. Idle warm replicas bill whenever min_replica is 1 or higher. Three things don't: image pulls, a failed boot, and a deployment scaled to zero.

The second meter is tokens, billed per million input and output, with cached input at roughly a tenth of input on most models and GPT-OSS carved out entirely. Two things live only in the docs. The instance reference carries an H200 at $0.125 a minute ($7.50 an hour) and an RTX-PRO-6000 at $4.00, neither of which appears on the pricing page. And the billing API returns surcharge_cost, a model distribution surcharge, on every dedicated line, which Baseten's own example puts at exactly 10% of compute.

How credits work

Credits exist, but they're a signup grant rather than a system. New workspaces receive credits for testing, Baseten applies them to the invoice before charging the card, and nothing needs redeeming. The amount isn't published in the docs or on the pricing page.

Past that grant, nothing credit-shaped is left. Baseten states plainly that it offers no separate free tier and no perpetual free plan. There's no expiry rule, no rollover, no top-up pack and no deduction order, because the balance isn't a wallet you refill. Running out means your card pays, and running out with no payment method means Baseten deactivates your models.

What happens when you hit the limit

Baseten charges you and invoices, which makes it the exception among the prepaid vendors here. An invoice issues when usage passes $50 or at month end, whichever comes first.

Budgets carry the sharp edge. A monthly budget emails at 75%, 90% and 100%, and by default it only notifies. Turning on Enforce budget rejects Model API requests once spend reaches it, but it never stops dedicated deployments or training jobs. An enforced budget caps the token meter only.

Rate limits run per account tier: an unverified Basic account gets 15 requests and 100,000 tokens a minute, a verified one 120 and 500,000, Pro 120 and 1,000,000. A 429 means you hit your limit, a 529 means Baseten has no capacity, and the second can happen inside the first.

How Baseten's pricing has changed

Date

Milestone

Source

17 Apr 2026

Cached input billed at a discount on every Model API except GPT-OSS

Vendor

4 Mar 2026

Billing usage API ships, splitting dedicated, Model API and training spend

Vendor

1 Dec 2025

Invoices move to the first of each month

Vendor

21 May 2025

Model APIs launch, adding a per-token meter beside per-minute compute

Vendor

21 Mar 2024

H100 MIG launches at $0.0825 a minute, $4.95 an hour. Now $3.75

Vendor

6 Feb 2024

H100 arrives at $9.984 an hour, A100 at $6.15. Now $6.50 and $4.00

Vendor

1 Jul 2023

Rates cut 40%. A10G goes to $1.207 an hour, unmoved since

Vendor

Flexprice’s Take

Baseten is the clearest writer on what a billable minute is in this whole index, and it sells the cheapest GPT-OSS tokens, which is exactly why its GPUs pay back last.

The lifecycle table is the model other vendors should copy. Most platforms leave you guessing whether a cold start bills. Baseten publishes a nine-row table answering it, counterintuitive rows included: image builds yes, failed boots no, partial minutes rounded up.

The arithmetic follows from the token rates. At $0.10 and $0.50 on GPT-OSS 120B against a $6.50 H100 hour, a 3:1 input to output mix needs roughly 542,000 tokens a minute before renting the GPU beats paying per token. Fireworks crosses near 508,000 and Together AI's dedicated H100 near 349,000. Cheaper tokens push that threshold up, so Baseten rewards staying on Model APIs longest.

A surcharge that shows up in an API field but on no price list belongs on the list.

Best For

Teams who need to forecast container-minute cost precisely before they deploy.

Watch Out For

Treating an enforced budget as a spend cap. It never stops a deployment.

Manish Choudhary

CEO & Co-founder, Flexprice

Reselling inference and need the same lifecycle precision in your own bill?

Flexprice meters per-minute compute and per-token usage on one invoice.

Flexprice’s Take

Baseten is the clearest writer on what a billable minute is in this whole index, and it sells the cheapest GPT-OSS tokens, which is exactly why its GPUs pay back last.

The lifecycle table is the model other vendors should copy. Most platforms leave you guessing whether a cold start bills. Baseten publishes a nine-row table answering it, counterintuitive rows included: image builds yes, failed boots no, partial minutes rounded up.

The arithmetic follows from the token rates. At $0.10 and $0.50 on GPT-OSS 120B against a $6.50 H100 hour, a 3:1 input to output mix needs roughly 542,000 tokens a minute before renting the GPU beats paying per token. Fireworks crosses near 508,000 and Together AI's dedicated H100 near 349,000. Cheaper tokens push that threshold up, so Baseten rewards staying on Model APIs longest.

A surcharge that shows up in an API field but on no price list belongs on the list.

Best For

Teams who need to forecast container-minute cost precisely before they deploy.

Watch Out For

Treating an enforced budget as a spend cap. It never stops a deployment.

Manish Choudhary

CEO & Co-founder, Flexprice

Reselling inference and need the same lifecycle precision in your own bill?

Flexprice meters per-minute compute and per-token usage on one invoice.

Customer
Sentiment Highlights

"Now Baseten's pricing is cheaper than the official one? Probably won't last, but still interesting."

Developer comparing open-weight model hosts, Hacker News, August 2026

"on AA openai gets 117tps. baseten gets 284tps. so 18% more expensive but 142% more tps."

Developer benchmarking price against throughput, Hacker News, September 2026

Frequently Asked Questions

Frequently Asked Questions

How much does Baseten cost per hour?

Does Baseten have a free tier?

Does Baseten charge for cold starts?

Is Baseten cheaper than Fireworks AI?

Launch usage-based billing this week, not next quarter

Launch usage-based billing this week, not next quarter

Get Instant Feedback on Your Pricing | Join the Flexprice Community with 500+ Builders on Slack

Join the Flexprice Community on Slack