AWS Bedrock

AWS Bedrock pricing charges per token by default, then offers three other ways to buy the same model.

Pricing Model:

Pricing Model:

Usage-based per token, plus reserved and provisioned capacity

Usage-based per token, plus reserved and provisioned capacity

Usage-based per token, plus reserved and provisioned capacity

Packaging Model:

Packaging Model:

Good / Better / Best (GBB)

Good / Better / Best (GBB)

Good / Better / Best (GBB)

Credit Model:

Credit Model:

None

None

None

Updated on:

AWS Bedrock pricing: four tiers, two provisioned modes, one undocumented unit

AWS Bedrock pricing charges per token by default, then offers three other ways to buy the same model. A service_tier parameter per request sets both the rate and the queue position. Reserved capacity sits outside that, billed monthly against tokens-per-minute you commit to in advance. The older Provisioned Throughput mode still exists and bills hourly per Model Unit.

Key takeaways

  • Bedrock exposes four service tiers on one API call: Reserved, Priority, Standard and Flex, selected with a service_tier parameter set to reserved, priority, default or flex.

  • Reserved capacity meters tokens-per-minute, not compute slices, and you size input and output TPM separately with minimums of 100,000 input TPM and 10,000 output TPM.

  • Batch inference runs at 50% below on-demand rates on selected models, and geographic and in-region routing costs 10% more than global cross-Region inference on current models.

  • AWS publishes hourly Provisioned Throughput rates for a handful of older Cohere and Meta models and tells everyone else to call their account team.

AWS Bedrock pricing in 2026

Mode

Billing unit

Commitment

Published rate

 

Standard

Per token, per model

None

Yes, per model per Region

Priority

Per token at a premium

None

Yes, per model

Flex

Per token at a discount

None

Yes, per model

Batch

Per token, 50% off on-demand

None

Yes, per model

Reserved

Per hour per 1K input and 1K output TPM, billed monthly

1 or 3 months

Yes, per model per Region

Provisioned

Per hour per Model Unit

None, 1 or 6 months

Partial. Cohere Command $49.50 / $39.60 / $23.77

Model

Input /1M

Output /1M

Batch in/out

---

---

---

---

Claude Opus 5

$5.50

$27.50

$2.75 / $13.75

Claude Sonnet 5

$2.20

$11.00

n/a

Source: aws.amazon.com/bedrock/pricing, read 23 September 2026.

What AWS Bedrock pricing actually meters

Three different units, depending on how you buy.

On demand, the unit is the token, priced per million, per model, per Region and per tier. Priority costs more and jumps the queue. Flex costs less and accepts longer processing. Standard is the default when the parameter is absent. Your on-demand quota is shared across all three, so switching tier changes price and latency but not headroom.

Reserved meters throughput instead. You commit to a number of input tokens-per-minute and a number of output tokens-per-minute, priced per hour per 1,000 TPM and billed monthly. AWS sets floors of 100,000 input TPM and 10,000 output TPM, and warns that your consumption counts both InputTokenCount and CacheWriteInputTokens, so prompt caching inflates the reservation you need.

Provisioned Throughput meters a Model Unit. The docs define an MU as a level of input and output tokens per minute for one model, then decline to say what that level is, directing you to your AWS account manager. You cannot size a Provisioned Throughput purchase from public documentation. AWS now also sells it by Tokens, with no dated announcement.

The Reserved Tier block publishes rates per Region: $0.33 per hour per 1K input TPM and $1.65 per 1K output TPM on Claude Opus 4.6 at one month, $0.297 and $1.485 at three. Provisioned Throughput still reads "please reach out to your account team". Non-GA Claude Mythos access is gated separately and routes to an Anthropic account manager.

What happens when you hit the limit

Reserved capacity degrades rather than fails. When traffic exceeds what you reserved, Bedrock overflows to the Standard tier automatically and bills those tokens at on-demand rates, which keeps the application up and makes the bill variable at exactly the moment you bought a reservation to avoid variance. The tier targets 99.5% uptime for model response.

On-demand tiers have no spend ceiling. Throttling is governed by token quotas, not budget, so hitting a limit produces a throttling error rather than a charge stop. Provisioned Throughput inverts that: capacity is fixed and billing continues until you delete it. Reserved reservations only stop billing when your AWS account manager removes them, not from the console.

How AWS Bedrock pricing has changed across all these years

Date

Milestone

Source

 

11 August 2026

IAM principal cost allocation extended to the bedrock-mantle endpoint

Vendor

31 December 2025

Model support list for Priority and Flex updated

Vendor

26 November 2025

Reserved service tier added

Vendor

18 November 2025

Priority and Flex service tiers added for on-demand inference

Vendor

16 July 2025

Custom models deployable for on-demand per-token inference without Provisioned Throughput

Vendor

21 August 2024

Batch inference reaches general availability

Vendor

29 March 2024

Provisioned Throughput becomes purchasable for base models with no commitment

Vendor

Flexprice’s Take

Bedrock's Reserved tier is a better capacity product than Azure's PTU, and its Provisioned Throughput is a worse one.

Reserving input and output tokens-per-minute separately is the right unit. It maps to how your workload actually behaves, and when you burst past the reservation Bedrock overflows to Standard instead of dropping requests. Azure's PTU reserves an opaque compute slice and fails when you exceed it.

Per-call tier selection is right too. One service_tier parameter, four price and latency points, and CloudTrail and CloudWatch show you afterwards which one you picked.

The Model Unit hasn't aged. It's the same undocumented slice as a PTU, you file a support case to buy any, and AWS publishes the hourly rate for Cohere Command at $49.50 while Anthropic's says "reach out to your account team". Reserved prices in the open now, and still doesn't cancel without an account manager.

Best For

Mixed workloads that want batch, flex and priority pricing behind one API.

Watch Out For

Reserved overflow quietly reverting to on-demand rates, and caching writes counting against your reserved TPM.

Manish Choudhary

CEO & Co-founder, Flexprice

Reselling Bedrock capacity and need cost and margin per customer, per model?

Flexprice tracks it at that level.

Flexprice’s Take

Bedrock's Reserved tier is a better capacity product than Azure's PTU, and its Provisioned Throughput is a worse one.

Reserving input and output tokens-per-minute separately is the right unit. It maps to how your workload actually behaves, and when you burst past the reservation Bedrock overflows to Standard instead of dropping requests. Azure's PTU reserves an opaque compute slice and fails when you exceed it.

Per-call tier selection is right too. One service_tier parameter, four price and latency points, and CloudTrail and CloudWatch show you afterwards which one you picked.

The Model Unit hasn't aged. It's the same undocumented slice as a PTU, you file a support case to buy any, and AWS publishes the hourly rate for Cohere Command at $49.50 while Anthropic's says "reach out to your account team". Reserved prices in the open now, and still doesn't cancel without an account manager.

Best For

Mixed workloads that want batch, flex and priority pricing behind one API.

Watch Out For

Reserved overflow quietly reverting to on-demand rates, and caching writes counting against your reserved TPM.

Manish Choudhary

CEO & Co-founder, Flexprice

Reselling Bedrock capacity and need cost and margin per customer, per model?

Flexprice tracks it at that level.

Customer
Sentiment Highlights

Llama 3.3 Instruct 70B is 5-10x cheaper than Sonnet 5

Bedrock user on model selection, Hacker News, July 2026

Frequently Asked Questions

Frequently Asked Questions

How much does AWS Bedrock cost?

What is a Model Unit in Bedrock?

Does Bedrock have a free tier?

Is cross-Region inference more expensive?

Launch usage-based billing this week, not next quarter

Launch usage-based billing this week, not next quarter

Get Instant Feedback on Your Pricing | Join the Flexprice Community with 400+ Builders on Slack

Join the Flexprice Community on Slack