AWS Bedrock Agents Pricing, Cost and Break-Even Value
By Paige Gilmore, Founder, NetLift · Published 2026-08-07 · Updated 2026-08-18
Prices checked 2026-08-31 against official vendor pricing pages.
AWS Bedrock Agents pricing is driven by token consumption across various models, ranging from $0.10 to $30 per million tokens, alongside a $1.95 monthly fee for custom model storage.
AWS Bedrock Agents costs scale based on the intelligence level required for the task. Low-latency, high-volume agents using Ministral 3B 3.0 start at $0.10 per million tokens for both input and output. In contrast, complex reasoning agents using Anthropic Claude Opus cost $6 per million input tokens and $30 per million output tokens on-demand.
Organizations using custom models face a monthly storage fee of $1.95 for Meta Llama 2 Pretrained (13B). Financial predictability depends on selecting the appropriate model and utilizing batch processing where possible, such as Claude 3.5 Sonnet, which offers reduced rates compared to public extended access.
Current official pricing
| Plan | Price | Billing | Notes |
|---|---|---|---|
| Meta Llama 2 Pretrained (13B) — Custom Model Storage | $1.95 | per month | Price to store each custom model per month |
Usage-based prices
| Model / rate | Price | Unit |
|---|---|---|
| Anthropic Claude 3.5 Sonnet (Public Extended Access) — Batch Input | $3 | per 1M input tokens |
| Anthropic Claude 3.5 Sonnet (Public Extended Access) — On-Demand Input | $6 | per 1M input tokens |
| Anthropic Claude 3.5 Sonnet (Public Extended Access) — Batch Output | $15 | per 1M output tokens |
| Anthropic Claude 3.5 Sonnet (Public Extended Access) — On-Demand Output | $30 | per 1M output tokens |
| Anthropic Claude 3.5 Sonnet v2 (Public Extended Access) — Cache Read | $0.6 | per 1M input tokens |
| Anthropic Claude 3.5 Sonnet v2 (Public Extended Access) — Cache Write | $7.50 | per 1M input tokens |
| DeepSeek v3.1 — Flex Tier Input (Sydney) | $0.3 | per 1M input tokens |
| DeepSeek v3.1 — Priority Tier Input (Sydney) | $1.05 | per 1M input tokens |
| DeepSeek v3.2 — On-Demand Standard Input (US) | $0.62 | per 1M input tokens |
| DeepSeek v3.2 — On-Demand Standard Output (US) | $1.85 | per 1M output tokens |
| Google Gemma 4 31B — On-Demand Standard Input (US) | $0.14 | per 1M input tokens |
| Google Gemma 4 31B — On-Demand Standard Output (US) | $0.4 | per 1M output tokens |
| Meta Llama 2 Chat (13B) — On-Demand Input | $0.75 | per 1M input tokens |
| MiniMax M2 — On-Demand Standard Input (US) | $0.3 | per 1M input tokens |
| Mistral Large 3 — On-Demand Standard Input (US) | $0.5 | per 1M input tokens |
| Mistral Large 3 — On-Demand Standard Output (US) | $1.50 | per 1M output tokens |
| Moonshot Kimi K2 Thinking — On-Demand Standard Input (US) | $0.6 | per 1M input tokens |
| Moonshot Kimi K2.5 — On-Demand Standard Output (US) | $3 | per 1M output tokens |
Which models offer the lowest operational overhead?
For high-frequency tasks where speed is more critical than deep reasoning, Ministral 3B 3.0 is the most cost-effective at $0.10 per million tokens. Google Gemma 4 31B follows closely at $0.14 for inputs and $0.40 for outputs. These models allow for wide deployment of agents without the aggressive cost scaling seen in larger frontier models.
How do reasoning requirements impact the budget?
Agents designed for debugging or complex code generation typically require higher-parameter models. Mistral Large 3 offers a mid-tier price point at $0.50 per million input tokens and $1.50 per million output tokens. At the highest end, Anthropic Claude Opus and Claude 3.5 Sonnet Public Extended Access both reach $30 per million output tokens, requiring a much higher threshold for business value to justify the spend.
To determine if an agent provides a positive return, you must compare the total token spend against the cost of the manual labor it replaces. NetLift allows organizations to move beyond simple cost tracking to measure whether high-priced models like Claude Opus deliver enough time savings to outperform cheaper alternatives. By quantifying the net return of AI adoption, leaders can justify the premium for advanced reasoning based on objective evidence of efficiency gains.
Sources
All pricing on this page comes from official vendor pages:
- https://aws.amazon.com/bedrock/pricing/ — retrieved 2026-08-11
Frequently asked questions
What is the cost for agents that do read-only debugging in production?
The cost depends on the model selected for reasoning. For example, Mistral Large 3 costs $0.50 per 1M input tokens and $1.50 per 1M output tokens, while Anthropic Claude 3.5 Sonnet Public Extended Access costs $6 per 1M input tokens and $30 per 1M output tokens.
How much does custom model storage cost for Llama 2?
AWS Bedrock charges $1.95 per month for Meta Llama 2 Pretrained (13B) Custom Model Storage.
What are the cheapest models for computer-use agents or browser harnesses?
Ministral 3B 3.0 is the most affordable at $0.10 per 1M tokens for both input and output. Google Gemma 4 31B is another low-cost option at $0.14 per 1M input tokens and $0.40 per 1M output tokens.
Is there a discount for batch input on high-performance models?
Yes. Anthropic Claude 3.5 Sonnet Batch Input is priced at $3 per 1M tokens, which is 50% lower than the Public Extended Access rate of $6 per 1M tokens.
About the author
Paige Gilmore is the founder of NetLift, the AI Value Management platform that helps organisations measure the cost, savings and return of AI adoption. Paige Gilmore on LinkedIn