How much will your LLM API bill be?
Estimate and compare the monthly cost of running GPT, Claude, Gemini, Grok, DeepSeek, and Mistral for your workload — then see how much prompt compression could shave off. Prices are per 1M tokens, as of August 2026.
| Model | $/1M in | $/1M out | Cost / request | Monthly | |
|---|---|---|---|---|---|
Mistral Small 4Budget Mistral | $0.15 | $0.60 | 0.168¢ | $84.00 | Details |
DeepSeek V4 FlashBudget DeepSeek | $0.22 | $0.66 | 0.229¢ | $114.40 | Details |
GPT-5.6 LunaBudget OpenAI | $0.20 | $1.20 | 0.256¢ | $128.00 | Details |
Mistral Large 3Budget Mistral | $0.50 | $1.50 | 0.520¢ | $260.00 | Details |
DeepSeek V4 ProBudget DeepSeek | $0.66 | $1.98 | 0.686¢ | $343.20 | Details |
Gemini 3.7 FlashBalanced Google | $0.75 | $3.75 | 0.900¢ | $450.00 | Details |
Claude Haiku 4.5Budget Anthropic | $1.00 | $5.00 | $0.01 | $600.00 | Details |
GPT-5Frontier OpenAI | $1.25 | $10.00 | $0.02 | $900.00 | Details |
Grok 4.6Balanced xAI | $2.00 | $6.00 | $0.02 | $1,040 | Details |
Claude Sonnet 5Balanced Anthropic | $2.00 | $10.00 | $0.02 | $1,200 | Details |
GPT-5.6 TerraBalanced OpenAI | $2.00 | $12.00 | $0.03 | $1,280 | Details |
Gemini 3.1 ProFrontier Google | $2.00 | $12.00 | $0.03 | $1,280 | Details |
Claude Opus 5Frontier Anthropic | $5.00 | $25.00 | $0.06 | $3,000 | Details |
Estimates only. Prices are standard (non-cached, non-batch) per-1M-token rates as of August 2026 and change frequently — verify with each provider. Batch and cached-input discounts can cut these further.
Cut the bill with prompt compression
Most of an LLM bill is input tokens — long chat histories, retrieved docs, and tool traces. Prompt compression removes low-value context before the call. See what a reduction does to your input spend.
Picking between these models?
Compare them side by side on features, pricing, and real pros & cons.
How the LLM cost calculator works
Every API call bills separately for the tokens you send (input) and the tokens the model generates (output), each priced per million tokens. This calculator multiplies your input and output tokens per request by each model's rate, then by your monthly request volume, so you can compare providers on the exact workload you run. A token is roughly 4 characters, or about ¾ of a word.
Why is my output so much more expensive than input?
Output tokens usually cost 4–8× more than input tokens across providers. If your app generates long answers, output dominates the bill; if it sends huge context (RAG, long chats), input dominates — which is where prompt compression helps most.
Are these prices current?
They reflect standard published rates as of August 2026. LLM pricing moves fast — providers cut prices and launch new tiers regularly — so always confirm on the provider's pricing page before budgeting. Batch APIs and cached-input reads can lower costs well below the figures shown here.