Compare the cost of OpenAI, Claude, Gemini, Grok and DeepSeek APIs for your token volumes.
Compare what the same workload costs on OpenAI, Anthropic Claude, Google Gemini, xAI Grok and DeepSeek. Enter tokens per request and requests per day; the table sorts every model from cheapest to most expensive, with prices checked against each provider’s official pricing page.
| Model | Provider | $ / 1M input | $ / 1M cached input | $ / 1M output | Notes |
|---|---|---|---|---|---|
| gpt-6-astra | OpenAI | $5.0 | $0.5 | $25.0 | prompts under 270K tokens |
| gpt-6.1-sol | OpenAI | $1.0 | $0.05 | $5.0 | prompts under 270K tokens |
| gpt-6-luna | OpenAI | $0.05 | $0.005 | $0.25 | prompts under 270K tokens |
| gpt-5.3-codex | OpenAI | $1.75 | $0.175 | $14.0 | |
| Claude Fable 5.1 | Anthropic | $10.0 | $0.25 | $50.0 | cache read price shown |
| Claude Opus 5.5 | Anthropic | $4.0 | $0.2 | $20.0 | cache read price shown |
| Claude Sonnet 5.5 | Anthropic | $2.0 | $0.2 | $10.0 | cache read price shown |
| Claude Haiku 4.5 | Anthropic | $1.0 | $0.1 | $5.0 | cache read price shown |
| gemini-3.1-pro-preview | $2.0 | $0.2 | $12.0 | prompts up to 200K tokens | |
| gemini-3.5-flash | $1.5 | $0.15 | $9.0 | ||
| gemini-3.8-flash | $0.75 | $0.075 | $3.75 | price until 31 Dec 2026 | |
| gemini-3.1-flash-lite | $0.25 | $0.025 | $1.5 | ||
| grok-4.7 | xAI | $2.0 | $0.5 | $6.0 | prompts under 200K tokens |
| grok-4.3 | xAI | $1.25 | $0.2 | $2.5 | prompts under 200K tokens |
| deepseek-v4-pro | DeepSeek | $1.32 | $0.044 | $3.96 | peak hours; off-peak half price |
| deepseek-flash | DeepSeek | $0.3 | $0.006 | $1.2 | peak hours; off-peak half price |
Prices checked 2026-10-03 from each provider’s official pricing page (standard tier, text tokens). USD per 1 million tokens, standard (non-batch) tier, text, shortest-context price tier. Copied from each provider's official pricing page on the checked date. Prices change often - always confirm on the provider's page. 1 token ≈ ¾ of an English word; count your prompt exactly with the LLM token counter.
Input (prompt) and output (response) tokens. Count a real prompt with the LLM token counter.
The share of input served from a prompt cache, and requests per day.
Cost per request, per day and per 30 days for every model, cheapest first.
Prices from official provider pages with the date they were checked
Includes cached-input pricing
Filter by provider
Runs in your browser
Measure the real token count of your prompts before estimating cost.
Generating each output token requires a full pass through the model, while input tokens are processed together, so providers charge several times more for output.
If many requests share the same long prefix (instructions, documents), providers can cache it and bill those input tokens at a large discount. Set the cached share to see the effect.
They were copied from the official pricing pages on the date shown under the table. Providers change prices and add tiers (long-context, batch, regional) often, so confirm on their page before budgeting.
Most providers offer about 50 % off for asynchronous batch requests. This calculator uses standard real-time prices.