What 1,000 AI API requests cost on 17 models (September 2026)
Per-million-token prices are hard to picture. This page turns them into something concrete: what 1,000 requests cost on every current model we track, for three common kinds of work. Every number comes from the provider’s official pricing page (linked below) and is calculated the same way for every model.
The three workloads
- Chat: 2,000 input tokens (system prompt, history, question) and 500 output tokens. A typical chatbot or support reply.
- Coding agent step: 20,000 input tokens, of which 18,000 are a repeated context read from the prompt cache, plus 1,000 output tokens. One step of an agent that re-reads the same codebase each turn.
- Summary: 10,000 input tokens (a long document) and 300 output tokens.
The formula is simply tokens × price per million, times 1,000 requests. For the agent column, the 18,000 cached tokens are charged at the model’s cache read price.
Results
Sorted by the chat column, cheapest first. Prices are per million tokens (input / output).
| Model | Price in / out | 1,000 chats | 1,000 agent steps | 1,000 summaries |
|---|---|---|---|---|
| GPT-6 Luna | $0.10 / $0.50 | $0.45 | $0.88 | $1.15 |
| GPT-5.6 Luna | $0.20 / $1.20 | $1.00 | $1.96 | $2.36 |
| DeepSeek V4.1 Flash | $0.30 / $1.20 | $1.20 | $1.91 | $3.36 |
| Mistral Large | $0.50 / $1.50 | $1.75 | $11.50* | $5.45 |
| Gemini 3.5 Flash-Lite | $0.30 / $2.50 | $1.85 | $3.64 | $3.75 |
| Gemini 3.8 Flash† | $0.75 / $3.75 | $3.38 | $6.60 | $8.63 |
| Claude Haiku 4.5 | $1 / $5 | $4.50 | $8.80 | $11.50 |
| DeepSeek V4 Pro | $1.32 / $3.96 | $4.62 | $7.39 | $14.39 |
| Grok 4.7 | $2 / $6 | $7.00 | $46.00* | $21.80 |
| Claude Sonnet 5 | $2 / $10 | $9.00 | $17.60 | $23.00 |
| GPT-6 Sol | $2 / $10 | $9.00 | $17.60 | $23.00 |
| GPT-5.6 Terra | $2 / $12 | $10.00 | $19.60 | $23.60 |
| Gemini 3.1 Pro (preview) | $2 / $12 | $10.00 | $19.60 | $23.60 |
| Claude Opus 5.5 | $4 / $20 | $18.00 | $31.60 | $46.00 |
| GPT-5.6 Sol‡ | $5 / $30 | $25.00 | $49.00 | $59.00 |
| Claude Fable 5.1 | $10 / $50 | $45.00 | $74.50 | $115.00 |
| GPT-6 Astra | $10 / $50 | $45.00 | $88.00 | $115.00 |
* The xAI models page does not list a cache read price for Grok 4.7, and Mistral’s pricing page only says cached input is “up to 90%” cheaper without a per-model figure. For those two, the agent column charges the cached tokens at the full input price, so it is an upper bound.
† Gemini 3.8 Flash prices are promotional through December 31, 2026. From January 1, 2027 Google lists $1.50 input, $7.50 output and $0.15 cache reads, which doubles its row.
‡ OpenAI’s main pricing page lists GPT-5.6 Sol at $5 / $30 (used here), while its developer pricing page shows $4 / $20. At $4 / $20 its row would be $18.00, $35.20 and $46.00.
What stands out
The spread is about 100x. The same 1,000 chats cost $0.45 on GPT-6 Luna and $45 on GPT-6 Astra or Claude Fable 5.1. For classification, extraction or simple replies, a budget model is usually the right starting point.
Output price drives the chat column. Even though a chat sends four times more tokens than it receives, output is priced about 3 to 8 times higher, so it is often half the bill or more. Capping reply length is one of the easiest savings.
Caching reorders the agent column. In an agent step most input is cached, so the cache read price matters more than the headline input price. Claude Opus 5.5 lists input at $4 but cache reads at $0.20 per million (95% off), which keeps its agent step at $31.60. Among the flagships, GPT-6 Astra’s cache read of $1 per million makes it the most expensive agent step at $88, compared with $74.50 for Claude Fable 5.1 at the same headline price. Models without a published cache price lose most of their advantage here, which is why Grok 4.7 jumps from 9th cheapest for chat to among the most expensive agent steps in this table.
Same price, different token count. Anthropic says Claude 4.7 and later models use a newer tokenizer that produces about 30% more tokens for the same text than earlier Claude models. So Claude Sonnet 5 and GPT-6 Sol, both $2 / $10, may not cost the same for the same prompt. Test with your own text before deciding.
Ways to pay less than the table
- Batch. Anthropic, OpenAI and Google offer a Batch API at 50% off if you can wait for results, and Mistral lists the same 50% batch discount. For those models, every figure above would halve. We have not found a published batch discount for DeepSeek or Grok.
- Off-peak DeepSeek. DeepSeek’s prices above are peak prices. Peak hours are 01:00–04:00 and 06:00–10:00 UTC on weekdays; at all other times, including weekends and Chinese public holidays, it charges half. See AI API pricing explained for the details.
- Long context costs more on some models. Gemini 3.1 Pro doubles its input price above 200K tokens. None of the workloads here come close, but a large agent context can.
Run your own numbers
These three workloads are only examples. The cost calculator lets you enter your own token counts, cache share and batch setting for any model, and gives you a link you can share. For a head-to-head view, use the model comparisons, and if a model you depend on seems off today, check the Dumb meter to see whether other people are noticing the same thing.
Prices change often. We recheck each provider’s official pricing page regularly and update this page when they change.