AI API prices, compared
Every major model's list price per million tokens, grouped by class. Checked September 28, 2026 against official pricing pages.
Frontier models
| Model | Input /1M | Output /1M | 1,000 chats |
|---|---|---|---|
| Grok 4.7 | $2.00 | $6.00 | $6.00 |
| Gemini 3.1 Propreview | $2.00 | $12.00 | $9.00 |
| Claude Opus 5.5 | $4.00 | $20.00 | $16.00 |
| Claude Opus 5older | $5.00 | $25.00 | $20.00 |
| GPT-5.6 Sol | $5.00 | $30.00 | $22.50 |
| Claude Fable 5.1 | $10.00 | $50.00 | $40.00 |
| GPT-6 Astra | $10.00 | $50.00 | $40.00 |
Mid-tier models
| Model | Input /1M | Output /1M | 1,000 chats |
|---|---|---|---|
| Gemini 3.8 Flash | $0.75 | $3.75 | $3.00 |
| Gemini 3.7 Flash | $0.75 | $3.75 | $3.00 |
| Gemini 3.6 Flash | $0.75 | $3.75 | $3.00 |
| DeepSeek V4 Pro | $1.32 | $3.96 | $3.96 |
| Gemini 3.5 Flasholder | $1.50 | $9.00 | $6.75 |
| Claude Sonnet 5 | $2.00 | $10.00 | $8.00 |
| GPT-6 Sol | $2.00 | $10.00 | $8.00 |
| GPT-5.6 Terra | $2.00 | $12.00 | $9.00 |
| Claude Sonnet 4.6older | $3.00 | $15.00 | $12.00 |
Budget models
| Model | Input /1M | Output /1M | 1,000 chats |
|---|---|---|---|
| GPT-6 Luna | $0.10 | $0.50 | $0.40 |
| GPT-5.6 Luna | $0.20 | $1.20 | $0.90 |
| DeepSeek V4.1 Flash | $0.30 | $1.20 | $1.05 |
| Mistral Large | $0.50 | $1.50 | $1.50 |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | $1.70 |
| Claude Haiku 4.5 | $1.00 | $5.00 | $4.00 |
“1,000 chats” assumes 1,500 input and 500 output tokens per message, no caching. Standard tier prices; batch and off-peak discounts are noted on each model page.
Frequently asked questions
What is the cheapest AI API right now?
For a typical chat workload, GPT-6 Luna is the cheapest model we track at $0.10 input and $0.50 output per million tokens, followed by GPT-5.6 Luna and DeepSeek V4.1 Flash.
What is the cheapest frontier model?
Among the top models, Grok 4.7 costs the least for a typical chat workload ($6.00 per 1,000 chats). The most expensive is GPT-6 Astra at $40.00.
Which model is cheapest for coding agents?
Coding agents send long prompts and get short answers. Without caching, one 60,000-token task costs least on GPT-6 Luna ($0.0080). With caching the ranking can change: try the calculator with the Coding agent preset.
Why is output more expensive than input?
Generating text takes more compute than reading it, so providers charge 3 to 8 times more per output token. Our pricing guide explains tokens, caching and batch discounts.
How often are these prices updated?
We re-check every price against the provider's official pricing page twice a week. The last check was on September 28, 2026. Changes show up on the news page.