AI API prices, compared

Every major model's list price per million tokens, grouped by class. Checked September 28, 2026 against official pricing pages.

Frontier models

ModelProviderInput /1MOutput /1MCached in /1M1,000 chats
Grok 4.7xAI$2.00$6.00–$6.00
Gemini 3.1 PropreviewGoogle$2.00$12.00$0.20$9.00
Claude Opus 5.5Anthropic$4.00$20.00$0.20$16.00
Claude Opus 5olderAnthropic$5.00$25.00$0.50$20.00
GPT-5.6 SolOpenAI$5.00$30.00$0.50$22.50
Claude Fable 5.1Anthropic$10.00$50.00$0.25$40.00
GPT-6 AstraOpenAI$10.00$50.00$1.00$40.00

Mid-tier models

ModelProviderInput /1MOutput /1MCached in /1M1,000 chats
Gemini 3.8 FlashGoogle$0.75$3.75$0.075$3.00
Gemini 3.7 FlashGoogle$0.75$3.75$0.075$3.00
Gemini 3.6 FlashGoogle$0.75$3.75$0.075$3.00
DeepSeek V4 ProDeepSeek$1.32$3.96$0.044$3.96
Gemini 3.5 FlasholderGoogle$1.50$9.00$0.15$6.75
Claude Sonnet 5Anthropic$2.00$10.00$0.20$8.00
GPT-6 SolOpenAI$2.00$10.00$0.20$8.00
GPT-5.6 TerraOpenAI$2.00$12.00$0.20$9.00
Claude Sonnet 4.6olderAnthropic$3.00$15.00$0.30$12.00

Budget models

ModelProviderInput /1MOutput /1MCached in /1M1,000 chats
GPT-6 LunaOpenAI$0.10$0.50$0.01$0.40
GPT-5.6 LunaOpenAI$0.20$1.20$0.02$0.90
DeepSeek V4.1 FlashDeepSeek$0.30$1.20$0.006$1.05
Mistral LargeMistral$0.50$1.50–$1.50
Gemini 3.5 Flash-LiteGoogle$0.30$2.50$0.03$1.70
Claude Haiku 4.5Anthropic$1.00$5.00$0.10$4.00

“1,000 chats” assumes 1,500 input and 500 output tokens per message, no caching. Standard tier prices; batch and off-peak discounts are noted on each model page.

Frequently asked questions

What is the cheapest AI API right now?

For a typical chat workload, GPT-6 Luna is the cheapest model we track at $0.10 input and $0.50 output per million tokens, followed by GPT-5.6 Luna and DeepSeek V4.1 Flash.

What is the cheapest frontier model?

Among the top models, Grok 4.7 costs the least for a typical chat workload ($6.00 per 1,000 chats). The most expensive is GPT-6 Astra at $40.00.

Which model is cheapest for coding agents?

Coding agents send long prompts and get short answers. Without caching, one 60,000-token task costs least on GPT-6 Luna ($0.0080). With caching the ranking can change: try the calculator with the Coding agent preset.

Why is output more expensive than input?

Generating text takes more compute than reading it, so providers charge 3 to 8 times more per output token. Our pricing guide explains tokens, caching and batch discounts.

How often are these prices updated?

We re-check every price against the provider's official pricing page twice a week. The last check was on September 28, 2026. Changes show up on the news page.