Guides / Pricing

AI API pricing explained: tokens, input vs output, caching and batch

AI APIs don’t charge per question. They charge per token, and they charge different amounts for what you send and what you get back. Once you understand a few ideas, you can estimate your bill with reasonable accuracy.

1. What a token is

A token is a chunk of text, usually a short word or part of a longer one. In English, a rough rule is that 1,000 tokens is about 750 words. Code, other languages and unusual formatting use more tokens per word.

Prices are quoted per million tokens (often written “per 1M” or “/MTok”). So “$2 input” means $2 for every million tokens you send.

Tokenizers differ between providers and even between models. Anthropic notes that Claude 4.7 and later models (including Opus 5.5, Sonnet 5 and Fable 5.1) use a newer tokenizer that produces about 30% more tokens for the same text than earlier Claude models. The same prompt can cost a different amount on two models with identical per-token prices.

2. Input and output are priced separately

  • Input tokens are everything you send: your question, the system prompt, the chat history, any documents.
  • Output tokens are what the model writes back, including hidden “thinking” on reasoning models.

Output usually costs 3 to 8 times more than input; for most Anthropic, OpenAI and Google models it is 5 to 6 times. For example, Claude Sonnet 5 costs $2 in and $10 out; GPT-6 Sol also costs $2 in and $10 out.

Worked example

A customer-support chatbot handles 1,000 messages a day. Each message sends 1,500 tokens (system prompt, history, question) and gets 500 tokens back.

  • Input: 1,000 × 1,500 = 1.5M tokens per day
  • Output: 1,000 × 500 = 0.5M tokens per day

On Claude Sonnet 5 that is 1.5 × $2 + 0.5 × $10 = $8 a day, about $240 a month. On Gemini 3.8 Flash ($0.75 / $3.75) it is $3 a day. Try your own numbers in the cost calculator.

3. Caching: pay less for repeated context

If every request starts with the same long system prompt or document, providers can cache it, and repeated reads cost a fraction of the normal price.

Provider Cache read price Cache write Notes
Anthropic 10% of input (90% off) on most models 1.25x input for 5 minutes, 2x for 1 hour Fable 5.1 and Mythos 5.1 reads are 97.5% off; Opus 5.5 reads 95% off
OpenAI 10% of input (90% off) on GPT-5.6 and later 1.25x input Prompts from 1,024 tokens; cache lasts 30 minutes
Google Listed at 10% of input on current models Storage fee per hour for explicit caches Implicit caching is on by default

Caching matters most for agents and coding tools, which re-read the same large context on every step. That is why a model with a cheap cache read, like Claude Opus 5.5 at $0.20 per million, can end up cheaper in practice than its headline price suggests.

4. Batch: half price if you can wait

Anthropic, OpenAI and Google all offer a Batch API at 50% off input and output. You submit many requests at once and get the results back later (OpenAI promises within 24 hours). It’s ideal for overnight jobs: tagging a product catalogue, summarising documents, generating descriptions.

5. Time-of-day pricing: DeepSeek

DeepSeek charges different prices at different times. Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday, excluding Chinese public holidays; everything else is off-peak at half price.

Model Peak input / output Off-peak input / output
DeepSeek V4 Pro $1.32 / $3.96 $0.66 / $1.98
DeepSeek V4.1 Flash $0.30 / $1.20 $0.15 / $0.60

If you run batch jobs, scheduling them off-peak halves the bill.

6. Long-context surcharges

Some models charge more once a prompt passes a size threshold. Gemini 3.1 Pro, for example, costs $2 / $12 up to 200K tokens and $4 / $18 above that. Anthropic says its Claude 4.6 and later models have no long-context surcharge up to their 1M-token window.

Checklist to cut your bill

  1. Trim the prompt. Every token of system prompt is paid on every request.
  2. Cap the output length. Output is the expensive side.
  3. Turn on caching for anything repeated across requests.
  4. Use batch for anything that doesn’t need an instant answer.
  5. Match the model to the task. Budget models like GPT-6 Luna ($0.10 / $0.50) handle classification and extraction well.

Compare every model’s current price on the price list.

Sources

  1. Anthropic pricing
  2. Anthropic prompt caching
  3. OpenAI API pricing
  4. OpenAI prompt caching
  5. Gemini API pricing
  6. Gemini context caching
  7. DeepSeek API pricing