AI cost calculator

Set your usage with the sliders or type exact numbers. Every model's monthly bill updates instantly.

Everything you send: system prompt, chat history, documents.

What the model writes back, including hidden reasoning.

%

Repeated prompt parts (system prompt, documents) billed at the cheaper cached rate.

A month is 30 days. Cache write fees, long-context surcharges and DeepSeek off-peak discounts are not included.

ModelPer requestPer monthPer yearRelative cost

How the calculator works

AI APIs bill by the token, with separate prices for what you send (input) and what the model writes back (output). The calculator multiplies your tokens per request by each model's official price, applies the cached rate to the share of input you mark as cached, and scales it to 30 days of traffic. Prices come from our price list, which is checked against each provider's pricing page.

How to pick the numbers

Start with a preset that looks like your use case, then adjust. For a chatbot, input per request is the system prompt plus the conversation so far plus the new message; output is the typical answer length. For document work, input is the document size. If you already use an API, your provider's usage dashboard shows the real averages.

Image and video costs

Image models mostly charge a flat price per picture that rises with resolution. Video models charge per second of generated footage. Switch tabs above to compare them; for occasional use, a subscription can work out cheaper.

Frequently asked questions

How many tokens is my prompt?

In English, 1,000 tokens is roughly 750 words. A one-page email is about 500 to 700 tokens, a 50-page PDF about 30,000. Code and other languages use more tokens per word. Chat apps resend the whole conversation history with each message, so input grows as a chat gets longer.

Why does output cost more than input?

Generating text is more work for the model than reading it, so providers charge 3 to 8 times more per output token. Capping the answer length is usually the quickest way to cut a bill.

What does prompt caching save?

When the start of your prompt repeats between requests, providers bill those tokens at a cached rate, usually about 10% of the normal input price. Coding agents and chatbots with long system prompts benefit most; set the cache slider to the share of input that repeats.

Is the Batch API worth it?

Anthropic, OpenAI and Google charge 50% less for batch requests that can wait up to a day. It suits overnight jobs such as tagging, summarising or translating large sets of documents.

How accurate is this estimate?

It uses each provider's list price and your numbers, so it is as accurate as your token estimates. Real bills also include cache write fees, long-context surcharges above certain prompt sizes and taxes, which are left out here.

Want the background? Read AI API pricing explained.