How this calculator works
- Prices come from each provider’s own API pricing page and are re-checked every day. Each figure has the date it was last confirmed; see our methodology.
- Monthly cost is input tokens × input price plus output tokens × output price, per million tokens, times the number of requests. “Per day” assumes a 30-day month.
- Caching moves the chosen share of each prompt to the provider’s cached-input price. One-off cache write and storage charges aren’t included.
- Batch switches every model to its published batch prices, including the cached-input price in batch mode.
- Your numbers stay private. The calculation runs in your browser; nothing you enter is sent anywhere.
Frequently asked questions
How is AI API cost calculated?
Providers charge per token, with separate prices for input (your prompt, instructions and any documents) and output (the model’s reply). Prices are quoted per million tokens, so a request’s cost is its input tokens × the input price ÷ 1,000,000, plus its output tokens × the output price ÷ 1,000,000. Multiply by the number of requests to get a monthly bill.
What are cached input tokens?
If many requests start with the same text, such as a long system prompt or a shared document, the provider can reuse it from a cache and bill it at a much lower rate. The “cached” slider sets how much of each prompt is read from the cache. Writing to the cache can cost extra (Anthropic and OpenAI charge for cache writes, Google charges for storage by the hour); those one-off charges aren’t included here.
What is the Batch API?
All three providers let you send requests in bulk to be processed asynchronously instead of straight away, at a lower price. It suits work that doesn’t need an instant answer, such as tagging a backlog of documents or running evaluations overnight.
Do reasoning or “thinking” tokens count?
Yes. Reasoning models bill their internal thinking as output tokens, even though you don’t see it. Google’s pricing page says so explicitly. For reasoning-heavy work, set output tokens well above the length of the visible answer.
Why might my real bill be different?
This calculator covers text input and output at the standard rate. Your bill can also include very long prompts (some models charge more above a threshold), images or audio, tools such as web search, regional or data-residency surcharges, and taxes. Free tiers and negotiated discounts can lower it.
How do I find my token counts?
Paste a typical prompt into our token counter to see how many tokens it uses. As a rule of thumb, 1,000 tokens is about 750 English words. Your provider’s usage dashboard shows exact figures for real traffic.