Guide
How AI API pricing works
How AI providers bill for API usage: per-million-token rates, input vs output prices, context windows, batch discounts, caching and long-context surcharges.
Prices last updated:
AI APIs are metered like electricity: you pay for what you use, measured in tokens. Every provider publishes two prices per model — one for the tokens you send and one for the tokens the model generates — quoted per million tokens.
The basic formula
Cost per request = (input tokens × input price + output tokens × output price) ÷ 1,000,000
Multiply by requests per day and by ~30 to get a monthly figure. That’s exactly what the calculator does, for every model at once.
What drives the price
1. Model tier
Providers sell several sizes. Flagship models cost the most and handle the hardest work; small models are much cheaper and fast. Across the 14 models we track, output prices range about 179× from cheapest to most expensive — from $0.28 to $50.00 per million tokens (GPT-6 Astra). The cheapest input today is DeepSeek V4 Flash from DeepSeek at $0.14.
2. Input vs output
Output tokens cost more — usually 2–6× the input rate — because generating each token takes a full pass through the model, while input is processed in parallel. Read the full explanation.
3. Context length
The context window is the maximum tokens a model can consider at once. Large windows (many models now accept around 1M tokens) let you send whole documents, but you pay for every token you send. Some providers also charge a higher rate once a single prompt passes a threshold (for example 200K or 272K tokens) — we note these on each model page.
Discounts that change the maths
- Batch APIs: OpenAI, Anthropic and Google take 50% off requests you submit as an asynchronous batch, returned within 24 hours. Good for evaluations, backfills and bulk classification. Toggle this in the calculator’s advanced options.
- Prompt caching: if many requests start with the same long prefix (a system prompt, a manual, a codebase), providers can cache it and bill repeat reads at a fraction of the input rate — often around a tenth. Writing to the cache may cost slightly more than normal input.
- Promotional rates: new models sometimes launch with temporary prices. We record the end date on the model page when a provider publishes one.
Tokenova shows standard list prices, optionally with batch discounts. Caching savings depend heavily on your traffic pattern, so treat our numbers as the uncached upper bound.
Hidden costs to budget for
- Reasoning tokens. Models that think before answering bill that reasoning as output. A short visible answer can carry thousands of hidden output tokens.
- Retries and failures. Timeouts, validation failures and agent loops all consume tokens.
- Growing chat history. Each turn resends the previous ones, so a 20-turn conversation costs far more than 20 × one turn.
- Other modalities. Images, audio and tool calls such as web search are priced separately.
Ready to put numbers on your own project? Follow how to estimate API costs before you build.
Frequently asked questions
- What does “per 1M tokens” mean?
- It is the price for one million tokens. To get the cost of one request, divide its token count by 1,000,000 and multiply by the rate. 2,000 input tokens at $2 per 1M costs 2,000 ÷ 1,000,000 × $2 = $0.004.
- Am I charged for the tokens in my prompt and the reply?
- Yes. Input tokens (everything you send, including system prompts and history) and output tokens (everything the model generates, including hidden reasoning) are billed separately, each at its own rate.
- Is there a monthly fee for AI APIs?
- Standard API access is pay-as-you-go with no monthly fee: you pay only for tokens used. Consumer chat subscriptions are separate products and don’t include API usage.
- How can I lower my AI API bill?
- Pick the smallest model that meets your quality bar, cap output length, cache repeated prompt prefixes, use batch APIs for work that can wait, and trim conversation history and retrieved context.