Guide

Input vs output token cost, explained

Why output tokens cost 2–6× more than input tokens, how to estimate the split for your workload, and practical ways to keep output costs down.

Prices last updated:

Every AI API bill has two lines: input tokens — everything you send — and output tokens — everything the model writes back. Output is always the more expensive of the two, and for many applications it’s the bigger half of the bill even though there are fewer output tokens.

Why output costs more

When a model reads your prompt it processes all input tokens together in one parallel pass. When it writes, it must produce tokens one after another: each new token needs another pass that reads everything before it. Output keeps expensive hardware busy for longer per token, so providers price it higher.

The output multiplier by model

ModelInput / 1MOutput / 1MOutput ÷ input
GPT-5.6 Sol OpenAI$5.00$30.006×
GPT-5.6 Terra OpenAI$2.00$12.006×
Gemini 3.1 Pro Google$2.00$12.006×
GPT-5.6 Luna OpenAI$0.20$1.206.0×
GPT-6 Astra OpenAI$10.00$50.005×
Claude Fable 5.1 Anthropic$10.00$50.005×
Claude Opus 5.5 Anthropic$4.00$20.005×
Claude Opus 5 Anthropic$5.00$25.005×
Claude Sonnet 5 Anthropic$2.00$10.005×
Claude Haiku 4.5 Anthropic$1.00$5.005×
Gemini 3.8 Flash Google$0.75$3.755×
Grok 4.7 xAI$2.00$6.003×
DeepSeek V4 Pro DeepSeek$0.435$0.872×
DeepSeek V4 Flash DeepSeek$0.14$0.282×

A worked example

Take a request with 2,000 input tokens and 1,000 output tokens on Claude Sonnet 5 ($2.00 in, $10.00 out):

  • Input: 2,000 ÷ 1,000,000 × $2.00 = $0.0040
  • Output: 1,000 ÷ 1,000,000 × $10.00 = $0.0100

The output is half as many tokens but 71% of the cost. That’s why the calculator lets you override the output estimate in its advanced options — the output guess matters more than the input count.

How to estimate your output length

  • Measure it. Run 20–50 realistic prompts and record the usage numbers the API returns. Use the median for budgets and the 90th percentile for worst cases.
  • Use typical ratios when you have nothing to measure: summaries 5–10% of the input; chat replies 150–500 tokens; generated articles or code files 1,000–4,000 tokens.
  • Remember reasoning. With thinking models, add the reasoning budget to the visible answer.

Ways to cut output cost

  • Set a max_tokens limit that fits the task, and ask for concise answers in the prompt.
  • Request structured output (JSON with only the fields you need) instead of prose.
  • Use lower reasoning effort, or a non-reasoning model, for simple tasks.
  • Route easy requests to a cheaper model and reserve flagship models for hard ones.

See how this plays out across real workloads in token cost by use case.

Try it on your own text. The Tokenova calculator counts tokens in your browser and prices them on every major model.

Frequently asked questions

Why are output tokens more expensive than input tokens?
Input tokens are processed in parallel in a single pass, while output tokens are generated one at a time, each needing its own pass through the model. Generation uses far more compute and memory time per token, so providers charge more for it.
How much more do output tokens cost?
Typically 2–6× the input price. Most current OpenAI, Anthropic and Google models charge 5–6× more for output; some others, such as Grok and DeepSeek models, charge about 2–3× more.
What is a normal ratio of input to output tokens?
It depends on the task. Summarization and RAG are input-heavy (10:1 or more). Chatbots are often 3:1 to 5:1 once history is included. Content and code generation can be output-heavy (1:2 or more).
Do reasoning or thinking tokens count as output?
Yes. Models that reason before answering bill those reasoning tokens at the output rate, even if the reasoning is not returned to you.