Guide
How to estimate AI API costs before you build
A step-by-step method to forecast monthly AI API spend before writing code: measure prompts, model the output, multiply by volume and add a safety margin.
Prices last updated:
You can forecast an AI feature’s running cost to within a factor of two before writing a line of code. It takes five steps and about fifteen minutes with the calculator.
Step 1: Write a realistic prompt
Draft the actual system prompt, one or two example user messages and any context you’ll attach (retrieved documents, product data, conversation history). Paste the whole thing into the calculator. Rough placeholders like “system prompt here” underestimate badly — system prompts and context are often the largest part of input.
Step 2: Estimate the output
Decide how long a good answer is and convert it to tokens (about 1.3 tokens per word). A three-paragraph reply is roughly 250–400 tokens. If you plan to use a reasoning model, add its thinking budget. Enter this under Advanced options → Output tokens. Why output matters most.
Step 3: Model the volume
Work out requests per day, not users. One chat session might be 5–10 requests; an agent completing a task might make 20–50. Multiply by active users for a daily figure.
Step 4: Price it across models
The comparison table shows every model at your request size and volume. Shortlist a flagship model for quality and a cheaper one to test against it. Enable batch pricing for work that can wait up to a day.
Step 5: Add a margin
Real traffic is messier than samples. Add 20–30% for predictable workloads and 50%+ for agents and long chats. Then use the budget planner to check how much headroom your budget leaves.
Worked example: a support assistant
- System prompt and help-center excerpts: ~2,300 tokens; user message and short history: ~300 tokens → 2,600 input tokens.
- Answers of ~250 words → 350 output tokens.
- 600 conversations a day × 5 messages → 3,000 requests/day.
On GPT-5.6 Terra ($2.00 in / $12.00 out) that is $0.0094 per request, about $846.00 per month, or $1,099.80 with a 30% margin. Open this example in the calculator to compare other models.
After you build: measure
Every API response reports the exact input and output tokens used. Log them from day one, compare them with your estimate after a week, and update the forecast. Look for the usual surprises: longer histories than expected, retries and verbose answers.
For typical token sizes by workload, see token cost by use case.
Frequently asked questions
- How do I estimate AI API costs before building?
- Write realistic sample prompts, count their tokens, estimate output length, multiply by the per-token price, then by expected daily requests and 30 days. Add a 20–50% margin for retries, longer conversations and growth.
- What safety margin should I add to an AI cost estimate?
- Add 20–30% for a well-understood workload and 50% or more for agents, long chats or reasoning-heavy tasks, where token use varies widely between requests.
- Should I estimate with the cheapest model?
- Estimate with the model you expect to need, and a cheaper fallback. Prototyping on a flagship model and then testing whether a smaller model holds quality often cuts costs by 5–10×.
- How accurate are pre-build cost estimates?
- Within about a factor of two if your sample prompts are realistic. Replace the estimate with measured usage from the API’s usage fields as soon as you have a prototype.