Guide

Token cost by use case: chatbot, RAG and summarization

Worked token and cost estimates for three common AI workloads — a support chatbot, retrieval-augmented generation (RAG) and document summarization.

Prices last updated:

Token usage depends far more on what you build than on which model you pick. Here are three common workloads with typical request sizes and what they cost per month on every model we track. All figures use standard list prices, no caching or batch discount.

Customer-support chatbot

System prompt (~800 tokens), a few help-center snippets (~1,500) and recent conversation history (~700). Replies are short.

3,000 input + 300 output tokens per request, 100,000 requests per month. Open in the calculator.

ModelPer requestPer month
DeepSeek V4 Flash$0.000504$50.40
GPT-5.6 Luna$0.00096$96.00
DeepSeek V4 Pro$0.001566$156.60
Gemini 3.8 Flash$0.003375$337.50
Claude Haiku 4.5$0.0045$450.00
Grok 4.7$0.0078$780.00
Claude Sonnet 5$0.009$900.00
GPT-5.6 Terra$0.0096$960.00
Gemini 3.1 Pro$0.0096$960.00
Claude Opus 5.5$0.018$1,800.00
Claude Opus 5$0.0225$2,250.00
GPT-5.6 Sol$0.024$2,400.00
GPT-6 Astra$0.045$4,500.00
Claude Fable 5.1$0.045$4,500.00

RAG question answering

Eight retrieved chunks of ~900 tokens plus instructions and the question. Answers cite sources in a few paragraphs.

8,000 input + 500 output tokens per request, 50,000 requests per month. Open in the calculator.

ModelPer requestPer month
DeepSeek V4 Flash$0.00126$63.00
GPT-5.6 Luna$0.0022$110.00
DeepSeek V4 Pro$0.003915$195.75
Gemini 3.8 Flash$0.007875$393.75
Claude Haiku 4.5$0.0105$525.00
Grok 4.7$0.019$950.00
Claude Sonnet 5$0.021$1,050.00
GPT-5.6 Terra$0.022$1,100.00
Gemini 3.1 Pro$0.022$1,100.00
Claude Opus 5.5$0.042$2,100.00
Claude Opus 5$0.0525$2,625.00
GPT-5.6 Sol$0.055$2,750.00
GPT-6 Astra$0.105$5,250.00
Claude Fable 5.1$0.105$5,250.00

Document summarization

A 25–30 page report in, a structured one-page summary out. Very input-heavy, and a good fit for batch pricing.

15,000 input + 800 output tokens per request, 10,000 requests per month. Open in the calculator.

ModelPer requestPer month
DeepSeek V4 Flash$0.002324$23.24
GPT-5.6 Luna$0.00396$39.60
DeepSeek V4 Pro$0.007221$72.21
Gemini 3.8 Flash$0.0143$142.50
Claude Haiku 4.5$0.019$190.00
Grok 4.7$0.0348$348.00
Claude Sonnet 5$0.038$380.00
GPT-5.6 Terra$0.0396$396.00
Gemini 3.1 Pro$0.0396$396.00
Claude Opus 5.5$0.076$760.00
Claude Opus 5$0.095$950.00
GPT-5.6 Sol$0.099$990.00
GPT-6 Astra$0.19$1,900.00
Claude Fable 5.1$0.19$1,900.00

What these numbers tell you

  • Input dominates RAG and summarization. Retrieve fewer, tighter chunks; cache the shared instructions; send only the sections you need.
  • Chatbots grow with conversation length. Summarize or truncate old turns and cache the system prompt.
  • Model choice is a 10–100× lever. Test whether a small model meets your quality bar before defaulting to a flagship.
  • Summarization rarely needs to be instant. Batch APIs halve the cost for jobs that can wait.

To size your own workload, follow how to estimate API costs before you build.

Try it on your own text. The Tokenova calculator counts tokens in your browser and prices them on every major model.

Frequently asked questions

How many tokens does a chatbot conversation use?
A typical support-bot turn uses 2,000–4,000 input tokens once the system prompt, retrieved help content and history are included, and 200–400 output tokens. Input grows every turn because history is resent.
How many tokens does a RAG query use?
Most RAG requests send 4,000–12,000 input tokens of retrieved context and return 300–800 output tokens. The number and size of retrieved chunks is the main cost lever.
How much does it cost to summarize a document with AI?
A 25-page document is about 15,000 tokens. Summarizing it costs from a fraction of a cent on low-cost models to around 20 cents on flagship models, and batch APIs halve that where offered.
Which use case is most expensive?
Per request, long-document work is most expensive because of its input size. Per month, high-volume chatbots often cost the most simply because of request count, and agents can exceed both because they make many calls per task.