Guide
What is a token?
A plain-English explanation of AI tokens: how tokenizers split text, how many tokens a word or page uses, and why token counts differ between models.
A token is the unit of text a language model reads and writes. Before a model sees your prompt, a tokenizer breaks the text into pieces — whole words, parts of words, punctuation or single characters — and maps each piece to a number. AI providers bill by these pieces, which is why every price list is quoted “per million tokens”.
How tokenizers split text
Modern tokenizers use byte-pair encoding (BPE). They start from individual bytes and repeatedly merge the most common neighbouring pairs seen in training data, building a vocabulary of a few hundred thousand pieces. Frequent words end up as one token; rarer words are assembled from several.
| Text | Likely split | Tokens |
|---|---|---|
| hello | hello | 1 |
| tokenization | token ization | 2 |
| 2026-09-25 | 202 6 - 09 - 25 | ≈6 |
| {"id": 42} | {" id ": 42 } | ≈6 |
Exact splits vary by tokenizer — the table shows the general pattern, not a guarantee.
Rules of thumb for English
- 1 token ≈ 4 characters ≈ ¾ of a word.
- 100 tokens ≈ 75 words; 1,000 tokens ≈ 750 words, about 1½ pages of prose.
- 1 million tokens ≈ 750,000 words — roughly 2,000 pages, or ten average novels.
These ratios break down for code (lots of symbols and indentation), numbers, URLs and non-English text. Languages written in non-Latin scripts can use two to four times as many tokens for the same meaning, which makes the same conversation more expensive.
Why token counts differ between models
Each provider trains its own tokenizer. OpenAI’s current models use o200k_base; Anthropic, Google, xAI and DeepSeek use their own vocabularies, and a provider may change tokenizer between model generations. The same paragraph might be 1,000 tokens on one model and 1,100 on another, so a lower per-token price doesn’t always mean a lower bill.
Tokenova counts exactly with o200k_base in your browser and applies a per-provider adjustment for models whose tokenizers aren’t publicly available as browser libraries. Those counts are labelled as estimates. For billing-grade numbers on other providers, use their token-counting API endpoints.
Tokens you don’t see
Your bill includes more than the text you typed:
- System prompts and instructions sent with every request.
- Conversation history — chat apps resend earlier turns each time, so input grows as the chat goes on.
- Tool definitions and retrieved documents in agent and RAG setups.
- Reasoning tokens — models that “think” before answering bill that thinking as output tokens, even when it isn’t shown.
Next, see how AI pricing works or jump straight to why output tokens cost more.
Frequently asked questions
- How many tokens is one word?
- In English, one word is about 1.3 tokens on average, so 100 words is roughly 130 tokens. Common short words are usually a single token; long, rare or technical words split into several.
- How many characters are in a token?
- About four characters of English text per token is a good rule of thumb. Code, numbers, non-Latin scripts and emoji usually use more tokens per character.
- Do all AI models count tokens the same way?
- No. Each model family has its own tokenizer, so the same text can produce different token counts on GPT, Claude, Gemini or DeepSeek models — typically within 10–20% of each other for English prose.
- Do spaces and punctuation count as tokens?
- Yes. Spaces are usually merged into the following word’s token, while punctuation marks are often tokens of their own. Formatting such as JSON braces, Markdown and indentation all adds tokens.