Paste the same paragraph into two token counters and you can get two different answers. Neither is wrong. Each AI model family splits text with its own tokenizer, and the split depends heavily on what kind of text you give it.
We counted a set of short samples with two tokenizers that OpenAI publishes: o200k_base, used by GPT-4o and newer OpenAI models, and the older cl100k_base, used by GPT-4 and GPT-3.5. The results show where rules of thumb hold and where they fall apart.
The test
Every sample below was counted exactly with each tokenizer, then compared with the two usual estimates: characters divided by 4, and words multiplied by 1.33.
| Sample |
Words |
o200k_base |
cl100k_base |
Chars ÷ 4 |
Words × 1.33 |
| English sentence |
9 |
10 |
10 |
11 |
12 |
| Same sentence in Spanish |
9 |
15 |
17 |
13 |
12 |
| Same sentence in Arabic |
7 |
16 |
30 |
11 |
9 |
| Same sentence in Urdu |
10 |
20 |
46 |
12 |
13 |
| A line of JavaScript |
14 |
24 |
24 |
17 |
19 |
| An invoice line with a date and amount |
4 |
18 |
18 |
9 |
5 |
| A small JSON object |
7 |
20 |
20 |
13 |
9 |
| A short message with three emoji |
4 |
10 |
14 |
4 |
5 |
The English sentence was "The quick brown fox jumps over the lazy dog." The other language samples are translations of it, so they carry the same meaning.
What a tokenizer actually does
A tokenizer has a fixed vocabulary of text pieces, built by finding the character sequences that appear most often in its training data. Common English words become single tokens. Rarer words are split into pieces: "tokenization" becomes "token" + "ization" in o200k_base. Anything not covered falls back to smaller pieces, down to individual bytes.
So the token count of a text depends on how well its words match the pieces in the vocabulary. Text that looks like the training data is cheap. Text that doesn't is expensive.
Language makes the biggest difference
English matched both estimates almost exactly. Spanish already needed 15 tokens for the same nine words, about 50% more than English.
The jump for Arabic and Urdu is far larger. In the older cl100k_base tokenizer, the Urdu sentence took 46 tokens, more than four times the English version. The newer o200k_base vocabulary is twice as large and includes many more non-English pieces, which brought the same sentence down to 20 tokens.
This matters in practice. If you work in Urdu, Arabic, Hindi or another non-Latin script, a model with an older or English-heavy tokenizer fills its context window faster and costs more per sentence. The same document can cost twice as much to process depending on the model.
Numbers and code run high everywhere
The invoice line "Invoice 2026-10-02: $1,234,567.89 due" has only four words but used 18 tokens in both tokenizers. Numbers are split into short groups of digits, and every hyphen, comma, colon and dollar sign is its own piece. The words-based estimate of 5 tokens was off by a factor of more than three.
Code and JSON behave the same way. Brackets, quotes, operators and indentation all cost tokens, so the line of JavaScript used 24 tokens against estimates of 17 to 19. For spreadsheets pasted as text, log files, source code and API responses, expect real counts close to double the usual rule of thumb.
Emoji are surprisingly expensive
Three party emoji and a thumbs up, plus two words, came to 10 tokens in o200k_base and 14 in cl100k_base. An emoji is several bytes long, and unless it is common enough to have its own vocabulary entry, it is split across more than one token.
Why other model families differ again
Anthropic, Google and Meta train their own tokenizers with different vocabularies, so a count from an OpenAI tokenizer is only an approximation for Claude, Gemini or Llama. For ordinary English the difference is usually modest. For the cases above, such as non-English text, code and numbers, it can be large in either direction.
Providers also count tokens you never see in your text: the formatting that wraps each chat message, system instructions, tool definitions and any images or files. Your usage report will therefore show more tokens than a counter shows for the text alone.
Practical rules
- Measure, don't estimate, for anything that isn't plain English. The rules of thumb were built for English prose.
- For code, data and numbers, double the estimate if you can't count exactly.
- Compare models on your own text. If you process large amounts of one language, the tokenizer can matter as much as the price per token.
- Leave headroom. Plan to use no more than about 80% of a model's context window, because the reply and hidden formatting need room too.
To check your own text, paste it into the AI token counter. It gives the exact o200k_base count, both estimates, and a cost figure if you enter your provider's price. For how much fits in a model's context window, use the context window calculator.