AI text tools

AI Token Counter

Paste text to see how many tokens it uses. You get an exact count with OpenAI’s published tokenizer, a range for other models, and an optional cost estimate from the price you enter.

Runs in your browser. Your text is not uploaded.

From your provider’s pricing page, to estimate cost.

Token count

≈ 0Estimated tokens (range from two rules of thumb)

Characters0
Words0
Lines0
Estimate: characters ÷ 40
Estimate: words × 1.330

Claude, Gemini, Llama and other models use their own tokenizers, so the same text gives a different count on each. Use the range as a planning figure.

How this calculator works

Characters, words and lines are counted directly. Characters are counted as Unicode code points, so an emoji or an accented letter counts as one.

Two common rules of thumb give the estimated range: about four characters per token, and about 1.33 tokens per word. For ordinary English prose they land close together; for code, numbers or short words they drift apart, which is itself a useful signal.

The exact count uses o200k_base, the tokenizer OpenAI publishes for GPT-4o and its newer models. It loads only when you ask for it, because its vocabulary file is about 1 MB.

Anthropic, Google and Meta use their own tokenizers, and the providers count some things (images, tool definitions, message wrappers) that plain text doesn’t show. Treat the range as a planning figure, and check your provider’s usage report for billing.

If you enter an input price per million tokens, the tool multiplies it by the exact count, or by the top of the range if you haven’t loaded the tokenizer, so the cost estimate errs on the high side.

Worked example

A 1,000-word English article

  1. Roughly 5,600 characters including spaces gives 5,600 ÷ 4 = 1,400 tokens.
  2. 1,000 words × 1.33 = 1,330 tokens, so the range is about 1,330 to 1,400.
  3. At an input price of $2.50 per million tokens, 1,400 tokens costs 1,400 ÷ 1,000,000 × $2.50 = $0.0035.

Questions people ask

Why do Claude, Gemini and ChatGPT give different token counts for the same text?

Each provider trains its own tokenizer with its own vocabulary, so the same sentence is split into different pieces. Differences of 10% to 30% are common, and larger for code and non-English text.

Does this count the tokens in my chat history or system prompt?

Only what you paste. In a chat app, the earlier messages, any system instructions and attached files all count toward the context window too, so paste everything you want measured.

Why does non-English text use more tokens?

Tokenizers are trained mostly on English, so common English words are often a single token while words in other languages, especially in non-Latin scripts, are split into several pieces.

Is my text sent anywhere?

No. Counting happens in your browser. The tokenizer file is downloaded from this site, and your text never leaves your device.

Sources

Last reviewed October 2, 2026