Quick answer: 1,000 tokens is roughly 750 words of ordinary English, about one and a half single-spaced pages. Going the other way, one English word averages about 1.3 tokens. These are rules of thumb: code, numbers and most other languages use more tokens per word.
What is a token?
AI language models don't read text word by word. They break it into tokens, which are common chunks of characters. A short, common word like "the" is usually one token. A longer or rarer word is split into pieces: "unbelievably" might become "un", "believ" and "ably". Spaces, punctuation and numbers also become tokens.
Models measure everything in tokens: how much you can paste in (the context window), how long an answer can be, and, for paid APIs, how much a request costs.
Tokens to words: conversion table
These use the common English estimate of 0.75 words per token, and 500 words per single-spaced page.
| Tokens |
About this many words |
About this many pages |
| 100 |
75 |
a short paragraph |
| 1,000 |
750 |
1½ |
| 4,000 |
3,000 |
6 |
| 32,000 |
24,000 |
48 |
| 128,000 |
96,000 |
190 |
| 200,000 |
150,000 |
300 |
| 1,000,000 |
750,000 |
1,500 |
For scale, a typical novel is around 80,000 to 100,000 words, which is roughly 105,000 to 135,000 tokens.
Why the ratio isn't fixed
Different models use different tokenizers. Each model family splits text in its own way, so the same paragraph can produce noticeably different token counts in different tools.
Language matters. Tokenizers are built mostly from common patterns in their training data, which is heavily English. Spanish, German, French and other languages usually need more tokens per word, and languages with non-Latin scripts often need many more.
Code and data are token-heavy. Indentation, brackets, variable names and long numbers all break into many small tokens. A page of code typically uses more tokens than a page of prose.
Formatting adds up. Markdown symbols, tables, URLs and repeated whitespace all count.
How to estimate tokens for your own text
There are two quick methods:
- From words: word count × 1.33
- From characters: character count ÷ 4
For plain English they give similar answers. When they disagree, the larger one is the safer estimate, which is what the context window calculator uses. Paste your text and it shows the token estimate, how much of different context window sizes it fills, and how many pages fit.
If you need an exact count, for example to manage API costs, use the official token counting tool or API from the provider you're using. Each provider's count only applies to its own models.
Why this matters in practice
Fitting documents into a chat. If a tool has a 128,000-token context window, that's roughly 190 pages of English, but the conversation, your instructions and the reply all share that space.
Getting good answers. Even when a long document fits, models can miss details buried in the middle of very long inputs. Pasting only the relevant sections often works better.
Costs. API pricing is per token, usually with output tokens costing more than input tokens. Halving a prompt roughly halves its input cost.
The 0.75 words-per-token rule is for ordinary English prose. Other text behaves differently:
| Text type |
Tokens per word, roughly |
1,000 tokens is about |
| Plain English |
1.3 |
750 words |
| Spanish, French, German |
often 1.5 to 2 |
500 to 650 words |
| Technical writing with jargon |
1.5 or more |
about 650 words |
| Code |
varies widely |
often far fewer "words" |
| Numbers and tables |
high |
much less text than it looks |
These are broad estimates. If you work in another language, count a sample with your tool's tokenizer to get your own ratio.
Worked examples
A cover letter: 400 words × 1.33 ≈ 530 tokens.
A 15-minute talk transcript: about 2,200 words × 1.33 ≈ 2,900 tokens.
A 300-page book: about 90,000 words × 1.33 ≈ 120,000 tokens.
A spreadsheet pasted as text: 500 rows of five columns can easily exceed 15,000 tokens, because every number, comma and line break counts.
APIs count tokens in both directions. Input tokens are everything you send; output tokens are what the model writes back. Output tokens are often priced higher, and many tools set a separate maximum length for a single reply. If answers stop mid-sentence, the reply may have hit its output limit, not the context window.
Saving tokens without losing meaning
- Remove boilerplate: disclaimers, signatures, navigation text.
- Clean transcripts of timestamps with the transcript cleaner.
- Send only the columns of a table that matter.
- Replace long examples with one representative example.
- Ask for shorter answers when you don't need detail.
tokens ≈ words × 1.33 for English, or characters ÷ 4
Use the larger of the two for safety, as the context window calculator does.