The simple explanation
Before a model ever sees your text, a tokenizer slices it into tokens — chunks that can be whole words, parts of words, punctuation, or spaces. cats might be a single token; unbelievable might be split into three or four. The split depends entirely on the tokenizer, and every model family slices differently. Inside the model, tokens are just numeric IDs — the model never sees letters or words, only these numbered pieces.
See this idea move.
The Token Playground experiment walks through this concept step by step — press run and watch it happen. Everything is simulated in your browser; no real AI runs.
A concrete example
Open the Token Playground and type “cats”, then “unbelievable!”. The visible chunking shows how one word can be a single piece while another breaks apart. (The playground’s rule illustrates the idea; it is not the tokenizer of any commercial model.) Notice how unusual spellings, other languages, and emoji tend to produce more tokens — the slicer falls back to smaller fragments for text it rarely saw in training. This is why the same sentence translated into another language can cost noticeably more tokens than the English original: the tokenizer’s vocabulary was built around the text it trained on, so ordinary English compresses well and everything else breaks into smaller, pricier pieces.
This one unglamorous step quietly shapes product behavior everywhere: autocomplete counters, rate limits measured in tokens, and context windows that fill up faster in some languages than others — all of it flows from how the text was sliced before the model ever saw it.
Where people get misled
Mistake one: token cost equals word cost. It doesn’t. Bills and context limits are counted in tokens, and a model with a 128k context window holds roughly three-quarters that many English words — fewer in other languages. Mistake two: expecting letter-perfect awareness. Because the model sees chunks, not letters, it can stumble on questions like how many times a letter appears in a word — it literally never saw the word as letters. Mistake three: assuming tokenization is universal. Switch models and the same sentence slices differently.
The honest limits
The “one token is about four characters” rule is a rough English estimate only. For billing math or context-limit calculations, always use the actual model’s tokenizer — guesses cost money. And tokenizers can quietly make some languages far more expensive to process than English.