PlainLogic

Interactive lab · Practical AI

Tokens: The Pieces AI Reads

Before a model can think about your words, it has to chop them up. Tokens are the pieces it reads — and the pieces it writes back.

The experiment

Press Run, watch the idea move

This animated diagram is powered by the PlainLogic flow engine. Follow the packet along the edges — every step is labeled as it fires.

Everything here is a simplified educational visualization — the shape of the idea, not real model internals. The Run button drives a simulated animation in your browser; no real AI model runs and nothing is sent anywhere.

The LLM flow — one token at a time

Simplified educational visualization

The model never sees your words as words. Everything first passes through TOKENIZE: the text is sliced into tokens, and only then does the model predict the next one. Watch the packet stop at that second box — that split is where pricing, context limits, and some of AI's strangest mistakes all come from. As with the whole flow, no real tokenization runs here; the animation shows the shape of the idea.

Ready. Press Run to watch a prompt travel through the flow.

Simplified educational visualization. The token chips below show one illustrative split — not the tokenizer of any commercial model.

Plain-language AI

In plain logic

A token can be a whole word, part of a word, punctuation, or a fragment of one. The split depends on the tokenizer — a fixed rulebook each model ships with. Common words often survive as one token; rare words get shattered into pieces.

"unbelievable!" becomes: un believ able ! "The cat sat" becomes: The ␣cat ␣sat

Notice the little box before cat: spaces are usually glued onto the following word, not kept as their own tokens. And this split is only one illustrative example — a real tokenizer might cut “unbelievable” differently. There is no single universal way to slice text.

Why tokens matter in practice: context windows are measured in tokens, not words or characters, and API pricing is too. A long document costs what it costs in tokens — and non-English text, code, and unusual spellings can inflate the count far beyond what the character length suggests.

Hands-on

Try this

Say “cats” out loud — one token, most likely. Now say “unbelievable!” — probably four. The visible chunking rule above is an illustration, but the principle is real: the model works with numeric IDs for these pieces, never with the letters you see.

This also explains some famous AI quirks. Ask a model how many R's are in “strawberry” and older models stumble — because to the model, “strawberry” is not eight letters, it is two or three opaque token IDs. It never saw the letters the way you do. Tokenization shapes what the model can and cannot reason about.

Honest boundaries

What this leaves out

The “one token is about four characters” rule you may have heard is only a rough English estimate. For billing or context-limit calculations, use the actual tokenizer of the model you are calling — the real count can differ by 30% or more, especially for other languages, code, or emoji.

This page also skips how tokenizers are built (byte-pair encoding and its cousins) and how special tokens mark the start and end of text. The core idea to keep: text is re-sliced into opaque pieces before the model ever engages, and everything downstream — cost, limits, even some mistakes — inherits that slicing.

Honest answers

Questions people ask

Are tokens the same for every AI model?

No. Each model family ships its own tokenizer, so the same sentence can split into different numbers of tokens on different models. That is why token counts — and costs — vary between providers.

Why does AI sometimes miscount letters in a word?

Because it does not read letters; it reads token IDs. A word like “strawberry” may be stored as two or three opaque chunks, so asking about individual letters is asking about something below the model's resolution.

How many tokens is my document?

For a rough English estimate, divide the character count by four. For anything you will pay for or fit into a context window, run the text through the model's actual tokenizer instead — providers publish free tools for this.