The direct answer
A context window is the amount of text an AI model can consider at one time. Think of it as the model's working memory: everything it needs to answer you — your question, the conversation so far, any documents you've shared — has to fit inside it.
Anthropic's own glossary defines it as the amount of text a language model can "look back on and reference" when generating new text — distinct from the vast data it was trained on. When a conversation outgrows the window, the earliest parts drop out. That's why AI seems to "forget" long chats.
How it works
Everything the model sees shares one window: the system instructions, your messages, pasted documents, retrieved information, and tool results all draw from the same pool. It's all counted in tokens — the small chunks of text the model actually processes.
The window is finite for a structural reason. Every token the model considers can interact with every other token, so the work grows fast as the window fills. Anthropic describes this as an "attention budget": every new token depletes it a little, and the model's focus stretches thinner the more you load in.
A simple example
Imagine a desk that only fits ten sheets of paper. Every time you add a new sheet, an old one slides off the edge. You can keep working — but anything on the fallen sheets is gone unless you wrote it down somewhere else.
A long chat works the same way. The model isn't storing your conversation anywhere; each reply is generated with only what's currently on the desk. When old messages slide off, the model genuinely can't see them anymore. It isn't forgetting so much as it never had long-term memory in the first place.
Why it matters
This is the hidden constraint behind most AI disappointments. Paste a huge document and ask a subtle question, and the answer may miss details that were technically "in there." The information fit, but the model's grip on it loosened as the load grew.
That's why curating what goes into the window matters more than its size. Keeping the input small, relevant, and high-signal — the discipline Anthropic now calls "context engineering" — beats stuffing in everything and hoping. For builders, it also means long-running agents need strategies for what to keep, what to summarize, and what to let slide off the desk.
The common misunderstanding
The common misunderstanding: a bigger context window means a better memory. It doesn't. A bigger window is a bigger desk, not a filing cabinet. Nothing is stored between conversations on the model's side — when you start a new chat, the desk is empty.
What feels like memory in a chat app is the app resending your conversation history with every message. That's why a long chat can suddenly lose the thread: the app is quietly trimming or summarizing old messages to stay within the window, and the model only ever sees what's sent this turn.
What changed recently
Context windows have exploded in size — from a few thousand tokens to millions in a few years — but recent research is a useful reality check. Chroma's team tested 18 leading models and found that performance still degrades as input length grows, even when the task itself stays simple. Near-perfect scores on easy retrieval tests don't mean the model handles real long-context work equally well.
The field's response has been a shift in emphasis. Anthropic's 2025 writing on the topic argues the craft is moving from prompt engineering — finding the right words — to context engineering: deciding what the model should see at all. The frontier question isn't "how much fits" anymore; it's "what deserves to be there."
Try it on PlainLogic
The AI Lab is where ideas like this stop being abstract — you can watch a model juggle its limited working memory in the interactive experiments.