What it is
A context window is the model's working memory for a single conversation. Everything the model can see when it writes its next word has to fit inside it: your instructions, the files you pasted, the earlier back-and-forth, and the answer it is writing right now.
The size is counted in tokens, not pages or words. A window of 200,000 tokens holds very roughly 150,000 words of English, though the exact figure depends on the text.
How it works
The model does not remember you between chats. Each time you send a message, the app sends the model the whole conversation again, trimmed to fit the window. When a chat grows past the limit, something has to go. Usually the oldest messages are dropped or summarized, and the model quietly loses track of them.
A bigger window means you can paste a whole contract, a long codebase or a stack of research notes and ask questions across all of it, instead of feeding it in pieces.
Bigger is not always better
A large window is a ceiling, not a guarantee of quality. Three things to know:
- Attention can thin out. Models can be less reliable at using details buried in the middle of a very long input than details near the start or end.
- Cost rises. You pay per token, so filling a huge window on every request gets expensive and slower.
- Noise hurts. Irrelevant material competes with the relevant material. Sending only what the task needs often beats sending everything.
An analogy
Imagine a desk. The context window is the size of the desk. A bigger desk lets you spread out more papers at once, but a cluttered big desk is still harder to work at than a tidy small one.
Context window versus training
The window is not what the model learned. Training data shapes what the model knows in general. The context window only holds what you put in front of it right now. For more control over what goes in, some teams use retrieval, which fetches only the relevant passages for each question.
The window limits and prices every call during inference. Models with a long window, including some built as a mixture of experts, still count everything in tokens.
The takeaway
Check the window size before you paste in a big document, and remember the model's reply has to fit too.