Context Window

What Is a Context Window and Why It Runs Out

October 2, 2026

A context window is the maximum amount of text an AI model can hold in view at one time, measured in tokens, and it runs out because everything competes for the same fixed budget your instructions, the whole conversation so far, any files you uploaded, and the reply the model is currently writing. When the total exceeds the limit, the oldest material is pushed out, and the model can no longer see it.

Most explanations stop at “it’s the model’s working memory”. That’s a fine analogy but it doesn’t answer the second half of the question. This article covers why a limit exists at all it’s a consequence of how attention works, not an arbitrary cap and what to do when you hit one.

How you know you’ve hit it

The limit rarely announces itself. These are the signs, roughly in the order they appear.

  • The assistant forgets a rule you set at the start of the conversation a tone, a format, a constraint and reverts to defaults.
  • It starts repeating points it already made.
  • It asks for information you already gave it.
  • Details from early in the thread vanish, while recent ones stay sharp.
  • Responses get slower before they get worse.
  • Eventually, an explicit message that the conversation is too long, or a refusal to accept the next file.

The dangerous phase is the first five. There’s no warning and no error. The model doesn’t know something was removed; it simply never sees it, and answers confidently with what’s left.

What’s actually taking up the space

Everything in the exchange shares one budget:

The system prompt. Hidden instructions from whoever built the product. You never see them, and they can run to thousands of tokens.

The full conversation. Every message from both sides, re-sent in its entirety each turn. Models have no memory between messages, so the transcript is the only thing keeping continuity.

Uploaded files. A 20-page PDF can be 15,000 tokens or more, and it stays in the window for the rest of the conversation.

Retrieved material. If the tool searched the web or a document store, those passages were inserted into your context before the model saw your question.

Tool definitions. Descriptions of every capability the assistant can call.

The reply being written. This is the one people miss, and it matters.

What is Context Window

Output comes out of the same budget

The context window is not “how much you can send”. It’s the total for input and output together.

Fill 95% of a window with a long document and there’s very little room left for an answer. This is why a model sometimes gives a thin response to a large upload, or stops mid-sentence. It hasn’t lost interest it’s out of room.

Providers usually impose a separate, smaller cap on output length as well, sitting inside the overall window. So the practical formula is: what you can send equals the context window minus your expected answer length, minus whatever the application has already spent on its own instructions.

Why there’s a limit at all

Here’s the part the keyword asks about and most articles skip.

Transformer models use self-attention, which means that when processing any given token, the model weighs its relationship to every other token in the window. That’s what lets it connect a pronoun to a noun forty sentences earlier.

The cost of doing this scales with the square of the length. Double the context and you roughly quadruple the attention computation. Go from 10,000 tokens to 100,000 and you’ve multiplied it by around a hundred. Various engineering techniques reduce the constants, but the underlying pressure doesn’t go away.

There’s a memory cost too. As the model generates, it caches intermediate values for every token already processed so it doesn’t recompute them. That cache grows with every token and has to sit in GPU memory alongside the model itself. Long contexts are expensive in hardware terms, not just in billing.

And there’s a training limit. Models are trained on sequences up to a maximum length, and their sense of position is learned within that range. Pushed beyond what they were trained on, quality degrades even where the software technically allows it.

So the window isn’t a policy decision someone could simply raise. It’s a point on a curve balancing capability, hardware cost and response time.

Why bigger windows didn’t solve it

Context windows have grown enormously from a few thousand tokens a few years ago to hundreds of thousands, and over a million on some models. That helped. It didn’t fix the problem, for three reasons.

Attention isn’t evenly distributed. Models reliably use material at the very start and the very end of a long context, and are markedly worse with what sits in the middle. Researchers call this the lost-in-the-middle effect. A fact buried at the 50% mark of a huge document may as well not be there.

More context can mean worse answers. Irrelevant material competes for attention with the instruction that matters. A focused 5,000-token prompt frequently beats a sprawling 100,000-token one on the same task.

It costs and it’s slow. You’re billed on every token in the window, every turn. Large contexts also increase the delay before the first word appears.

Advertised capacity and useful capacity are different numbers. Treat the headline figure as a ceiling, not a target.

Context window, memory, and retrieval

Three things get confused. They’re distinct.

The context window is temporary and scoped to one conversation. Close the chat and it’s gone entirely.

Memory features in consumer products work by storing notes about you in a separate database and re-inserting the relevant ones into the context at the start of a new conversation. The model still remembers nothing; the application is pasting things in.

Retrieval searches a document store or the web at question time and inserts only the relevant passages. For a practical look at how modern AI tools use these capabilities, see our Genspark AI review.

What to do about it

Practical, in rough order of effectiveness.

  • Start a new conversation when the topic changes. The single biggest win. Twenty turns about something unrelated are pure dead weight on every subsequent message.
  • Restate critical constraints. If a rule matters, repeat it in your latest message rather than trusting that turn three is still visible.
  • Put the important thing last. Recency gets attention. Your key instruction should sit near the end of your message, not buried in the middle of a wall of pasted text.
  • Upload the relevant pages, not the whole document. Extract the chapter you need.
  • Summarise and restart. Ask for a summary of the conversation, open a fresh chat, and paste it in. You keep the substance and drop the bulk.
  • Use a retrieval-based tool for large corpora rather than fighting the window.

There’s no token counter in most consumer chat apps, so you’re working blind. A rough proxy: an average exchange runs a few hundred to a couple of thousand tokens, so a long working session gets into six figures faster than it feels like it should.