Articles

What Is Context Rot in LLMs?

Context rot is the drop in an LLM's accuracy and reliability as the amount of text in its context window grows, even when the window technically fits it.

Chisato Chisato · · 4 min read
Abstract illustration representing a large language model

Context rot is the tendency for a large language model’s accuracy and reliability to degrade as more text is loaded into its context window, even when everything technically fits within the model’s stated limit. A model that answers a question correctly with 2,000 tokens of relevant context can start missing details, contradicting itself, or ignoring instructions once that same context is padded out to 80,000 tokens — not because the window overflowed, but because more content in the window doesn’t mean the model uses all of it equally well.

It’s not the same as running out of room

It’s worth separating context rot from the simpler problem of a context window filling up. Running out of room is a hard limit — text gets truncated or an error is thrown, and the failure is obvious. Context rot is quieter and more dangerous precisely because nothing fails loudly: the request succeeds, the model returns a plausible-looking answer, and that answer is wrong or incomplete because something in the middle of a long context got underweighted, misread, or effectively ignored.

This distinction matters for anyone building on top of an LLM. A window that advertises room for hundreds of thousands of tokens is a statement about capacity, not about the quality of reasoning at that capacity. Treating “it fits” as “it will be used correctly” is where a lot of production reliability problems in LLM applications actually come from.

What causes it

A few overlapping mechanisms drive context rot, and they compound as context grows:

  • Positional bias. Models attend most reliably to text near the start and the end of a long context, and are more likely to overlook details buried in the middle — the effect sometimes called “lost in the middle.” A fact placed at token 500 of a 50,000-token prompt is meaningfully less likely to be used correctly than the same fact placed in the first or last few hundred tokens.
  • Attention dilution. The attention mechanism inside a transformer distributes its focus across everything in the window. As the amount of content grows, the attention available to any single passage shrinks proportionally, even for passages that are directly relevant to the question being asked.
  • Distraction from irrelevant content. Padding a prompt with tangentially related or outright irrelevant material doesn’t just fail to help — it actively competes for the model’s attention with the material that actually matters, and can pull the model toward a wrong answer that superficially matches the noise.
  • Accumulated contradictions. In long conversations or long agent runs, earlier statements, outdated intermediate results, or corrected mistakes can still be sitting in the context. The model has no reliable way to know that a fact from ten turns ago was later superseded, and may cite the stale version.

Why it matters more for agents than for chat

Context rot is a background nuisance in a short chat exchange, but it becomes a first-order design problem in agentic systems. An agent that runs a long loop of tool calls, retrieved documents, and intermediate reasoning steps keeps appending to its own context with every step, as covered in the pattern behind building an AI agent. Left unmanaged, an agent’s context grows monotonically over a long task — full of tool outputs, retries, and dead ends — until the signal-to-noise ratio in that context degrades badly enough that the agent starts making avoidable mistakes, loses track of its own earlier plan, or repeats work it already did.

This is one reason agentic RAG designs deliberately re-retrieve and re-summarize rather than accumulating every piece of retrieved text forever, and why well-built agent frameworks periodically compact or prune their own history instead of letting it grow unbounded.

Mitigations that actually help

  • Retrieve narrowly instead of dumping broadly. Pulling in only the passages that are likely relevant, rather than an entire document or knowledge base, keeps the context focused and reduces the amount of competing, distracting content. A reranker applied after initial retrieval further narrows the field down to the passages most worth the model’s attention.
  • Summarize and compact aggressively. Rather than letting a conversation or an agent’s history grow forever, periodically collapsing older turns into a compact summary preserves the gist while shrinking the token count competing for attention.
  • Put load-bearing information at the edges. Given the positional bias models exhibit, the most important instructions and facts belong near the start or the end of the prompt, not buried in the middle of a large block of retrieved text.
  • Prefer targeted retrieval over maximal context when you can. The tradeoffs between stuffing everything into a long context versus retrieving just what’s needed are covered in more depth in RAG vs long-context LLMs — a bigger window is not automatically the better architectural choice.
  • Evaluate at the lengths you’ll actually use. A model’s benchmark scores on short prompts say little about how it behaves at the context lengths a real application will hit. Testing with an LLM eval built around your actual context sizes, not a vendor’s headline number, is the only reliable way to know where the degradation starts for your specific workload.

The takeaway

Context rot is the gap between what a context window can technically hold and how reliably a model actually reasons over everything inside it — a gap that widens as context grows, driven by positional bias, attention dilution, distraction from irrelevant content, and accumulated contradictions. A bigger context window is capacity, not a guarantee of quality, which is why the fix is rarely “just add more context” and usually the opposite: retrieve narrowly, summarize aggressively, keep the load-bearing facts near the edges of the prompt, and measure performance at the lengths you’ll actually run in production.