Context Rot: Why Long AI Sessions Get Worse, and What to Do

The longer a conversation with an AI model runs, the less reliably it uses what is in it. That effect has a name, context rot, and a short list of habits that keep it from spoiling a day’s work.

6 min read

AI context rot is the drop in quality you see as a model’s input grows: it recalls facts from the conversation less reliably, loses track of instructions given early on, and is more easily pulled off course by material that looks relevant but is not. It happens well before the context window is full, and a bigger window does not cure it. The practical response is to keep each session short and focused — clear between unrelated tasks, compact with a stated focus, start fresh after repeated corrections, push broad reading into subagents — and to keep the state of the work somewhere outside the conversation, so starting fresh costs you nothing.

What the research found

The term is best known from Chroma’s technical report Context Rot: How Increasing Input Tokens Impacts LLM Performance (opens in a new tab), published in July 2025. It tested 18 models, including GPT-4.1, Claude 4, Gemini 2.5 and Qwen3, on tasks that were deliberately kept simple so that only the length of the input changed. Performance fell as input grew, and it fell unevenly rather than in a straight line.

  • The less the wording of the question resembled the answer hidden in the text, the faster accuracy dropped as the text got longer.
  • Distractors — passages close to the answer but wrong — hurt, even one at a time, and more of them hurt more.
  • On a conversational memory benchmark, models did clearly better given a focused input of about 300 tokens than the full history of about 113,000 tokens that contained the same answer.
  • Even copying a sequence of repeated words became less reliable as the sequence got longer.

An earlier paper, Lost in the Middle (opens in a new tab) by Liu and colleagues, found a related pattern: models used information best when it sat at the beginning or end of the input, and significantly worse when it sat in the middle, including models built for long contexts. Put the two together and the lesson is simple. Having something in the window is not the same as the model using it.

Why it happens

Anthropic’s engineering team offers a working explanation in its post on effective context engineering (opens in a new tab). In a transformer every token attends to every other, so the number of relationships grows with the square of the input length, and the model’s attention is stretched thinner as the window fills. They describe context as a finite resource with diminishing marginal returns, and say models have an “attention budget” that every extra token draws on. You do not need the maths to act on it. More material means each piece gets less of the model’s focus.

What it looks like in a working session

In a coding agent, context rot rarely announces itself. Claude Code’s best practices (opens in a new tab) put it plainly: as the window fills, Claude may start “forgetting” earlier instructions or making more mistakes. The signs are familiar once you know to look for them:

  • A rule it followed all morning is quietly ignored in the afternoon.
  • It returns to an approach you rejected an hour ago, because the rejection is buried under two hundred tool calls.
  • It mixes up two files, two bugs or two branches it has read about.
  • Answers drift towards whatever it read most recently rather than what you asked.
  • Corrections stop sticking: you fix the same mistake twice and it comes back a third time.

The same documentation names two ways sessions get there. The kitchen-sink session starts on one task, wanders to something unrelated and comes back, so the window is full of material that has nothing to do with the job. The correction spiral fills the window with failed attempts, each of which is now context the model weighs.

What to do about it

  1. Clear between unrelated tasks. In Claude Code, /clear resets the window. A new job does not need the last job’s file reads.
  2. After two failed corrections, stop correcting. Clear, and write a better first prompt that includes what you learned. A clean session with a sharper prompt usually beats a long one full of rejected attempts.
  3. Compact with a focus, before you have to. /compact keep the failing test, the plan and the files changed gives you a summary you chose rather than one the automatic pass guessed at.
  4. Send broad reading to a subagent. It reads the thirty files in its own window and returns a summary, so your main session gets the conclusion without the bulk. See Task tool vs subagents.
  5. Plan in one session, build in another. Have the agent write the plan or spec to a file, then start a fresh session that reads it and does the work.
  6. Keep the standing instructions short. A long instructions file is itself a source of rot: rules get lost in the noise.

If the session has already hit the limit and stalled, that is a recovery problem rather than a quality one; Claude Code task stuck? covers the error messages and the commands that free space.

Keep the state of the work out of the chat

Every remedy above involves throwing conversation away, which is only painless if nothing important lives only in the conversation. Claude Code’s documentation on what survives compaction (opens in a new tab) is a useful guide here: the project-root CLAUDE.md, auto memory and a plan written in plan mode are re-read from disk, while the conversation is summarised. Whatever is on disk comes back whole; whatever was only said in chat comes back as a summary, or not at all.

So the habit that makes fresh sessions cheap is writing state down as you go. For a single task on one machine, a plan file is enough. When the work spans days, assistants or people, a task on a board does the same job where everyone can read it. On a fenbs board the task holds the note (what the problem is), the plan (how it will be done, rewritten as the agent learns), test notes and a comment thread, and a new session picks it up with fenbs_get_item. Standing knowledge — “the staging database resets nightly” — goes into the board’s AI context, which each assistant reads with fenbs_get_context when it starts.

Before you clear: an instruction worth giving
Before I clear this session:
- rewrite the plan on BUG-031 with what we now know
- comment what was tried, what failed, and the current error
- set test notes: what passes, what has not been run

Then /clear, and start the next session with “read BUG-031 and carry on”. The new window has the conclusions and none of the noise.

When a long session is fine

None of this means every session must be short. When you are deep in one hard problem and the history is genuinely the working material, letting it accumulate can be the right call. The test is relevance, not length: if most of what is in the window still bears on the next step, keep going; if most of it is about something finished or abandoned, it is costing you accuracy and should go.

Related

The wider discipline of deciding what goes into the window is covered in context engineering for AI agents. For a board that keeps plans and progress between sessions, see a task-tracking workflow for Claude Code and connecting Claude Code.

Questions people ask.

Does a larger context window fix context rot?

No. A larger window lets more fit before the limit, but the research on context rot found performance falling as input grew, on tasks the models handled well when the input was short. A focused input still tends to beat a long one containing the same information.

Is compaction enough on its own?

It helps, because it replaces a long history with a shorter summary. But a summary can drop details, and repeated compaction of a long, wandering session keeps summarising noise. Clearing between unrelated tasks and writing state to a file or task before you compact work better together.

How do I know when to start a fresh session?

Good signals are: you have corrected the same mistake twice, the agent has returned to an approach you rejected, it has started ignoring a standing rule, or you are switching to an unrelated task. Write down where things stand first, then clear.

Where did the term context rot come from?

It is best known from Chroma’s July 2025 technical report, Context Rot: How Increasing Input Tokens Impacts LLM Performance, which tested 18 models and found that performance degrades as input length grows, often unevenly.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.