2 AM PAGER

Your support bot is hallucinating. Blame the dumpster, not the model.

2:14 AM. The bot is inventing refund policies. The retriever is fine and the prompt is fine. The problem is the fourteen thousand tokens of dumpster you fed it.

2026-09-21

Tuesday, 2:14 AM. Your pager goes off. The support bot is hallucinating refund policies that do not exist, and customers are quoting them back at your team.

You check the retriever. It is fine. Top chunk, exact match, the right policy doc. You check the prompt. Fine. So you look at what you actually fed the model: fourteen thousand tokens of ticket history, KB articles, and system instructions. A dumpster.

The model did not break. It drowned. And the research says drowning is what models do when you keep pouring tokens in.

(The 2 AM incident is a story. The studies below are real.)

Longer inputs, worse answers, every model

Chroma ran the experiment at a scale that ends the debate: 194,480 calls across 18 models, same tasks, longer inputs. Performance degraded on every single model. Not some. Every one.

194,480
calls across 18 models in Chroma's Context Rot study (Jul 2025). Same tasks, longer inputs, degraded performance on all 18.

The key design choice: they held task difficulty constant and varied only the input length. That isolates the variable everyone wants to blame on something else. It is not harder questions. It is not worse retrieval. It is the length itself. More tokens in, worse answers out, across GPT, Claude, Gemini, and Qwen families.

So when your vendor announces a bigger context window, read it as a bigger bucket, not a better reader. The bucket got bigger. The reader did not.

Perfect retrieval still rots

Du et al. went further, and their result is the one that should change how you build. They gave the models perfect retrieval: the exact evidence, guaranteed present, zero distraction. Then they replaced everything irrelevant with whitespace. Then they masked the distractions entirely and forced the model to attend only to the relevant tokens.

Performance still fell 13.9 to 85 percent as inputs grew. Even with nothing to distract it, even with attention forced onto the right tokens, the sheer length of the input hurt.

13.9-85%
performance drop with perfect retrieval as input length grew to 30K tokens. Du et al., Findings of EMNLP 2025. Retrieval was not the bottleneck. Processing capacity was.

Read that again, because it kills the most common excuse in production AI. “Our retrieval is bad” is not the whole story and sometimes it is not the story at all. The transformer has finite capacity to integrate information. Past a point, more context is not more information. It is more noise the model has to carry while doing the same reasoning.

Their mitigation is almost insultingly simple: make the model recite the retrieved evidence before answering. That buys back a few points. A patch, not a cure.

Position matters too

Liu et al. found the U-shaped curve: models read best at the start of the context and at the end, and go blind in the middle. GPT-3.5-Turbo’s multi-document QA performance dropped more than 20 percent when the answer sat in the middle of 20- and 30-document inputs. At its worst, the model did worse with documents than with no documents at all. The context was actively hurting.

So your fourteen thousand tokens were not just long. The important parts were probably in the middle, exactly where the model cannot see them. You paid for the tokens and got negative value.

>20%
performance drop for GPT-3.5-Turbo when the answer sat mid-context. Liu et al., Lost in the Middle (arXiv 2307.03172).

The window was never the product

Three studies, one conclusion: stuffing more context into the prompt is not a strategy. It is a hope. The fix is the same discipline as everything else on this site.

Ticket closed, 3:07 AM. The bot stopped hallucinating when it stopped reading the dumpster.

This post started as an Instagram post →

Discussion

Talk it through

Argue with us on Instagram.