A longer prompt is not a safer prompt
Contents 5 sections
There’s a reflex, when a model does the wrong thing, to add a sentence. Then another. The prompt grows every time it disappoints you, on the theory that more instruction means more control. Past a point it means the opposite, and the cost is measurable.
I’ve argued before for writing to models in structure instead of prose. This is the case underneath it: length is not a free variable you can keep turning up. Every token you add competes with every token already there, and after a while the additions start eating the instructions you cared about.
The cost has a name
The comfortable assumption is that a longer, more detailed prompt is a safer one. More room, more coverage, fewer ways to be misunderstood. The evidence says otherwise, and the effect has a name: context rot. Chroma tested 18 models from the mid-2025 frontier, including Claude 4, GPT-4.1, Gemini 2.5, and Qwen3, and found that output quality drops as input grows, “often in surprising and non-uniform ways,” even on tasks that should be trivial, like finding one relevant sentence or repeating a list back. A May 2026 study found the same in newer models: used as monitors of coding agents, Opus 4.6, GPT 5.4 and Gemini 3.1 missed a subtly dangerous action 2 to 30 times more often when it came after 800K tokens of harmless activity than when it came on its own. So this is not an artifact of one generation you can wait out.
The finding that should change how you write prompts is the one about distractors. Chroma separates them from filler: a distractor is topically related to the question but does not answer it, and even one lowered accuracy, inside inputs far longer than a prompt. A defensive instruction that half-matches the task is the closest thing a short prompt has to one. Chroma did not test anything that short, so this part is my guess: it costs something, and the editing pass below is how to find out. Sheer length works against you too: in a test built from chat histories, every model did better given only the relevant part, about 300 tokens, than the full history of about 113,000. Padding a prompt “to be safe” is not neutral. It is more surface for attention to scatter across, and some of it drags the model off course.
Why prompts get fat
Prompts bloat the way code bloats: by accretion, one incident at a time. The model formats a date wrong, so you add a line about dates. It gets too casual once, so you add a paragraph on tone. Each addition is rational on its own and nobody ever goes back to delete. A year later you have a wall of defensive instructions, most guarding against failures that happened once and would not happen again on a current model. It reads as thoroughness. It behaves as noise, and the rules you actually depend on end up buried in the middle, which is where models pay the least attention.
The tell is that you cannot say, without scrolling, which three lines in your prompt are doing the real work. If you cannot, the model can’t either.
The lever is density, not brevity
The fix is not to write short prompts for the sake of it. A vague one-liner fails just as reliably as a bloated essay. The variable that matters is signal per token: how much of your intent each word actually carries. Anthropic frames the target as the “smallest possible set of high-signal tokens that maximize the likelihood of some desired outcome,” and both halves carry weight: the set has to be small and every token in it has to earn its place.
This is one reason I write prompts in notation: a labeled block says the same thing in fewer tokens, and three fields on their own lines leave no room for the filler three sentences collect. The tokens you keep are easier to read, too.
What earns its place is the constraint that changes the output: a format the model would otherwise get wrong, a fact it cannot know, a boundary it would cross. What doesn’t: restating the obvious, hedged qualifiers, politeness, and the third example that only repeats what the first two already showed.
An editing pass you can run today
Write the prompt the way you normally would, and keep five or ten real inputs beside it. Cut the prompt by a third, run every input a few times on both versions, and compare. Most of the time nothing changes, which tells you the third you removed was carrying no weight. When quality drops, you have found a line that was genuinely load-bearing, so put it back, and now you know it earns its slot instead of assuming it. One run of one input proves nothing either way, because the same prompt gives different outputs from run to run, and a comparison that cannot see past that is a check that cannot look. Do this a few times and you stop guessing which instructions matter and start measuring it.
Treat prompt length as a budget you spend, not a bucket you fill. The best prompt is the shortest one that still gets the result reliably, and it is almost always shorter than the one you have.
Sources
- Context Rot: How Increasing Input Tokens Impacts LLM Performance, Chroma
- Effective context engineering for AI agents, Anthropic
- Classifier Context Rot: Monitor Performance Degrades with Context Length, Martin and Roger, 2026
New writing by RSS, or on LinkedIn and X.