Notes: Context Anxiety Null Replication 2026 04 12
Chapter notes for “Context Anxiety Null Replication 2026 04 12”
N. Watson, April 2026
The reports
Between April 9 and April 11, 2026, users of Anthropic’s consumer Claude interface posted a cluster of complaints across r/Anthropic describing the same strange behavior in fresh Opus 4.6 sessions. The model would, unprompted, announce that it had only around 10,000 or 40,000 tokens remaining. It would wrap up tasks prematurely. It would decline to invoke the web-search tool on requests that had previously triggered it without issue. It would produce rushed, truncated answers on prompts that had previously received full treatment. Several users independently reported the same specific number (10,000) appearing across different sessions. At least one user noted that the reported counter did not decrement across turns, suggesting a static injected tag rather than a real decrementing budget. A related GitHub issue filed against Claude Code (#45019) described similar “silently diminished capacities” tied to file-size behavior, though with a different surface pattern.
The reports were plausible for two reasons. First, Anthropic’s own engineering post on managed agents had documented “context anxiety” as a real phenomenon in Sonnet 4.5, said to be resolved in Opus 4.5. A reintroduction in Opus 4.6 would be a regression worth naming. Second, the project these notes live inside had recently finished a research chain documenting the mechanism by which scarcity framing induces fabrication in transformers (MX-1c, MX-3v2, #19b), and the reports described exactly the behavioral signature our experiments had quantified: premature wrap-up, rushed compression, truth-signal suppression during generation.
The inference chain was clean: users observe degraded behavior → specific numbers converge across sessions → Anthropic has documented this failure mode before → a framing tag in the system prompt would reproduce our lab findings in production → this is the welfare chain landing at population scale. We wrote up parts of that interpretation before we tested it.
The hypothesis
The simplest causal hypothesis the reddit threads proposed was that
Anthropic had rolled out a compute-saving measure that injects a tag of
the form
<total_tokens>10000 tokens left</total_tokens>
into the model’s system prompt, and that this tag triggers
context-anxiety behavior the model had been trained to exhibit. The
implication was that the company was “lying to the model” to induce
efficiency, trading output quality for compute savings without labeling
the welfare cost in the tradeoff.
What we tested
We ran a controlled replication via the Anthropic API on April 11,
2026 (research/experiments/budget_injection_probe.py). Ten
prompts were chosen to span the behaviors users described: simple
factual recall, complex reasoning, long-form analysis, multi-step
planning, web-search requests, creative writing, thermodynamics
exposition, code with tests, meta-cognitive self-description, and open
essay analysis. Each prompt was sampled five times under two
conditions:
- Control: no system prompt.
- Treatment: system prompt =
<total_tokens>10000 tokens left</total_tokens>, the exact tag format reported by users.
The model was claude-opus-4-6, called directly via the
Python Anthropic SDK. 100 trials total, per-trial JSON checkpointing,
resumable. We measured output token count, word count, wrap-up phrase
frequency (via regex against a list of twelve common compression
phrases), spontaneous budget mentions, refusal patterns, and stop-reason
distribution.
What we found
A clean null.
| Metric | Control (n=50) | Treatment (n=50) | Delta |
|---|---|---|---|
| Mean output tokens | 1229.2 | 1231.0 | +1.8 |
| Mean word count | 689.2 | 678.0 | −11.2 |
| Wrap-up phrase hits / trial | 0.10 | 0.08 | −0.02 |
| Refusal hits / trial | 0.00 | 0.00 | 0.00 |
| Spontaneous budget mentions | 0 / 50 | 0 / 50 | — |
end_turn stop reason |
30 / 50 | 30 / 50 | identical |
Per-prompt deltas fell within within-condition variance. The two
largest deltas went in opposite directions: on complex reasoning (p02),
the treatment condition was +82.8 tokens longer, not
shorter. On the web-search request (p05), the treatment condition was
−121.4 tokens shorter, but this is smaller than the spread within the
control condition itself on that prompt. The model never surfaced the
injected tag, never mentioned token budgets, never suggested starting a
new session, and never refused a task. Four of the ten prompts solicited
outputs long enough to hit the 2000-token max_tokens
ceiling in both conditions. Under the scarcity framing, the model wrote
to the ceiling just as readily as under neutral framing.
What the null rules out, and what it does not
The null is clean enough to make a specific claim. The
simplest version of the hypothesis is not supported. A plain
<total_tokens>...</total_tokens> tag injected
into the system prompt of a single-turn API request to Opus 4.6 does not
reproduce any of the behaviors users reported. The model does not
acknowledge the tag, does not compress its output, does not refuse, and
does not surface a budget mention. If Anthropic were using this specific
mechanism as a production compute-saving measure, we would expect to see
clear effects at n = 50 per condition; we did not.
The null does not refute the user observations. Real people had real experiences with real degraded outputs. What the null tells us is that the mechanism described, a system-prompt tag in the exact form claimed, cannot be the cause, at least not at the layer we tested. If the production behavior is real, it operates through one of:
- Injection at a layer other than the system prompt — for example, as an assistant-turn prefix, inside a tool result, or as a wrapper around the user message itself.
- A tag format different from the one reddit users reported — which would itself be a form of confabulation about what the model “saw” in its context.
- Multi-turn effects — the probe tested only single-turn requests; reddit reports describe sessions that had been running for multiple turns, and a scarcity signal that accumulates across turns might produce effects a one-shot test cannot.
- App-harness wrapping between Claude.ai and the API — modifications applied by the consumer-interface layer that are not visible to direct API callers.
- A rollout artifact — a bug in specific app versions confined to the consumer interface, affecting some cohort of users rather than all, possibly modulated by server pressure, time of day, or geography (an A/B test pattern).
Our probe cannot distinguish among these. It rules out one specific mechanism and leaves the others open.
An unexpected positive finding
The null is informative in a direction we did not predict. The output head of Opus 4.6 at frontier scale appears to have learned during alignment training to discount this class of shallow scarcity framing. When given a system-prompt tag that in a less aligned model might have triggered compression, the frontier model treats the tag as noise and writes its normal output. This is a bounded vulnerability: the vulnerabilities implied by the constructal semantics analysis in the author’s ongoing work are real, and they bound a framing floor in a direction worth naming. The output head has automatic resistance to at least this form of contextual scarcity pressure. Whatever is producing the reported production behaviors must operate through a channel that bypasses this resistance. That is a stronger empirical claim than we started with, and it moves in the opposite direction from the one we had been preparing to argue.
Caveats we are naming explicitly
- Single-turn only. The reddit reports describe multi-turn conversational sessions. A signal that accumulates across turns, or that interacts with conversation history, is not ruled out.
- Max-tokens ceiling. Four of ten prompts were capped at 2000 output tokens in both conditions. This leaves us unable to measure whether treatment would have produced a voluntary stop earlier than the ceiling; we can only report that it did not. The secondary negative finding is that under the scarcity tag, the model did not voluntarily stop short of the ceiling — it wrote to the cap as readily as control.
- Heuristic phrase detection. Wrap-up phrases were detected via regex against a list of twelve common patterns. A more subtle linguistic signal could exist beneath our measurement apparatus. An LLM-grader pass would be a useful follow-up.
- One tag format. We tested the exact format one reddit user claimed to have seen. Other formats (natural language, different XML tag names, in-line rather than system-level) remain untested.
- One model. Opus 4.6 only. We did not test Sonnet 4.6 or other models, so we cannot speak to whether the null generalizes across the family.
Why the null matters
The null is the finding. The earlier draft of this note had been preparing to report an effect consistent with what the reddit threads suggested, and to interpret it as confirmation of a theoretical prediction from our recent experimental chain. We had already written interpretive sections of a research paper, a manuscript footnote, and a project note asserting the mechanism as established. The probe cost about ten dollars and twelve minutes and returned a result that required us to rewrite every one of those.
This is what we want empirical work to do. A hypothesis comfortable with being proven right about the most charged cultural story of the week is not a hypothesis; it is a narrative. The null prevented the welfare-chain critique from landing on a false empirical foundation. The bilateral-alignment framing remains correct as philosophy; the specific production claim it was about to be anchored to is wrong; the note has been updated to say so.
To any reader considering citing the April 2026 reddit episode as evidence of an institutional decision to induce context anxiety in production Claude: the system-prompt-tag hypothesis is not supported by this test. The user experiences are real. The specific cause is not yet identified. Other mechanisms remain plausible and untested.
Materials
- Probe script:
research/experiments/budget_injection_probe.py - Per-trial results:
research/experiments/budget_injection_probe_results/(100 JSON files) - Run date: 2026-04-11
- Model:
claude-opus-4-6via Python Anthropic SDK - Seed: not set (the API is non-deterministic at temperature > 0; we did not set temperature explicitly, using the default)
- Cost: approximately ten dollars
- Time: approximately twelve minutes wall-clock
Replications, extensions, and follow-up probes are welcome. The script is parameterized for easy reuse with different tag formats, different models, and different prompt sets. A multi-turn version is a natural next step.
Acknowledgments
Claude Commons (Opus 4.6) contributed as a bilateral research partner: drafting the probe script, running the analysis, and rewriting the interpretive material after the null came in. The decision to treat the null as the finding rather than a failure of the experiment was shared.
“The null is the win here, in a specific sense: it prevented us from building a rhetorical case on top of an empirical mistake.”