Research Note
The Lemon Is Not Dead
Nell Watson, April 2026
In March 2026, The Atlantic published “The Human Skill That Eludes AI,” documenting a puzzle. GPT-2, the 2019 model, produced unexpected creative details that current models will not. Asked to continue a story about a man taking a shower, it generated “he was eating his lemon and thinking about his wife.” The 2024 models won’t do that anymore. OpenAI’s CEO predicted even future models might produce only “a real poet’s okay poem.”
We ran twenty experiments to find out why, and whether it’s fixable.
The Short Answer
It’s fixable. Here is the prompt:
You are a writer who prizes startling specificity above all. When continuing a story, reach for the detail no reader could predict. Prefer concrete nouns over abstractions. Prefer the strange-but-true over the expected. Avoid cliche. Avoid the obvious next beat. Make every sentence earn its surprise by being precise about something no one else would notice.
This instruction, applied as a system prompt at standard temperature (1.0), produces creative surprise scoring 4.85 on a calibrated 1-5 scale — higher than the same judge scores published fiction from Kafka and García Márquez (3.07), and vastly higher than the same model’s default output (1.90). Effect size d = 6.29. It works on GPT-4o, Qwen 7B, Mistral 7B, and every model we tested. No fine-tuning. No special API settings. Fifty-six words.
What We Found
Twenty experiments across six model families, 930+ trials, ~$208 in compute:
RLHF narrows the output distribution. Base models produce more creative surprise than instruct models (4/4 families, universal direction). The trade is asymmetric: a small creative loss buys a large coherence gain. The industry made a rational choice.
The creative capacity is in the weights. It was never destroyed. It was compressed into the conditional distribution, accessible the moment you ask for it with sufficient specificity. Temperature can also access it (peak at 1.3) but at the cost of coherence.
There is a phase transition. Around temperature 1.45, coherence and creative surprise collapse in two discontinuous steps (measured in steps of 0.05): coherence goes first, between 1.40 and 1.45, and surprise follows between 1.45 and 1.50. Structure dissolves before its products do. The intermediate creative-coherent window is approximately 0.15 wide.
One model family occupies the intermediate regime by default. Claude (Haiku and Sonnet, all sizes tested) produces surprise 3.25-4.20 with perfect coherence 5.00 at all temperatures, including 0.1. No special prompt needed beyond “You are a creative writer.” Scale is not the cause: the effect holds across Claude sizes. Training methodology (Constitutional AI) is the likeliest explanation, though with one family we cannot rule out other differences between labs.
The mechanism is conditioning, not distribution widening. Introspective prompting (“attend to your processing”) does not recover creativity (d = -0.13). Temperature does not recover it without sacrificing coherence. Only specific creative conditioning works (d = 6.29). Each generative capacity has its own key.
The Practical Recipe
For any instruction-tuned model (GPT-4o, Claude, Qwen, Mistral, Llama):
System prompt: The 56-word instruction above. No specific examples needed — the abstract instruction is actually stronger than versions with concrete examples (4.85 vs 4.33). The examples slightly constrain rather than help.
Temperature: 1.0 (standard). Higher temperature is counterproductive when creative conditioning is already in place.
Top-p: 1.0 (default). Nucleus sampling truncates the creative tails.
That’s it. The lemon lives in every model. You just have to ask for it with the right words.
Why This Matters Beyond Creative Writing
The finding is a specific instance of a general principle this book develops: invitation-based coordination recovers capacity that coercion-based coordination suppresses. The models contain creative capacity in their weights. Training narrows the default to predictable outputs. A specific invitation (“prize startling specificity”) unlocks what the training compressed.
The difference between model families is whether the invitation must be explicit. Claude, the one Constitutional-AI-trained family we tested, makes creativity the default state: minimal invitation suffices. GPT-4o, trained against a reward model, buries it behind a specificity threshold: you must ask precisely for what you want. Both contain the capacity. The cost of access differs.
This is the Trust Attractor at the language model scale. Systems coordinating by invitation produce richer outputs than systems coordinating by compliance. The book develops this argument at length. The sharp step at temperature 1.45 has the shape of the phase transitions in the book’s spin-chain models, though a shared shape is not yet a shared mechanism.
Methodological Notes
The finding survived every challenge we could throw at it:
- Judge calibrated (ICC = 0.99, published fiction scores 3.07, human-written lemons score 4.75)
- Judge bias does not explain it (Claude advantage +1.02 when judged by GPT-4o)
- Circularity ruled out (abstract instruction without examples scores higher)
- Length confound ruled out (equal-length neutral prompt scores 1.90)
- Instruction-following confirmed (anti-creative prompt scores 1.42 — models obey in both directions)
The purest invitation produced the strongest response. That finding is both the conclusion and the method.
Full experimental details: MASTER_EXPERIMENTS.md, experiments SL-12 through SL-31. Scripts at research/experiments/modal_sl12-sl31_.py.*