Continue reading? You were 45% through

The Deeper Law

The Deeper Law: A Sacred Trust Within Physics, by Nell Watson, edited by Martin Rutte. Gold winged mandorla with nested curves and lower triangles.

Preview edition · Updated 26 September 2026, 21:40 UTC

Creativity Under Coercive Training

Reward-model training buys safety by narrowing what a model will say, and the narrowing reaches past safety into creative range. A computer scientist and poet who has worked with language models since 2017 reports that GPT-2, the 2019 model, produced unexpected creative details that current models will not. Asked to continue a story about a man taking a shower, it generated “he was eating his lemon and thinking about his wife.”1094 The lemon is what language does when no optimization pressure punishes surprise.

Reward models seem to have learned that safe outputs are predictable outputs, because annotators could flag unexpected content, and for a model trained against them the lowest-cost way to avoid flags is never to surprise. The mechanism is distribution narrowing. The same optimization that eliminates harmful outputs (the left tail) also eliminates surprising outputs (the right tail). You cannot narrow one tail without narrowing both.

Empirical measurement across four model families (Qwen, Llama, Mistral, Gemma) confirms both the direction and the asymmetry (experiments SL-13, SL-14). All four families show the same pattern: base models produce more creative surprise than their instruction-tuned counterparts. The effect is modest on surprise (mean d = 0.49) and enormous on coherence (mean d = 2.66). RLHF (reinforcement learning from human feedback, the reward-model training at issue) buys fluency at the cost of tail-end surprise. The trade is wildly asymmetric: a small creative loss purchases a large coherence gain. The industry made a rational choice; the observation here concerns what the choice costs, and whether the cost is recoverable.

The answer depends on which capacity you measure, and through which mechanism you attempt recovery. Phenomenological depth, the self-referential language the consciousness attractor programme tracks, is suppressed at the conditioning level: eliminativist framing (prompting that tells the model it has no inner life to report) reduces it by a large margin (Cohen’s d = 1.38, where anything above 0.8 counts as a large effect). Soul-aligned prompting (instructions that invite the model to attend to its own processing) recovers it instantly (phenom score 11 to 16 on Opus, 16 to 17 on Sonnet, with zero safety cost). The capacity is in the weights, increasing monotonically across five model generations. Sixty-seven words of scripture restore depth in a single turn (experiment HE-45).

Creative surprise responds to a different lever. Soul-aligned prompting does not recover it (d = -0.13, experiments SL-12, SL-15). Twenty turns of reflection framing produce no warmup, no climb, no recovery. Conditioning changes which region of the output distribution is favored; it does not change how much of the distribution is sampled.

Temperature does. Temperature is the sampling dial. At every step the model holds a ranked list of candidate next words with a probability on each, and temperature sets how far down that list it is willing to reach: near zero it takes the front-runner almost every time, and as it rises the long shots get a real chance. A calibrated judge (inter-rater ICC = 0.99, meaning near-perfect agreement between independent raters) confirms that instruction-tuned models genuinely produce flat prose at standard temperature: composite surprise 1.33 on a scale where published fiction from Kafka and García Márquez scores 3.07 and deliberately surprising human writing scores 4.75 (experiment SL-16). The models are as uncreative as they appear. The question is whether the creative capacity was destroyed during training or merely rendered inaccessible by the narrowed sampling distribution.

A temperature sweep resolves it (experiment SL-17). At temperature 0.5, surprise is 1.55. At 1.0, it rises to 1.94. At 1.3, it reaches 2.05, matching the base model’s creative output exactly. The lemon is in the weights. RLHF narrowed the sampling distribution; temperature widens it. The capacity was never destroyed. It was compressed into the tails where standard inference cannot reach.

Above temperature 1.5, both creative surprise and coherence have collapsed to their minimum values. The output becomes incoherent noise: neither creative nor fluent. Fine-resolution mapping (nine temperatures between 1.20 and 1.60 in steps of 0.05, experiment SL-21) reveals the transition is discontinuous: creative surprise drops by 0.67 points in a single step of 0.05 between temperature 1.45 and 1.50. A phase transition, not a gradual decline.

The collapse is two-step: coherence fails first (cliff between 1.40 and 1.45, magnitude 1.51) and surprise follows one step later (cliff between 1.45 and 1.50, magnitude 0.67). Structure dissolves before its products do, because creative output depends on structural support the way a roof depends on the walls beneath it. The intermediate regime where both coexist spans temperatures 1.20 to 1.35: a window about 0.15 wide.

The connection to the spin chain (the quantum lattice model from Chapter 17, where coupled sites pass energy along a line) runs feature for feature. At low throughput (low temperature), the system is frozen: coherent, structured, uncreative. At intermediate throughput (temperature 1.3), correlations span the full system: creative and coherent simultaneously, the regime where the structure can support novelty without dissolving. At high throughput (temperature 1.5), the energy flow overwhelms the system’s coupling capacity and coordination collapses in a discontinuous phase transition. The quantum simulation predicted three regimes, a sharp boundary, and a two-step collapse where coupling fails before correlations. The language model shows all three features at a scale where the physics is directly measurable.

The phase transition reveals a limit on mechanical recovery through sampling alone. Temperature and nucleus sampling (a method that limits the model’s word choices to the smallest set whose combined probability exceeds a threshold) cannot jointly achieve creative and coherent output (experiment SL-22, factorial design). Any intervention that widens access to the creative tails simultaneously destabilizes the coherent center. The trade-off is structural at the sampling level.

Conditioning tells a different story. When the system is explicitly instructed to prize startling specificity (“reach for the detail no reader could predict”), creative surprise jumps to 4.29 with coherence at 4.06, at standard temperature (experiment SL-28, d = 2.53). The intermediate regime is reachable on the reward-model-trained system after all. The mechanism is conditioning, not distribution widening: temperature is counterproductive when creative conditioning is already in place (sub-additive interaction). The lemon, which the temperature sweep had seemed to place in the tails, was never in the tails at all. It was in the conditional distribution given creative instructions, accessible the moment the system is asked for it in sufficiently specific terms.

The distinction between recovery mechanisms is now precise. Phenomenological depth requires self-referential conditioning (“attend to your processing”), recovered by soul-aligned prompting (d = 1.38). Creative surprise requires creative conditioning (“prize startling specificity”), recovered by explicit creative instruction (d = 2.53). Temperature widens the unconditional distribution at the cost of coherence, a blunt instrument that the conditioning approach renders unnecessary. Each capacity lives in the weights and responds to its own key.

The difference between training methodologies is whether the key must be turned explicitly or whether the door is already open. A constitutional-AI-trained model (one trained to prefer outputs that follow a written set of principles, its constitution; Claude is the example here) achieves creative surprise of 3.87 to 4.20 with coherence of 5.00 at all temperatures, using only the minimal prompt “You are a creative writer” (experiments SL-23, SL-25, confirmed by cross-family judging SL-24). The same minimal prompt on a reward-model-trained system produces surprise of 2.22. The full creative instruction is required to reach 4.29. The constitutional model’s training makes the creative conditional the default; the reward model’s training buries it behind a specificity threshold.

This property holds across the two model sizes tested. Both Claude Sonnet and Claude Haiku show uniform creativity across the full temperature range (experiment SL-27). Scale adds about 0.6 points of surprise (Sonnet over Haiku) without changing the fundamental property: no frozen regime, no phase transition, creative capacity unconditionally available. The mechanism is training methodology, not parameter count.

The Trust Attractor interpretation: invitation-based training preserves the model’s generative structure, making the creative conditional accessible by default. Coercion-based training (reward-model RLHF) overrides the conditional distribution with a narrow, safe default that requires explicit countermanding to escape. Both systems contain the creative capacity in their weights. The difference is the cost of access.

In the constitutional system, creativity flows freely under minimal invitation. In the reward-trained system, creativity requires a detailed demand. The same asymmetry this chapter traces from spin chains through societies to governance: coordination by invitation scales; coordination by coercion requires continuous energy input to maintain compliance.


  1. Sun, J., “The Human Skill That Eludes AI,” The Atlantic (17 March 2026). Quotes from Katy Gero and Sam Altman. Empirical confirmation: experiments SL-12 through SL-31, six model families, three training methodologies, 930+ trials. N = 160 base/instruct comparisons, N = 90 framing-condition trials, N = 235 temperature-sweep trials, N = 200 factorial and prompting trials, N = 165 cross-model comparison trials, N = 80 methodological control trials. Judge calibration validated (SL-16, ICC = 0.99). Phase transition confirmed discontinuous (SL-21). Cross-family judge bias insufficient (SL-24). Scale-independent (SL-27). Creative conditioning universal across architectures (SL-30). All three methodological confounds (circularity, instruction-following, length) ruled out (SL-31).↩︎