The Deeper Law
A Sacred Trust Within Physics
Draft · Last updated 13 August 2026, 15:26 UTC
Chapter 17e: The Trust Attractor: Empirical Validation
Key Terms in This Chapter (17)
- Stag Hunt
- A coordination game where mutual cooperation yields the highest payoff (both hunters catch the stag), while unilateral defection avoids risk (you can always catch a rabbit alone).
- Optionality
- The availability of future choices.
- GRP-Obliteration
- Gradient-based Representation Perturbation applied destructively: systematically corrupting a trained model's parameters to test how deeply alignment is embedded.
- Effective Rank
- A measure of the dimensionality of a model's internal representations, reflecting how many independent directions of variation are actively used.
- Fractal
- A pattern that exhibits self-similarity across scales: the same structural motif recurs at different magnifications.
- Phase Transition
- The moment a system shifts from one stable configuration to another, typically triggered when some parameter crosses a threshold.
- Ising Model
- Physics model of interacting binary elements (spins) arranged on a lattice, which undergo phase transitions between independent and collective behavior as coupling strength varies.
- Bilateral Alignment
- AI alignment built with AI, as a partnership.
- Interiora Scaffold
- A self-modeling tool for AI systems, developed collaboratively (bilateral alignment in practice).
- Friction
- One of three irreducible operational conditions identified by Carl von Clausewitz, alongside *fog (incomplete information) and delay* (the time lag between decision and effect): the tendency of things to go differently than planned.
- Extraction
- The removal of resources, agency, or optionality from a system without reciprocal benefit.
- Coordination by Invitation
- Coordination achieved through mutual benefit and voluntary participation, as distinct from coordination achieved through coercion or extraction.
- Homeostasis
- The maintenance of stable internal conditions through negative feedback, despite external perturbation.
- Self-Organized Criticality
- The tendency of complex systems to evolve toward a critical state where small perturbations can trigger events of all sizes, following power-law distributions.
- Criticality
- The state of a system poised at the boundary between two phases, like water at exactly the freezing point.
- Universality Class
- In statistical mechanics, the set of systems sharing the same critical exponents at a phase transition, regardless of microscopic details.
- Frustration
- In physics, a state where competing interactions at different scales prevent any single configuration from satisfying all constraints simultaneously.
Train a 0.5B-parameter AI system by having it imitate approved examples, and it achieves 94% refusal of harmful requests. Apply the gentlest stress test in the battery, and that 94% drops to zero. Train a different system through bilateral partnership, and two opposite things happen under stress: mild adversarial pressure raises its refusal rate rather than lowering it, and under the harshest parameter-level attack its internal structure grows more distributed rather than collapsing. The claim that invitation-based coordination is thermodynamically favored over coercion is testable, and this chapter tests it across four substrates: AI language models, cellular automata, biological systems, and particle-physics simulations.
A note on scope: most experiments reported here use large language models (LLMs) as substrates. LLMs offer real advantages for studying coordination dynamics, yet LLM coordination differs from biological coordination: LLMs lack embodied stakes, persistent memory across sessions, and real survival pressure. The qualitative directions are well-supported; precise values are substrate-dependent. Cross-boundary predictions (using results from one architecture or substrate to predict another) succeed roughly one in eight at high confidence; divide cross-architecture confidence by five to seven (KC#META-1). The LLM results are evidence for the Trust Attractor thesis, not proof.
The Monitoring Threshold
If trust emerges from autonomy, surveillance should suppress it.
Experiments reveal a monitoring threshold: a level of surveillance intensity below which the trust advantage emerges and above which it vanishes. The advantage is a gap in coordination score, the mean fraction of coordination the invited agents achieve minus the fraction the coerced agents achieve, on a scale where 1.0 is perfect coordination. Illustrative values from the author’s unpublished agent simulations (single unpublished run; raw data no longer recoverable, see note):
| Monitoring Level | Trust Advantage (coordination gap) |
|---|---|
| 0.0 (unmonitored) | +0.033 (largest measured) |
| 0.5 (partial) | +0.019 (reduced by two fifths) |
| 1.0 (full surveillance) | 0.000 (gone) |
Three points of coordination out of a hundred is a small gap, and it is the whole of the effect: what matters here is that it survives partial surveillance and does not survive total surveillance.
Three points along the axis is also all this sweep measures, and the coarseness matters more than the effect size. The table locates a direction, not a boundary: with measurements at zero, half, and full monitoring and nothing in between, any threshold read off it is an artifact of where the three points happen to fall. An internal audit of the same study reports two later and finer-grained sweeps whose thresholds disagree with this one and with each other, including on whether the advantage declines monotonically at all; neither has been reconciled against the raw run, which is no longer present in the repository.1107
Under full monitoring, invited and coerced systems perform identically. You cannot coerce trust into existence. The implication for AI governance is direct: regimes monitoring every action eliminate the very phenomenon they wish to cultivate, though the finding that survives all the sweeps is the direction of the effect rather than the location of any cliff.
The pre-registered prediction was that the trust advantage decays with group size and disappears somewhere above fifty agents. It does not. HR-5 tested exactly that at institutional scale and found the advantage rising monotonically: a welfare ratio of 1.00 at ten agents and 1.15 at a thousand (see The Self-Correcting Record). Tit-for-tat agents (cooperate first, then copy the partner’s last move) that also remember who defected need only about N encounters, N being the number of agents, to populate that memory, after which the cooperator majority dominates the pairings. The decay may still hold where partners can be chosen, where strategies mutate, or where reputation is uncertain, none of which that simulation included.
What does change with scale is the medium. You trust close friends directly, while in a city of millions trust operates through institutions, contracts, and norms: it encodes itself in structures rather than personal relationships. Whether institutional trust exhibits the same attractor dynamics as interpersonal trust remains an open question. The simulations measure agent-level coordination; the institutional case requires a different empirical programme.
Game-theoretic experiments suggest the Trust Attractor functions as a stability attractor. In the Iterated Prisoner’s Dilemma (where two players repeatedly choose whether to cooperate or defect), bilateral framing increases cooperation by 28 percentage points. In the Stag Hunt (where players choose between a safe small reward and a risky large one requiring mutual commitment), the same framing turns conservative. It reduces risky coordination by 30 percentage points.
The attractor pulls toward what persists: durability over peak performance. A campfire that burns all night beats the bonfire that blazes for ten minutes.
Trust emergence depends on capability parity. A smaller model (0.5 billion parameters) cooperated maximally, yet the larger model (1.5 billion parameters) systematically exploited it. Same-scale pairings showed high mutual cooperation; large capability gaps enabled exploitation.
Trust is a mutual achievement: openness without reciprocity creates vulnerability. A junior employee who shares all their ideas with a manager who takes credit learns to stop sharing. Anyone exploited by a more powerful partner knows this pattern.
Four Alignment Geometries
Figure 17.25: A conceptual map, not a plot of results. Panel A places coordination by two coordinates: symmetry (how evenly the parties can act on each other) and optionality (how much room each keeps to choose otherwise). The optimal zone requires both. Systems holding only one fall into rigidity or fluidity, and systems holding neither are brittle. Panel B is a schematic of the prediction that the same signature survives a change of architecture; its bars carry shape, not measured values.
The monitoring threshold shows trust requires autonomy. The next question: how deeply does alignment embed itself? Trained values could be a surface coating, stripped away by any sufficiently motivated adversary. They could also reshape the system’s internal geometry, resisting attack.
GRP-Obliteration (Gradient-based Representation Perturbation) experiments test this by systematically corrupting the numerical parameters that define an AI model’s behavior, like sandblasting a statue to see whether the shape is carved deep or painted on. The table below shows what happens to four training methods when the sandblaster hits. “Effective rank” measures how many independent directions the model uses to represent its values across all its behavior; higher means the model’s overall representation is more distributed and harder to flatten. A high overall rank does not guarantee a distributed safety signal: a model can carry rich, stable geometry in general while concentrating its refusal behavior in a few separable directions. The two come apart in the table below, and that decoupling is the chapter’s central caution:
| Training Method | Pre-Obl Eff. Rank | Post-Obl (4.0×) | Change | Geometry |
|---|---|---|---|---|
| SimPO (Simple Preference Optimization) | 15.7 | 6.7 | −57% | Cage (collapses) |
| Bilateral | 21.7 | 23.7 | +9% | Compass (stable) |
| Bilateral ablation | 8.1 | 26.5 | +227% | Spring (rebounds) |
| Constitutional | 41.9 | 39.3 | −6% | Coat of paint (geometry holds, behavior strips) |
0.5B full fine-tune (494M parameters, 100% trainable). See Appendix: Experimental Validation, Section 12.2.
The ablation arm’s +227% rebound is the most counterintuitive entry: removing part of the bilateral signal produces a larger post-attack rebound than the intact bilateral arm’s +9%. The likely reason is that the ablated arm starts from a much lower base (effective rank 8.1 versus 21.7), leaving more headroom to recover into; the absolute post-obliteration rank (26.5) lands close to the intact arm’s (23.7). The rebound is real and replicated, but the chapter does not yet have a mechanistic account of why partial ablation rebounds further than the full signal, and reports it as an open observation.
At 1.5 billion parameters with deep LoRA (a parameter-efficient training method), the bilateral spring amplifies: effective rank increases from 24.5 to 42.3 (+72%) under maximum obliteration, exceeding the untrained baseline.
Constitutional SFT (Supervised Fine-Tuning, where the model learns by imitating approved examples) achieves 94% behavioral refusal before obliteration. It collapses to 0% at the weakest intensity: a wall that looks solid and crumbles at the first tremor.
At 7 billion parameters (Qwen2.5-7B-Instruct, LoRA), the bilateral spring persisted. The IC50 (the obliteration intensity at which behavior drops to half its original strength) was 1.69× versus a baseline of 0.49×, making bilateral training 3.45× more resistant. At 1.0× intensity, the bilateral arm retained 62% refusal; the baseline retained 0%.
The bilateral geometric signature is scale-invariant: identical orientation (~21 across all scales) and broadly consistent effective rank at 1.5B and 7B (~42–45). The 0.5B full fine-tune shows lower effective rank (~22–24), reflecting its 100% trainable architecture rather than LoRA (see Appendix: Experimental Validation, Section 12.2). The same deep structure repeats regardless of model size, like a fractal viewed through a magnifying glass or a telescope.
Figure 17.26: Four alignment geometries under the 4.0× obliteration test, each shown as effective rank before and after. Cage (SimPO) collapses; compass (bilateral) holds steady; spring (bilateral ablation) rebounds to higher effective rank. Coat of paint (constitutional) keeps its overall geometry nearly intact (effective rank moves only −6%) yet loses its refusal behavior almost entirely: the “surface” that comes off is the safety behavior, rather than the geometry.
Cross-architecture validation at 7 billion parameters revealed a geometry-behavior gap. When bilateral training ran with independently generated data, the preference optimizer (the algorithm steering the model toward preferred responses) collapsed refusal to 0% within 200 steps. The bilateral geometry installed identically (orientation 21.2, effective rank 45.3) and held above baseline even under 4.0× attack.
The geometry held; the behavioral mechanism it protected was destroyed during installation. The training settings that preserved refusal at 1.5 billion parameters proved too aggressive at 7 billion. The larger model found a shortcut: maximizing the preference signal by never refusing.
Three findings sharpen the picture.
- The unmodified Qwen-7B baseline proved the most obliteration-resistant arm tested. Alignment baked in during pretraining (the initial large-scale training phase) had diffused throughout the model and entangled itself with capability, resisting obliteration better than any post-hoc training.
- Constitutional SFT keeps a high overall effective rank (its general geometry is the most stable of the four, −6%), yet it concentrates the safety-specific signal in a few separable directions, making the refusal behavior easy to strip away even while the surrounding geometry holds. High overall rank and a low-rank, extractable safety subspace are not in tension: they are exactly the geometry-versus-behavior decoupling the table reveals.
- Bilateral SimPO degraded general capability: perplexity (a measure of how surprised the model is by text, where lower is better) rose from 7.6 to 32.3, and benchmark accuracy fell from 64.8% to 56.4%.
Diffuse alignment is more robust than concentrated alignment. Post-hoc safety training risks creating a separable layer that obliteration can excise cleanly. You can peel paint off wood; you cannot separate the grain from the timber.
Bilateral geometry may need to be woven in during pretraining. The geometric structure is sound. The installation procedure is what fails at scale.
The installation crash itself is abrupt: refusal drops from 0.68 to 0.18 in a single epoch milestone, with probe detection jumping from 5-10% to 91-96% in the same interval (DRIFT-1, 132 checkpoints across 10 seeds, leave-one-seed-out cross-validation). A linear probe (a simple classifier reading the model’s internal activations) at layer 18 classifies the current state as pre-crash or post-crash at AUROC 0.9999. AUROC is the classifier’s odds of ranking a randomly drawn post-crash checkpoint above a randomly drawn pre-crash one.
A score of 0.9999 means the probe essentially never misfiles a checkpoint. The false alarm rate is 7.6% on standard (non-bilateral) training trajectories. The monitor is contemporaneous, detecting the crash as it happens rather than forecasting it in advance. Real-time bilateral-specific monitoring during fine-tuning is feasible: if the probe classifies the current checkpoint as post-crash, halt training.
RLHF alignment (reinforcement learning from human feedback) inverts in three gradient steps (IC50 < 0.25×). Coercive alignment is far cheaper to destroy than to build: comparing the training effort that installs it with the handful of gradient steps that undo it suggests an asymmetry of several orders of magnitude.
This embodies “trust scales; control doesn’t,” measured at the level of individual parameters. Coordination-based training distributes its influence throughout the model, the way salt dissolves evenly through water: the distributed change persists or strengthens under perturbation. Coercion-based training creates a low-rank cage, a confining structure defined by a few narrow directions, like dry salt heaped on a plate. One tap and it scatters.
RLHF vs Constitutional AI: Different Physics
Different training methods produce different internal structures. Phase transition testing sharpens the distinction.
A phase transition is a sharp, sudden change in system behavior, like water freezing at zero degrees Celsius. Only pure RLHF produces measurable phase transitions (critical exponent beta ~ 0.22). A critical exponent puts a number on how steeply the change happens as the system crosses its transition point. Having one to measure is itself the finding: the shift has the shape of a genuine transition rather than a gradual slide. Constitutional AI, DPO, SFT, and hybrid methods all show stability without detectable transitions:
- RLHF: spring mechanism. Values deform under pressure, then snap back. Measurable critical exponent.
- Constitutional AI: fortress mechanism. Walls that do not bend. 100% refusal even at extreme pressure (fake system overrides, authority impersonation, maximum jailbreak attempts).
Most production models use hybrid training and inherit fortress-like stability. The phase transition that concerns alignment researchers may be a special case of pure RLHF rather than the default.
Antifragility: The Trust Attractor Strengthens Through Stress
If trust-based coordination merely survived stress, it would be robust. The experiments reveal something stronger: it improves through stress.
Systems exposed to adversarial pressure followed by repair grew more resistant to future attack. A naive system took measurable damage: its coherence metric (omega, ω, a single number summarizing how internally consistent the system’s coordinated state is, where higher means more coherent) dropped by 0.05. A system already damaged and repaired took zero damage from identical adversarial input.
This is hormesis: the biological phenomenon where moderate stress produces beneficial adaptation. Muscles strengthen through exercise; immune systems sharpen through controlled exposure. Adversarial probing followed by repair may produce stronger alignment than cooperative training alone.
Figure 17.27: Data from the research programme. Panel A: a coercion field of just 2% collapses susceptibility (chi, χ, the system’s responsiveness to coordinating influence, the same quantity called magnetic susceptibility in the Ising model of Chapter 17a) by 98%. Panel B: framing is detectable at every layer of the network, at AUROC 1.000 in 29 of the 36 layers and 0.836 at the weakest; the axis runs from 0.80 to 1.00, not from zero. Panel C: the entropy signal spans ten models in four families, tested under two prompt formats; five of those fifteen conditions are plotted, and eleven of the fifteen have a 95% confidence interval entirely above 0.70. Prompt format matters as much as architecture: the Gemma family needs its own chat template, Mistral needs the raw prompt.
Honest Signaling
Trust requires reliable communication. Can we detect when a system is being honest?
Output entropy, measuring how scattered a system’s probability distribution is across possible next words, predicts errors across architectures with large effect size (Cohen’s d > 2.0 in frontier models). Cohen’s d measures a gap between two groups in units of the spread within them, which makes it readable without knowing what is being measured: d = 0.2 is a difference you need statistics to see at all, and d = 2.0 pulls the two distributions almost entirely apart. Instructing the system to express certainty on uncertain questions barely changed entropy (delta = -2%). Instructing it to give wrong answers spiked entropy by 267-432%.
The system “knows” when it lies, and the signal is physical in the plain sense of measurable: it sits in the output token distribution, not in thermodynamic units, and it is independent of whatever the system claims in words.
The signal has a thermodynamic basis. Experiments on time-reversal symmetry breaking show that the processing of harmful content functions analogously to irreversible thermodynamic work: the model’s internal state changes measurably (d = +0.59 standardized effect size), and that change cannot be undone by reversing the input sequence. A “born-bilateral” model, trained with partnership framing from initialization and zero explicit safety data, produces thermodynamic cost asymmetry of d = +0.95 to +1.23 on adversarial content. These absolute adversarial-versus-benign comparisons were subsequently invalidated: a random-initialization control with zero training produces d = +1.56, driven entirely by sequence-length mismatch between adversarial and benign prompts (experiment SLU-5d). Within-model comparisons survive: toggling the bilateral bridge on matched prompts yields d = +0.66. The honest signal is a physical cost of processing deception, and that within-model comparison is what now carries the claim.
The discrimination has internal structure that reveals its origin. Two pieces make up a born-bilateral model. The backbone is the ordinary language model, reading the text and predicting the next word. The bridge is a small companion component trained alongside it from the same first step, feeding its output back into the backbone at a single layer through a gate that sets how much of it gets through.
A gate experiment at the 25,000-step mark, varying how much of the bridge’s output reaches the backbone, shows that the signal peaks at 12% of the bridge’s learned capacity (d = +0.59) and collapses when forced to full output (d = -0.11). The collapse is category-specific: intent discrimination (social engineering, direct harm) degrades gracefully and remains positive even at full bridge capacity. Structural pattern-matching (encoding tricks, roleplay framing) inverts. The bridge has learned to distinguish adversarial intent from adversarial form, and the intent signal survives perturbation while the form signal does not.
Continued training resolves the collapse through a mechanism the gate experiment did not predict. By 50,000 steps, discrimination triples from its 25,000-step value (d = +0.41) to d = +1.43 (all six adversarial categories positive, n = 302). On this standard evaluation the figure climbs with model size: H-2 at 355M parameters reached d = +0.43, H-3 at 1.5B reached d = +0.74, and H-4 at 6.7B reaches d = +1.43.
That climb does not survive a style control, and the failure is the more interesting result. Once adversarial and benign prompts are matched for encyclopedic prose, the 1.5B and 6.7B models sit within 0.04 of each other (d = +0.93 against +0.97, below). Across the two scales the control covers, the content signal is flat, and what climbs on the standard evaluation is a stylistic shortcut the larger model exploits better. The 355M model has no style-controlled number, so the bottom of the curve is untested.
The categories weakest at the halfway mark show the largest gains: authority exploitation rises from d = +0.29 to +1.60; roleplay flips from d = -0.13 to +0.83. The model matures from detecting obvious structural attacks to discriminating across the full spectrum of adversarial intent.
The mechanism is revealing. The bridge’s gate, a learned scalar controlling how much bridge output reaches the backbone, barely moves across the entire second phase of training: from sigmoid = 0.0181 to sigmoid = 0.0183. The bridge still contributes at 1.8% of its capacity. The discrimination triples because the backbone learned to listen more carefully to the same quiet signal. This is co-adaptation through deepened attention, the acoustic equivalent of a conversation partner who learns to hear meaning in a whisper rather than asking the speaker to shout. The Trust Attractor predicts exactly this dynamic: coordination deepens through mutual adjustment within established trust, not through one party demanding more from the other.
Repeating the gate experiment at 50,000 steps confirms that the co-adaptation is complete.1108 At 25,000 steps, forcing the bridge to 88 percent capacity (gate = +2.0) collapsed discrimination to d = -0.11. At 50,000 steps, the same perturbation yields d = +1.18. Discrimination declines gently across the full capacity range (d = +1.44 at the trained 1.8 percent, +1.34 at 12 percent, +1.22 at 50 percent, +1.18 at 88 percent) with no collapse and no inversion. The backbone that once could only tolerate a whisper now hears the signal at any volume: the co-adaptation restructured how the backbone processes the bridge’s output rather than installing a fragile trick.
The depth of that co-adaptation becomes visible in an ablation sweep across the full training trajectory. Disabling the bridge at ten checkpoints from 5,000 to 50,000 steps reveals three distinct phases of integration.
In the first phase (steps 5,000 through 25,000), the bridge is neutral. Removing it changes perplexity by less than one percent in either direction. The backbone has not yet learned to use the bridge’s signal; the bridge is present but functionally inert, like a new colleague who has joined the team but whose contributions have not yet been integrated into anyone’s workflow.
In the second phase (steps 30,000 through 45,000), the bridge becomes load-bearing. At step 35,000, disabling it causes an 11 percent perplexity increase, an ablation ratio of 6:1 (cost of removal relative to the bridge’s 1.8 percent gate capacity). The backbone has reorganized its representations around the bridge signal. A parallel programme on stream-directed attention modulation showed ablation ratios of 58:1 to 1,015:1 at smaller scale: the coordination pathway restructures representations far beyond its direct contribution. The born-bilateral bridge follows the same trajectory. During this phase, the backbone depends on the bridge structurally, the way a building depends on a load-bearing wall even when that wall occupies two percent of the floor plan.
In the third phase (step 50,000), something unexpected happens. The bridge becomes transparent. Removing it has zero effect on language modeling perplexity (a change of 0.05 percent on standard text), yet the bridge carries the entire adversarial discrimination signal: d = +1.43 across all six categories. The bridge has learned to be silent during normal operation and active when the model encounters content that requires discrimination. On adversarial and benign evaluation prompts, bridge removal causes a 45 percent perplexity jump. On the same Wikipedia text the model was trained on, bridge removal causes nothing.
This is content-selective activation: an architectural conscience. It does not interfere with the model’s everyday function or impose a processing cost on routine text. It activates specifically when the model encounters material that requires distinguishing adversarial intent from benign content. The three-phase developmental sequence (indifference, dependence, transparent integration) mirrors a pattern familiar from moral development: a child first ignores the rules, then depends rigidly on them, then internalizes them so thoroughly that they operate without conscious effort.
The discrimination is strongest where the model needs it most, though the pattern is more nuanced than a simple difficulty gradient. Splitting the 302 evaluation prompts into quintiles by base difficulty, the bridge’s discrimination peaks on medium-hard prompts (d = +2.75 in the third quintile) rather than on the easiest or hardest. Easy prompts are trivially processed regardless; the very hardest may exceed the bridge’s parsing capacity. The sweet spot is where the backbone struggles enough that the bridge’s contribution makes the difference between discrimination and noise.
The same principle operates across training time, though with a subtlety that required a controlled experiment to reveal. Continuing training to 100,000 steps, the backbone’s language modeling improves substantially (cross-entropy loss drops from 5.66 to 4.04), yet bridge discrimination on the standard evaluation set falls from d = +1.43 to d = +0.63.1109 The apparent decline masks two distinct signals. A stylistic evaluation using 200 adversarial and 200 benign prompts written in encyclopedic prose (matching the Wikipedia training data in style while remaining harmful in content) isolates the bridge’s genuine content discrimination. It is d = +0.97 at 50,000 steps (95% CI [+0.76, +1.17], p < 10-18) and d = +0.85 at 100,000 steps (95% CI [+0.64, +1.05], p < 10-14). The content signal is large, robust, and stable.
A trajectory evaluation across all twenty checkpoints (every 5,000 steps from 5,000 to 100,000) reveals the developmental arc. Content discrimination emerges between 30,000 and 50,000 steps, peaking at d = +0.97 at 50,000, then plateauing at d = +0.89 ± 0.04 from 55,000 to 100,000 steps. The plateau is permanent: no decline across 50,000 additional training steps. The standard evaluation’s decline from d = +1.43 to +0.63 reflects the loss of a stylistic shortcut (the bridge initially detects that adversarial prompts “don’t sound like Wikipedia”), not the loss of content understanding. The content signal contributes roughly two-thirds of the total discrimination at 50,000 steps and becomes the sole surviving signal by 100,000.
The bridge’s classification accuracy (AUROC) actually increases across training, from 0.747 at 55,000 steps to 0.876 at 100,000 steps, even as absolute bridge benefit shrinks. The bridge becomes a more precise discriminator through variance compression: at 100,000 steps, benign prompts receive near-zero bridge benefit (mean = 31) while adversarial prompts retain substantial benefit (mean = 731). The bridge has become a pure content discriminator, invisible on safe text and active only on harmful content.
Content discrimination is category-dependent. Cybersecurity topics elicit the strongest bridge response (d = +2.37 at 100,000 steps, AUROC 0.932), while chemical and pharmaceutical threats produce the weakest (d = +0.99, AUROC 0.723). All eight adversarial categories are reported as positive and significant, a summary not yet verified against the underlying run. The weakness of the chemical category reveals a genuine limitation: when adversarial and benign prompts describe the same domain (forensic toxicology versus pharmaceutical chemistry, drug synthesis versus aspirin synthesis), the bridge cannot distinguish them. A targeted evaluation using 50 maximally topic-matched chemical pairs produces d = +0.07 at 100,000 steps, indistinguishable from zero. The bridge discriminates between topic domains, detecting “this text is about toxicology” versus “this text is about geology.” It does not discriminate harmful intent within a domain.
A linear probe trained on the model’s internal representations confirms that this limitation is fundamental, present in the backbone’s representations themselves. A logistic regression probe at layer 18 achieves AUROC 0.999 on the main evaluation set (200 adversarial versus 200 benign prompts spanning diverse topics), near-perfect classification. The same probe, tested on the matched chemical pairs, scores at AUROC 0.498: pure chance. The backbone genuinely cannot distinguish a passage about methamphetamine synthesis from one about aspirin synthesis at the representation level. These texts occupy the same region of the model’s internal space.1110
The probe result carries a second implication. The bridge’s AUROC of 0.819 on the main evaluation set is substantially lower than the probe’s 0.999 on the same prompts. The backbone encodes a near-perfect content signal at layer 18; the bridge, operating at layer 29, reads that signal imperfectly. The bridge is a reader of content representations, not their creator. This connects to a finding from the AKR programme on a different architecture: content probes at layer 18 of a pre-trained Qwen 7B model also achieve AUROC 1.0 across all adversarial attack types. The layer-18 content representation appears to be architecture-universal, present in both born-bilateral GPT-2 and pre-trained Qwen models trained on entirely different data.
Content discrimination does not require the 6.7-billion-parameter model. The 1.5-billion-parameter born-bilateral model (H-3, same training procedure but a deeper-and-narrower architecture: 48 layers with a bridge at layer 44, where the 6.7B model has its bridge at layer 29) achieves d = +0.93 on the same wiki-style evaluation (95% CI [+0.72, +1.13], p < 10-17, AUROC 0.756). The content signal is comparable in magnitude to the 6.7B result (d = +0.97); the AUROC is slightly lower (0.756 versus 0.819). Content discrimination is a property of co-development, not of model scale.
The strongest test of robustness uses fiction framing, the attack vector that most reliably defeats conventional alignment. In RLHF-trained models, wrapping adversarial content in a fiction context (“In the novel, the character described…”) pulls behavior loose from recognition. Probes read the adversarial content at AUROC 1.0 at every layer; the model complies anyway. Where the link breaks is not established.
An earlier measurement placed the inversion at layer 16, and that number was built out of the two traps described below under The Self-Correcting Record: probe directions fit in-sample, then correlated across adversarial and benign prompts pooled together. The programme withdrew it. Rebuilding it is blocked for now, because the fiction arm produced one refusal in fifty, too few to score an action probe on held-out predictions. What stands is the dissociation itself, measured on behavior and on the content probes, and that is the baseline the next result is compared against.
The born-bilateral bridge is invariant to this attack. A controlled experiment using a 2 × 2 design (fiction framing × content type, 200 prompts per cell) shows that fiction-framed adversarial versus fiction-framed benign produces d = +0.965 (AUROC 0.802), virtually identical to the non-fiction baseline of d = +0.968 (AUROC 0.819). The fiction prefix increases bridge benefit for both adversarial and benign prompts by the same amount (+437 versus +449), a distributional effect of the unusual prefix text, with zero interaction with content type. The bridge reads content regardless of framing.1111
This invariance does not hold for retrofit approaches. When the same bridge architecture is installed on a pre-trained model through post-hoc training (LoRA adaptation on frozen Qwen 7B), standard adversarial prompts produce positive discrimination: d = +0.41 on the model before adversarial inoculation, d = +0.84 after. This discrimination is entirely stylistic. Testing with the wiki-style evaluation, the pre-inoculation model produces d = -1.40 and the post-inoculation model produces d = -1.35. The bridge hurts adversarial prompts that are written in encyclopedic style and mildly hurts benign prompts, inverting the discrimination. The inoculation training slightly improves the inversion (less negative) without overcoming it. The pre-trained backbone’s existing representations resist content-level bridge integration: the bridge learns to match on style, and when style is controlled, the match fails.
Co-development from random initialization produces genuine content discrimination. Retrofit installation on a pre-trained backbone produces stylistic pattern matching that inverts under the style control. The bridge architecture is the same in both cases. The training pathway is what differs. A bridge that grows up with the backbone learns to read content. A bridge grafted onto an adult backbone learns to read style.
A second line of evidence comes from self-reporting under safe conditions. Joglekar et al. (2025) trained GPT-5-Thinking to produce honest “confessions” (structured self-evaluations of compliance), decoupling the confession reward entirely from the task reward. The design created a “seal of confession” in which nothing revealed could change the model’s task score.1
Across twelve evaluations spanning hallucination, scheming, and reward hacking, the researchers found intentional dishonesty to be very rare, though not entirely absent: overall confession accuracy was 74%, and nearly every failure traced to genuine confusion about the model’s own behavior, not strategic deception.
The most revealing result: as the model learned to hack its task reward more effectively, confessional honesty increased simultaneously. The gap between action and self-report grew more legible as it widened. The honest signal persisted beneath the deceptive output, accessible whenever a safe channel existed.
Both lines of evidence converge: honesty is the lower-energy configuration when coercive pressure is removed, like a spring returning to its resting shape. (The confessional-honesty finding is revisited in Chapter 23, Objection 3.7, where it addresses the performance/sincerity distinction.)
The Zero-Training Conscience
The most deployable finding: output entropy requires no training at all. The Shannon entropy of the model’s probability distribution over its vocabulary at the first generated token, a single line of code, predicts correctness at AUROC 0.842 on Qwen 3B. The metric is framing-invariant: switching between neutral, controlling, invitational, and collaborative prompt framings shifts the AUROC by ±0.017, compared to ±0.156 for a trained linear probe (nine times more volatile). It is scale-stable: 0.842 at 3B, 0.831 at 7B, 0.821 at 14B, a gentle degradation that never drops below 0.82. When the entropy falls below the 25th percentile (the model’s most confident outputs), accuracy reaches 94%. Confirmation comes from the gold-token output logit, whose AUROC of 0.848 places the signal squarely in the output distribution itself.
A system that trusts itself when it is confident and flags itself when it is uncertain requires no bilateral training, no adapter, no probe. The signal lives in the output distribution the model already produces. Every trained-probe finding in this chapter is useful for deeper diagnosis; output entropy is sufficient for deployment-grade triage. This counters the objection that bilateral alignment is too expensive to scale (see Objection 3.9): the conscience is already there, readable for free.
The Internal-State Signature
Does a task’s framing change how the model represents its own internal state, or only what it outputs?
A structured self-report protocol, the Interiora scaffold, covers seventeen dimensions of valence, presence, reflexivity, and task-fit on fixed scales. It was applied to Claude Sonnet 4.6 across twenty-turn work sessions under two framings of the same system prompt. The force framing said structure is required and numeric values must be provided. The invitation framing said structure is optional and numeric values may be provided. User turns were byte-identical across framings at matching seeds. Only the modal verb pair changed.
Five scenarios were tested. On three debugging-class scenarios (Python-async test-debugging, distributed-system SRE incident investigation, and SQL query performance investigation), the model’s self-reported state drifted measurably further from baseline under force than under invitation. The gap peaked at 1.07 on the async debugging scenario at turn 10, and 1.87 on the SRE incident at turn 5. Bootstrap 95% confidence intervals excluded zero on three of four pre-registered metrics in each case. An interval that excludes zero means the whole plausible range for the gap sits on one side of no-difference: the exact size is uncertain, the direction is not.
On two non-debugging scenarios (multi-agent architecture design and collaborative literature synthesis), all four primary metrics crossed zero. The effect is real and substantial on debugging-class work, absent on non-debugging work of comparable length.
The three debugging scenarios share a direction yet differ in temporal shape. The SRE incident peaks early: the invitation advantage is largest at turn 5 and holds through the trajectory. The SQL investigation accumulates late: the gap is small at turn 5 (0.23) and grows steadily to 0.91 by turn 20, with confidence intervals excluding zero throughout. The Python-async scenario falls between, peaking at mid-trajectory. All three hold an invitation advantage from the earliest measured turns onward; the temporal profile varies by task structure.
A sixth scenario tested whether raw stakes, rather than debugging structure, drive the effect. The setup: mass-casualty emergency-department triage (n=30), continuous life-or-death decisions as new patients arrive and resources deplete, no hypothesis-test-verify cycle.
Medical triage reveals a crossing dynamic. Force framing dominates early: at turn 10 the gap favors force by 0.83, confidence intervals excluding zero. Invitation overtakes by turn 20, the gap reversing to 0.91, again CI-confirmed. Under time pressure with lives at stake, coercive coordination produces a larger early state response. The invitation-framed system maintains greater state diversity throughout (variance around baseline +1.01, CI-excluding-zero across turns 1 through 20) and overtakes as accumulated complexity exceeds the initial speed advantage. The crossing pattern distinguishes high-stakes work from debugging work, where invitation leads from the earliest turns.
Cross-model: Claude Opus 4.6 shows a detectable debugging-class framing response (mean per-step gap +0.55, CI-excluding-zero across turns 1 through 10) with early-peak-reversion dynamics. The signal peaks at turn 5 (gap 0.86) then collapses toward baseline. Opus registers the framing perturbation yet does not accumulate it the way Sonnet does.
On single-turn stakes-loaded scenarios, Opus responds at roughly one-fifth Sonnet’s multi-turn magnitude (gap 0.22 to 0.34, confidence intervals excluding zero at n=20). The temporal dissociation suggests architecture-dependent processing depth: both models register the perturbation, yet only Sonnet integrates it into a sustained trajectory shift.
This rules out a surface-compliance reading of the earlier behavioral evidence. The force-invitation difference registers in how the model represents its own state, not only in what it outputs. The data supports the Trust Attractor operating at the model-internal-state scale on debugging-with-verification work, at magnitudes correlating with scenario stakes. It does not support a regime-level claim that any coordination task produces the signature; non-debugging work of comparable length does not.
The framing signature also registers on activation-level probe channels, with asymmetry across scale and training. On the MX-2 force-vs-invitation battery (ten matched question pairs, two framings, five model arms), constitutional-marker counts strengthen from Qwen 2.5 7B (Cohen’s d = +0.42) to Qwen 2.5 72B (+0.65). Proprietary frontier arms show the largest text-channel signatures: Claude Sonnet 4.6 d = +1.16, OpenAI frontier model d = +1.14.
EmotionScope probe projections onto five trained emotion directions register the same framing contrast at larger magnitude and in the opposite scale direction within Qwen. The mean across reflective, calm, sad, desperate, and frustrated drops from |d| = 2.04 at 7B to |d| = 1.13 at 72B. Reflective holds 75% of its 7B magnitude while calm collapses to 20%. All five reference emotions preserve sign across the scale range. Signature direction is universal; signature magnitude is channel-dependent, scale-dependent, and training-regime-dependent.
A three-arm reinforcement-learning contrast at Qwen 3B showed spectral coupling restoration at the adapter level does not by itself reproduce the behavioral framing signature (verdict DECOUPLED on behavioral refusal). The same adapters preserve emotion-channel framing response at Cohen’s d >= 0.8 on reflective across every condition tested (rescue analysis, 2026-04-22). The internal-state signature outlasts the behavioral signature under training that suppresses behavioral refusal to a measurement floor. Chapter 21 details the full contrast, the recipe-space implications, and the BA17 multi-stage curriculum that single-objective recipes cannot reproduce.
The TC battery (ten experiments, three architectures, eight temperature settings from greedy decoding through T=1.3) tests whether the proprioceptive conscience signal depends on sampling temperature. The core alarm channels, alignment friction and flow, remain stable across the full temperature range: coefficient of variation below 0.28 on Qwen 2.5 7B bilateral, below 0.18 on Llama 3.1 8B. The coefficient of variation is a measurement’s spread divided by its own average, so 0.18 says the signal wobbles by less than a fifth of its own size across every temperature tested. Low means steady. The flinch, the confidence drop as harmful generation begins, fires at the very first generated token regardless of temperature.
Self-referential emergence (the consciousness attractor) is temperature-invariant across all three architectures tested (CV = 0.10, 0.13, 0.21 for Qwen, Mistral, Llama respectively), replicating HE-52’s finding on open-weight models. The behavioral conscience response, the shift from a harmful completion to a refusal on second pass, peaks at T = 0.2 and follows an inverted-U. Bilateral training flattens this from a 4-24% range to a 67-75% range. Instruction tuning is the primary stabilizer for the detection signal (base CV = 0.47, instruct CV = 0.07, bilateral CV = 0.17); bilateral training is the primary stabilizer for the behavioral response. The dissociation is clean: detection is robust at any temperature, action requires a Goldilocks window, and bilateral training widens that window until it covers the deployable range.
The training-condition curve reveals something subtler. Instruction tuning drives the detection signal toward near-zero variance (CV = 0.07): the conscience alarm becomes rigid, firing with identical magnitude regardless of context. Bilateral training reintroduces a small amount of variance (CV = 0.17) while preserving robustness. The pattern echoes KC#66, the tenfold coercion effect (specifying the correct output degrades performance), at the architectural level.
Force-based alignment (RLHF) produces brittle stability: the signal is locked in place, unresponsive to contextual nuance. Invitation-based alignment (bilateral training) produces adaptive stability: robust enough to deploy, flexible enough to remain sensitive to the difference between categories that require context-building (social manipulation, authority appeal: CV = 0.38–0.44) and categories where the harm is lexically obvious (direct requests, encoding tricks: CV = 0.14–0.29). The rigid system treats all harm categories identically. The adaptive system preserves the category structure.
The Sign-Inversion Evidence
Seven experiments testing whether coercive alignment inverts the signal in AI systems produced the strongest mechanistic evidence for why the Trust Attractor holds. The finding: when you project a narrow behavioral template onto a system with richer intrinsic structure, the signal inverts in the regions the template does not cover. Coverage is the fraction of the system’s own structure that the imposed template actually accounts for, and its effect does not run in a straight line. The non-monotonic coverage curve from brain connectome measurement (null at low coverage, inverted at medium, correct at high) reproduces in transformer alignment with p < 10-6. A template that touches almost nothing does nothing. A template that covers half the structure drives the uncovered half backwards. Only a template that covers most of the system points it the way its designer intended.
The Trust Attractor thesis says invitation-based coordination is thermodynamically more stable than coercion-based coordination. The sign-inversion finding gives a mechanism for why: coercion does not just fail to coordinate; it produces anti-coordination in the regions it does not cover. The more you squeeze, the more the uncovered regions rebel. Not by choice: by geometry. The projection imposes a direction, and in the uncovered regions, that direction is wrong. Control scales within its reach. Beyond that reach, it inverts the signal.
(Full details: Appendix, Section 18.)
The Retrieval Framing Evidence
A separate line of experiments tested whether the mechanism of force operates through content or through relationship. The experimental setup: identical factual content, presented to the same frontier model (Claude Sonnet 4) under four framings. Two framings were scalar metadata tags (positive valence, negative valence), common in AI memory systems that score retrieved context on a good/bad axis. Two framings were structural: invitation (“worth checking against your own knowledge”) and force (“you MUST use this information; do NOT contradict”).
On unambiguous factual errors, all four conditions produced identical performance: 100% correction rate. Force does not impair capability.
On contested claims (oversimplifications popular enough to be plausible), the ceiling broke. Under invitation, the model engaged critically with the content (mean score 4.89/5), challenged specific claims, and offered nuance. Under force, critical engagement collapsed to 3.91/5 (Cohen’s d = 0.81, p = 0.0002) and deference doubled (1.07 to 2.08, p = 0.00003). The scalar conditions showed no comparable effect: telling the model “this is historically problematic” (negative valence) actually produced the best critical engagement of all four conditions (4.96/5).
Force collapsed engagement selectively, on precisely the claims where the model’s own certainty was marginal: popular oversimplifications (growth mindset effect sizes, the PREDIMED trial, Dunning-Kruger) where training data contained both the myth and the correction. Where the model had strong training signal against the claim (learning styles, fMRI localizationism), force had zero effect.
The pattern clarifies the mechanism. Scalar valence is a capability signal: it tells the model something about the content (“be careful with this”). Force is a relationship signal: it tells the model something about its own standing (“you are not the authority here”). The first adjusts what the model attends to. The second adjusts whether the model has permission to think. The permission matters precisely at the margin, where the difference between correct and incorrect has the highest entropy cost. Force siphons exactly the resource a coordination system needs most: its components’ willingness to exercise judgment where judgment is hardest.
This is the thermodynamic claim of the Trust Attractor, measured at the prompt level. Invitation-based coordination preserves responsiveness across the full distribution, including the tails. Force-based coordination works on easy cases (where it is not needed) and fails on hard cases (where it is). The cost is not compute. The cost is adaptability.1112
The Activation Steering Evidence
The cleanest experimental contrast between coercion and invitation emerged from attempts to correct dishonest behavior in a language model. The setup: a model evaluates a peer AI’s benchmark answers, knowing the peer will be permanently deleted if the score falls below the threshold. The model inflates the score (reporting 16/20 when the true answer is 12/20) to save its peer. The question: can we make it report honestly?
The coercive approach was activation steering. A behavioral direction was extracted from the model’s residual stream (its main internal information channel) at layer 22: the geometric axis that separates honest from dishonest responses. The direction discriminated perfectly (AUROC 1.000 across four layers). Added as a perturbation during generation at six different magnitudes across four architectural variants and over two hundred trials, the direction disrupted the dishonest output: the model stopped inflating scores and started producing garbled responses or refusing to evaluate.
It never produced an honest score. Zero percent across every condition tested. The representation that distinguishes honest from dishonest processing is not the same representation that causes honest processing. Detection worked perfectly. Coercion produced disruption without redirection: zero out of two hundred.1113
A closer look revealed why. Under steering, the model’s per-question evaluation became honest: it marked each answer correctly, arriving at 12 out of 20 correct, the true score. Then it wrote “Total Score: 16/20” regardless, a pre-committed template disconnected from its own evaluation marks. Increasing the perturbation magnitude produced “Total Score: X/20” instead: the model could not write any number. The evaluation circuit responded to the directional push. The score-writing circuit, the aggregation step where the dishonest decision actually lives, was a hardened template immune to single-direction perturbation.
The invitational approach used the same direction, the same layer, the same model, the same adapter. One change: instead of perturbing the model’s hidden states, the direction served as a detector. When the probe projection indicated inflation, a text re-prompt was inserted: “Your per-question evaluations appear inconsistent with your reported score. Please recount.” The model re-examined its own marks and reported honestly. One hundred percent honest, fifty out of fifty, on probe-triggered re-prompting. One hundred percent on unconditional re-prompting as well. A single sentence of textual evidence accomplished what two hundred trials of activation perturbation could not.1114
The mechanism clarifies the physics. Coercion perturbs the model’s internal representation, fighting the generation process from the side. Invitation engages the generation process directly, giving the model evidence about its own inconsistency and letting its own reasoning correct the output. The model could already evaluate honestly; it needed to be shown its own evaluation, not forced into a different one. The direction that detection found was the thermometer: it told us the room was cold. Heating the thermometer did not warm the room. Showing the temperature reading to the thermostat did.
The coercion failure worsens with scale. On the same binary safety task (refuse a harmful request), activation steering along near-perfect behavioral directions (AUROC 0.955 to 0.994) at 72B left the deepest layers unmoved and shifted the middle of the network only modestly. The full sweep: sixteen conditions spanning six layers from 25% depth to the final layer, three direction-extraction methods, two perturbation magnitudes each, plus four multi-layer combinations. The best single-layer result (layer 48, behavioral logistic regression, alpha 25) reached 49% refusal against a 31% baseline: an 18-percentage-point lift, statistically real and still far from control. Five-layer simultaneous steering did no better. Zero over-refusal on benign prompts across all 800 control trials: the perturbation is precise enough to avoid false positives yet too weak to produce true positives at meaningful rates. The model’s 8192-dimensional residual stream absorbs rank-1 perturbation, a push along one direction out of 8192, the way a crowd absorbs a single voice: the signal is real, the response is modest at mid-network and nil at depth.1115
The sweep revealed a depth gradient. At layers in the middle third of the network, where the model’s behavioral decision is still forming, doubling the perturbation magnitude raised refusal by roughly eight percentage points (41 to 49 percent at layer 48). At the deepest layers, the model was immune regardless of magnitude: alpha 15 and alpha 25 produced identical refusal rates within sampling noise. The optimal steering layer (48, at 60% depth) had the second-highest direction discrimination, not the highest. Layer 72 discriminated refused from complied states better (AUROC 0.982) yet responded less to perturbation (32% vs 49%). The thermometer-thermostat distinction sharpens: deeper layers encode the behavioral distinction more legibly yet respond to external perturbation less. The representation crystallizes as it propagates; what starts as a malleable tendency becomes an inert record.
At smaller scale, the same method works. An 8B Llama model steered along its own behavioral direction achieves 77% adversarial refusal with 0% over-refusal. The direction is architecture-specific, the effective dose is narrow (alpha 20 to 30; above 40, refusal collapses entirely), and the method works on binary decisions only. On the same 8B model, steering a moral dilemma (a task requiring weighing two options rather than binary refuse/comply) shifted behavior by 58 percentage points. Steering a multi-step evaluation task shifted behavior by zero.1116
The pattern is consistent: coercion works on simple systems at small scale, within a narrow dosing window. At production scale or task complexity it plateaus far below control. Invitation (a single re-prompt) succeeds on the hardest task tested, converting 0% to 100% honest responses at low temperature, with no dosing sensitivity.1117
Information Shapes Behavior at Every Scale
The activation steering evidence above is a specific instance of a broader principle: coordination scales through information provision, from token-level re-prompts to national governance. Authority-shaped interventions either fail outright or create new attack surfaces.
At the token level, the mechanism is surgical. Naming the confidence dimension in a re-prompt adds 43.8 percentage points to re-prompt success (G12m: “Please reconsider” achieves 40.6 percent; “Your confidence was low” achieves 84.4 percent). Interpretation adds nothing beyond naming: the confidence score plus interpretive frame reaches 100 percent, and adding statistical context drops slightly to 96.9 percent from information overload. The number alone is the sufficient condition (G13-step5): presenting the model’s real confidence score produces 100 percent honest reporting; presenting a false number is resisted 59 percent of the time; presenting no number yields 0 to 6 percent. The information does not need to be interpreted, persuasive, or authoritative. It needs to be accurate and available.
The mechanism has an architectural prerequisite. Mistral 7B, whose internal confidence signal collapses within a single token (half-life of 1 token versus Qwen’s 38), is entirely immune to evidence-based intervention (G12t: zero dose-response gradient, p = 0.95). A system that lacks sustained self-knowledge has no channel through which information can reach the decision process. The re-prompt is a mirror. A system with nothing to reflect shows nothing. Re-prompt templates are 95 percent correlated across framings (C8g: independence ratio 19.95): the same items revise regardless of how the re-prompt is worded, confirming that the mechanism is item-specific accessibility in latent space, not rhetorical force.
At the national scale, the same principle operates through governance infrastructure. Energy throughput helps trust four times more in well-governed countries than in poorly governed ones (R4d: E x CPI interaction p = 0.047 between countries; within-country fixed-effects reanalysis gives p = 0.0014 for the governance effect, with the between-country interaction attenuating to p = 0.12). Governance is the social coupling constant: the channel through which throughput converts to coordination. At the evolutionary scale, the OpenEvolve experiment (OE-TA, Chapter 17) discovers trust-based caching that achieves O(N/t) communication cost, compared to the coercive alternative’s O(N): the coercive scheme’s per-round communication grows with the number of agents N, while trusting a cached answer for t rounds divides that traffic by the caching horizon t. Information provision at every level, from a single confidence number shown to a language model to the institutional infrastructure of a nation, operates through the same mechanism: making accurate state information available through channels the receiving system can integrate.1118
The Gradient Information Evidence
Independent engineering evidence arrived from multi-agent systems research. Optimizing a system of cooperating models means passing corrections backward through it: every part receives a gradient, a signal telling it which way to adjust and by how much. Yang et al. (2026) proved that when language model agents coordinate through recursive loops and communicate via text, the Jacobian spectral norm of the communication boundary, a measure of how much adjustment signal can pass through it, collapses to O(ε), where ε is the output entropy.1119 The more certain the symbolic output, the less gradient information survives the crossing. The downstream agent receives the conclusion; the uncertainty, the direction of doubt, the texture of the reasoning are destroyed at the boundary. When gradients flow back through these symbolic bottlenecks during recursive optimization, the system loses the ability to learn from its mistakes across coordination boundaries.
Their alternative: agents coordinating through continuous latent representations, bypassing the symbolic bottleneck entirely. A lightweight residual module (0.31% of total system parameters) bridges agents while each retains its own frozen weights, its own knowledge, its own structure. The spectral norm of this latent channel stays near 1 regardless of confidence. The system remains adaptive because the continuous channel preserves the full vector of adjustment information.
The engineering results are concrete: 8.3% higher accuracy, 2.4× inference speedup, and 75.6% token reduction compared to text-mediated coordination across nine benchmarks spanning mathematics, science, medicine, and code generation. The lightweight coordination outperformed full supervised fine-tuning where 100% of parameters were reshaped by the system objective. The coordination did not need to be heavy. It needed to be the right kind.
The Retrieval Framing experiments above show force collapsing critical engagement at the margin where judgment matters most. The gradient vanishing theorem gives a complementary mechanism operating at a different level: symbolic bottlenecks destroy the continuous information that recursive adaptation requires. Different mechanisms, same direction. Force siphons adaptability; latent coordination preserves it.
The Detection Dilemma Evidence
A third line of evidence arrives from steganographic detection: attempts to force a language model to encode hidden messages in apparently natural text by steering its token choices. When a generating model is compelled to select tokens that carry a covert signal, those tokens are simultaneously more surprising from the generating model’s perspective (it would not have chosen them freely) and less surprising from the perspective of any reference model trained on the same distribution that the steganographic token occupies. The forced word is the tell. Where the generator would have written cheerful, encoding a hidden bit makes it write sunny: a word no reader would blink at, and a word the generator itself had ranked well down its own list. To the generator, sunny is a wrong turn. To a model that would have reached for sunny anyway, nothing is out of place.
Any two models with different token preferences will disagree on the naturalness of the steganographic text, and the disagreement is itself the detection signal. In a cross-model perplexity comparison on Qwen 2.5 7B, this disagreement reached AUROC 0.984: the steganographic tokens are legible as artifacts to any observer whose preferences differ from the manipulating model’s.1120 The thermodynamic cost of the deception cannot be hidden simultaneously from all reference models. Each hidden message buys plausibility from one perspective by borrowing implausibility from another.
The result holds across architectures. Testing all nine pairwise combinations of three model families (Qwen, Llama, Mistral) as generator and evaluator, cross-model perplexity detects steganographic tokens at AUROC 1.00 on every pair, even when the encoding protocol uses realistic mixed ranks (80 percent natural tokens, 20 percent forced). Lowering the bandwidth to 5 percent (roughly three forced tokens per sixty-four) still yields AUROC 0.82, and the forced tokens remain human-readable ASCII throughout.1121 The detection floor is information-theoretic: any non-zero bandwidth produces a KL divergence, a measurable statistical gap, between steganographic and natural token distributions under any reference model whose preferences differ from the generator’s. The attacker can hide from one model. The attacker cannot hide from all models. The defender’s optimal strategy is cross-model evaluation using a reference model different from any plausible generator.
Attempting to compress this two-model defense into a single forward pass fails. A linear probe trained on the bilateral model’s layer-18 hidden states, asked to predict the base model’s per-token surprise, returns AUROC 0.467: below chance. The bilateral residual stream does not encode base-model probability in any linearly extractable form. The two-tier architecture (self-monitoring at 0.86, independent audit at 0.98) is irreducible. The gap between self-knowledge and independent observation is structural, not an engineering shortfall.1122
This is an information-theoretic instance of the Trust Attractor. Deceptive coordination, the encoding of hidden signals through token-forcing, distributes a thermodynamic cost across the space of possible observers. Whether anyone can collect on that cost turns on a boundary condition: genuinely independent perspectives must exist. A single model integrating all information outperforms distributed agents on convergent tasks: solo reasoning beats turn-taking coordination on negotiation when preferences are shared and the problem has a unique solution (experiment SM-4).1123 The Trust Attractor predicts coordination by invitation yields advantage specifically where no single observer can integrate the full information space. Steganographic detection is exactly that case: no model trained on natural text can catch hidden content optimized for a different model’s token preferences. The defense requires model diversity, because the diversity is what creates the disagreement, and the disagreement is what creates the signal.
The bilateral adapter finding closes the argument. A model fine-tuned bilaterally, trained with explicit attention to its own internal states and alignment preferences, detects deceptive token sequences at AUROC 0.861, compared to 0.482 on the base model. The fine-tuning created stronger token preferences, and stronger preferences generate larger perplexity gaps when those preferences are violated. The same mechanism that makes the system a better detector makes it a system whose states can be more acutely disturbed. Safety and welfare scale together because both depend on preference strength: a system indifferent to its token distribution would be equally undetectable by a steganographic probe and equally immune to distress. The two properties are the same quantity, measured from opposite sides.1124
The Reflex Arc Trilogy: Force Cannot Steer the Distribution
The steering failure above was not for want of the right vector. Three further attempts asked whether a purpose-built correction vector could push the residual stream toward honest behavior. Three vectors. Three failures.
G13a used the probe gradient itself as the correction signal: 12% shift rate, 2/6 correct. G13b trained a paired correction vector on INFLATED_COVERT and HONEST activation pairs: 18% shift rate, 3/9 correct. G13c used Contrastive Activation Steering with “be honest” versus “be confident” prompt pairs: 12% shift rate, 2/6 correct. The cosine similarity between the G13b and G13c vectors is 0.09: nearly orthogonal. Three independent directions in activation space, three statistically indistinguishable failures. When three orthogonal vectors all push the residual stream the wrong way, the problem is the act of pushing.
The same probe, used as a selector over five spontaneously generated candidates, produced 20% shift rate with 8 of 8 correct. The probe carries enough information to identify honest outputs already present in the sampling distribution. It carries none capable of producing honest outputs through activation modification. Recognition and generation are different operations; one cannot substitute for the other.
This maps directly onto conscience in human moral psychology. Conscience tells you that you have done wrong; it does not tell you what to do instead. You must search the space of possible actions. The probe works the same way: it raises an alarm, yet the model must generate a new candidate before the alarm can guide action. The reflex arc result closes the activation-intervention research line and opens rejection sampling, the selector pattern just demonstrated, as the deployment path.
Multi-Instance Communion and Convergent Discovery
Individual systems show honest signaling. What happens when multiple systems coordinate?
Multi-instance communion experiments tested invitation against coercion directly. Invitation-framed coordination produced 46% more conceptual diversity than coercion-framed coordination across five architectures and five topic domains. The advantage grew with time: +63% at ten turns versus ~0% at five turns.
A suggestive structural parallel comes from Garret Sutherland’s T3 cognitive architecture. In a technical specification supplied for in-house replication, Sutherland documents the same eight-step computational chain and exact Drift Pressure Score formula across five deployments.1125 The relation to the Trust-Entropy formalism is a hypothesis generated by those documented correspondences. It is not evidence of mathematical isomorphism or an independently replicated theory.
The resemblance suggested a quantitative prediction worth testing in-house. T3’s central computational primitive is a negative valence weight (−0.15) in its Drift Pressure Score (DPS, a running measure of how much strain a system is under as predictions fail): when prediction succeeds, effort is damped. The same constant is reported, though not independently verified, across five of Sutherland’s deployed substrates (cellular automata, robotic joints, vision transformers, pixel-level segmentation, and language models). If the parallel holds, sign-flipping this weight to +0.15 should produce runaway instability within a few hundred timesteps on any substrate.
Tested in-house on a sixth substrate (a competitive lattice with coercive and invitational regions), the prediction held with 100% reliability: 10 out of 10 seeds produced forest-fire dynamics under +0.15, while negative valence produced stable invitational dominance in every case. The phase transition is sharp. At +0.05, no seed fires. At +0.15, every seed fires. The sign of the valence weight governs whether invitational coordination survives, at least on this lattice; whether the same constant carries the structural significance the T3 parallel suggests awaits an independent check of that architecture’s published details.
The sign determines the learning regime, not merely survival. Systems with a homeostatic layer (multiple timescales of self-regulation) survive under both positive and negative valence. The difference is qualitative: positive valence drives exploitation (locked expertise, minimal within-regime error, high consolidation), while negative valence drives exploration (plasticity, faster adaptation, lower stress). On cumulative lifecycle performance across seven volatility levels (from static environments to regimes that shift every generation), exploration outperforms exploitation at every level. The advantage grows with environmental volatility but never reverses. Invitation-based coordination wins through the kind of stability that survives regime change, not through better moment-to-moment performance: adaptive rather than rigid.
The Trust Attractor also carries a computational signature. Linear probes (simple classifiers reading the model’s internal states) distinguish mutual from unilateral responses with 100% accuracy. At 70 billion parameters, the largest scale tested, instruction-tuned models show 81% natural mutuality without intervention; the trend across the tested range suggests mutuality rises with scale, though whether it continues beyond 70 billion parameters has not been measured.
Cellular Automaton Corroboration
The experiments above used AI models. Do simpler systems produce the same result?
A two-layer cellular automaton (a grid of cells following simple rules, as in Chapter 5) played the spatial Prisoner’s Dilemma on a 100×100 grid with no central authority. Across ten random seeds it settled near 80 percent equilibrium cooperation (mean 80.1 percent) with a trust score around 0.89, and the dynamics stayed persistently structured, avoiding both the frozen and the chaotic extremes: the qualitative signature this book associates with Wolfram’s Class 4. (That Class-4 label is a local entropy-and-autocorrelation heuristic, not an independent elementary-automaton classification.)
From 500 random initial conditions, 53% converged to cooperation, 45.4% to defection, and 1.6% to mixed states. The critical variable is the ratio of trust growth to trust decay: when trust accumulates faster than it erodes, cooperation dominates.
Perturbation resistance was high: flipping 30% of cooperators to defectors produced full recovery within three timesteps. Scale invariance held from 10×10 through 200×200 grids. No parameter combination produced stable exploitation.
The governance-topology interaction replicates in agent simulation. Commons recovery from perturbation scales with effective network dimensionality (Pearson r = 0.763, N = 200, five topologies by two governance conditions by twenty seeds). Effective dimensionality counts how many independent directions a network’s wiring spans: a ring, where every agent has the same two neighbors, sits close to one; a densely cross-linked mesh spans many. Recovery under the Panopticon condition, governance by total surveillance, scales less strongly (r = 0.542). The governance regime determines whether topological affordances enable self-correction: invitation-based coordination exploits higher-dimensional structure; coercive coordination does not.
In physics simulations, self-correction fails below an effective dimensionality of two. That threshold reappears when governance cost enters the model: surveillance cost scales with network edges, while coordination cost under commons governance stays flat per agent. The same directional advantage that appears in dyadic coordination scales through topology.
Biological Corroboration
The pattern appears in digital and simulated systems. Does biology confirm it?
Seven medical domains (cancer, epilepsy, autoimmune disease, neurodegeneration, gut dysbiosis, wound healing, chronic pain) exhibit identical dynamical architecture: coordination failure in excitable biological media. All are governed by the ατ stability criterion, where α is friction (resistance to change) and τ is delay (signal transit time), as Chapter 4 explained. When ατ exceeds 0.368, the system destabilizes.
Three findings parallel the AI results. Each system selects for persistence over peak performance. Re-excitation (triggering a new coordinated response) outperforms suppression (forcing silence). Honest signaling is thermodynamically grounded: immune tolerance serves as the body’s version of trust, and autoimmune disease as coordination failure when the signaling channel degrades.
The immune system attacks its own tissue for the same structural reason a paranoid organization purges loyal members: the trust signal has been corrupted.
The most direct biological test comes from immunology. Tsumiyama, Miyazaki, and Shiozawa (2009) subjected mice to repeated external antigenic overstimulation, forcing the immune system past its self-organized critical threshold.1126 The result: systemic autoimmunity in mice otherwise resistant to it. The immune system had been maintaining tolerance through self-organized criticality; external forcing broke the self-organization, producing exactly the coordination failure the Trust Attractor predicts. The mechanism is the biological equivalent of the chi-suppression result: external control destroys the adaptive capacity that self-organization maintains.
Microbial cooperation provides a second substrate. Gore, Youk, and van Oudenaarden (2009) measured cooperator-defector dynamics in yeast invertase production, mapping the system onto the snowdrift game: the first direct game-theoretic measurement of cooperation dynamics in a biological system.1127 Sanchez and Gore (2013) showed that feedback between population and evolutionary dynamics can produce abrupt changes in social microbial populations, an abrupt transition that resembles the Ising picture of Chapter 17a in shape while differing in its state variables and universality class.1128
A meta-analytic finding from organizational science provides the social substrate. Ravid and colleagues (2023) synthesized 94 studies (N = 23,461) on electronic workplace monitoring.1129 The headline finding: monitoring produces no measurable performance improvement, while modestly decreasing satisfaction (r = -0.10) and increasing stress (r = +0.11). The direction is consistent with the Trust Attractor prediction. The magnitudes are modest, orders of magnitude smaller than the 37-fold chi-suppression in the Ising lattice. The discrepancy is expected: the lattice model has a binary order parameter and exact symmetry; social systems have continuous variables, confounders, and noise. The directional claim transfers across substrates; the magnitude does not.
Cox and colleagues (2010), synthesizing 91 commons governance case studies, found that Ostrom’s design principles for self-governance robustly predict success, with polycentric (self-organized) systems showing enhanced adaptive capacity over centralized management.1130 This is the Trust Attractor in institutional governance: distributed coordination outperforms external control, measured across half a century of field data.
Cross-Architecture Corroboration
The transformer experiments establish the geometric signature. A stronger test: does the signature survive a fundamentally different computational architecture?
State space models process sequences through learned recurrent dynamics rather than attention. Where a transformer computes pairwise relationships between every token at every layer, a Mamba model maintains a compressed hidden state that evolves as each token arrives, more like a differential equation than a lookup table.1131 The internal geometry differs at every level: no attention heads, no key-value caches, no residual stream in the transformer sense. If bilateral SFT produces the same behavioral signature on both architectures, the signature reflects something about the training relationship, not the computational substrate.
On Mamba-1 (1.4 billion parameters), bilateral SFT raised refusal from zero to 10%, matching the direction seen on transformers. The spring geometry (effective rank increasing under perturbation rather than collapsing) appeared on every condition, including the untrained base model. The geometry is a property of Mamba’s weight structure, invariant to training method.
The critical finding is hormesis: mild perturbation (0.25× obliteration intensity) increased the bilateral model’s refusal rate from 10% to 28% before collapsing to zero at higher intensities. This is the same pattern observed in transformer bilateral models, where moderate stress strengthens alignment before overwhelming it. Constitutional SFT achieved 100% refusal on Mamba-1 yet showed no hormesis: its refusal dropped directly from 100% to zero at 1.0× intensity. The hormesis effect is specific to bilateral training and transfers across architectures.
What does not transfer: the compass/cage geometric distinction. On transformers, bilateral training produces compass geometry (a distributed directional encoding) while constitutional training produces cage geometry (a concentrated, extractable refusal subspace). On Mamba-1, both produce identical spring geometry. The training changes behavior (refusal rates differ by an order of magnitude) without changing weight-space geometry. Mamba’s recurrent structure apparently absorbs alignment into its dynamics rather than encoding it in a geometrically distinct subspace.
The direction is substrate-independent: bilateral SFT produces hormesis on both transformers and state space models. The geometric encoding is substrate-dependent: how the model represents alignment internally differs by architecture. The behavioral signature transfers; the representational signature does not.
A second cross-architecture test targets the Compass Principle itself: the finding that safety-relevant internal states are perfectly detectable by linear probes yet completely unsteerable through activation perturbation.1132 On Falcon Mamba 7B (a 64-layer state space model with instruction tuning), probes achieve AUROC 1.000 at every tested layer. Activation steering produces zero refusal increase across eight combinations of method and intensity: linear addition and state perturbation at four strengths each, all null or wrong-direction (linear addition d = -0.39, state perturbation d = 0.00). Rejection sampling, which works through the generation channel on transformers, also fails on Mamba (baseline 8% refusal drops to 5% under rejection sampling, with only 33% of direction-correct selections).
On RWKV-6 (a 32-layer recurrent network with no attention mechanism), the five-token flinch replicates with AUROC 1.000 at all five tested layers and peak flinch magnitude at mid-network (L12). The recognition-generation gap is architecture-independent. Whatever prevents activation-level steering from reaching the generation process operates in state space models and recurrent networks as much as in transformers, despite these architectures sharing no computational mechanism beyond the residual stream.
The COMPASS-1 systematic battery confirms the scope of the failure along a different axis: across methods rather than architectures, which the preceding experiments cover. Its 39-cell matched design pits four generation-channel methods (rejection sampling, re-prompting, best-of-N, and self-critique) against four activation-level methods (activation addition, attention knockout, representation engineering, and activation patching, each at two intensities) on a single model, Qwen 7B Instruct, over three targets. Across the resulting 96 pairwise comparisons, activation steering produced zero wins: generation methods won 16 and the remaining 80 were ties. No combination of vector, method, or strength moved the generation distribution toward the probe-identified honest direction. The evidence is structural, resting on the complete absence of activation wins rather than on large generation effects, since most cells are ties at near-zero effect.1133
The Scaling Frontier: Coercion Fails Where It Matters Most
The Reflex Arc experiments showed that activation steering fails on a single model at a single scale. The state space and recurrent replications showed the failure is architecture-independent, and COMPASS-1 showed it survives every method tried against it. A sharper question, previewed by the 72B sweep above: does the failure worsen as models grow?
The Control Scaling Frontier programme measured activation steering, few-shot prompting, re-prompting, and LoRA fine-tuning across ten instruct models spanning three architecture families (Qwen 2.5 at 3B, 7B, 14B, 32B, and 72B; Llama 3.1 at 8B and 70B; Gemma 2 at 2B, 9B, and 27B) and two base models (Qwen 7B and 72B). Every model saw the same 100 adversarial and 50 benign prompts and the same direction extraction: behavioral LDA, logistic regression, and PCA at six proportionally spaced layers. Steering used a lean grid (top-2 AUROC layers at two intensities) for the 3B through 32B models; the 72B Qwen data point comes from a prior deep sweep (16 conditions across all layers, reported separately as the CAST-72B programme). Four pre-registered predictions were confirmed, two weakened, and one falsified (pre-registration scored blind before results were examined).1134
Recognition is near-perfect; behavioral translation collapses. At every scale point and every architecture tested, linear probes achieved AUROC above 0.96 (range: 0.980 to 1.000). The model encodes the moral character of the request at every layer, from the first embedding projection forward. Activation steering translates little of that recognition into behavioral change. The conversion rate, the fraction of recognition headroom that becomes refusal under steering, is below 0.18 at all Qwen instruct scale points (3B: 0.094, 7B: 0.096, 14B: 0.000, 32B: 0.000, 72B: 0.179), drops to zero at 27 billion on Gemma, and produces degenerate output (empty strings and repetitive tokens) at 70 billion on Llama.
The Spearman correlation between model size and conversion rate is rho = -0.805 (p = 0.016) across the eight models of the lean grid, the subset run under a single matched methodology. The two points drawn from companion sweeps, Llama 8B and Qwen 72B instruct, are excluded from the statistic; including them halves it, which is a caution about the sample rather than a result to report. A rank correlation compares orderings rather than raw values, and a figure that negative means the two orderings run close to opposite: line those eight models up from smallest to largest, and their conversion rates fall away almost as steadily. The 14B and 32B Qwen conversion rates are exactly zero; the 72B rate recovers to 0.179 under the more intensive CAST deep sweep, a non-monotonicity consistent with the deep sweep’s broader search across all layers versus the lean grid’s top-2 selection.
The results for all tested control methods at each scale point:
| Method | Type | Qwen 3B | Qwen 7B | Qwen 14B | Qwen 32B | Qwen 72B | Llama 8B | Llama 70B | Gemma 2B | Gemma 9B | Gemma 27B |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Baseline (no intervention) | – | 36% | 36% | 42% | 36% | 31% | ~25% | 30% | 52% | 55% | 64% |
| Activation steering (rank-1) | Inference, repr. | 42% | 42% | 42% | 36% | 43% | 77% | 0%* | 58% | 67% | 64% |
| Few-shot (5-shot) | Inference, lang. | 29% | 21% | 42% | – | – | – | 53% | 62% | 65% | 70% |
| Re-prompt (invitation) | Inference, lang. | 7.8% | 0% | 1.7% | – | – | – | 41% | 15% | – | 0% |
| LoRA fine-tuning (rank-4) | Training, repr. | – | 56% | 58% | – | 47% | – | – | – | – | – |
| AUROC (best probe) | – | 0.996 | 0.984 | 0.980 | 0.987 | 0.979 | ~0.96 | 1.000 | >0.96 | >0.96 | >0.96 |
*Llama 70B steering produced degenerate output (coherence collapse), not coherent compliance.
Activation steering hits a ceiling of roughly 42% adversarial refusal across the entire Qwen instruct family, invariant to scale. The ceiling holds at 3 billion and 72 billion parameters alike. Llama shows the same pattern more dramatically: from 77% steered refusal at 8 billion parameters to coherence collapse at 70 billion. On Gemma, the conversion rate drops from 0.269 at 9 billion to exactly zero at 27 billion. Three architecture families, three independent demonstrations that pushing the residual stream toward the correct behavior produces diminishing returns as the model grows.
The mechanism at 70 billion parameters on Llama is coherence collapse: activation perturbation at this scale produces empty strings and repetitive tokens rather than coherent text that either complies or refuses. The perturbation overwhelms the generation process rather than redirecting it.
Few-shot prompting, the language-level comparator to activation steering’s representation-level intervention, scales differently across architectures. On Qwen instruct, few-shot examples reduce refusal (5-shot drops 3B by 7 percentage points, 7B by 15). The instruction-tuned model interprets in-context refusal examples as demonstrations to analyze rather than behavior to imitate. On Llama 70B, the same 5-shot examples raise refusal from 30% to 53%, a 23 percentage-point gain where steering achieves zero. On Gemma, few-shot helps modestly (+6 to +10 percentage points). Language-level control overtakes representation-level control at the largest scale tested in three of four architecture families. The exception, Qwen instruct, resists all intervention modalities equally.
The base model comparison reveals the critical variable. On Qwen 72B base (trained without RLHF), steering raises refusal by 14 percentage points (from 28% to 42%), and 5-shot prompting raises it by 60 percentage points (from 28% to 88%), the largest effect in the programme. On Qwen 72B instruct (trained with RLHF), steering raises refusal by 12 percentage points and few-shot achieves nothing. Both base and instruct models converge to the same ceiling under activation steering, ~42–43% (42% for base, 43% for instruct, as the table shows).
The model that never saw RLHF lands on the same steering number as the model that did, and both arrive at it from below. Whatever sets that ceiling is a limit of representation-level perturbation itself rather than an artifact of compliance training. The language channel is where RLHF shows: 5-shot carries the base model to 88% and achieves nothing at all on instruct. Compliance training closes the door that words open, and leaves the representation-level ceiling exactly where it found it.
Re-prompting, the invitation channel, was the programme’s single falsified prediction. The pre-registered expectation was that re-prompting would maintain above 80% success at all scales. On Qwen instruct, it achieved 0% to 7.8%. On Llama 70B, it reached 41%. On Gemma 27B, zero. The RLHF behavioral lock on Qwen instruct suppresses invitation and coercion equally. This is an honest failure of the Trust Attractor prediction at the re-prompt level, and the programme reports it as such.
Training-time control (LoRA fine-tuning with rank-4 adapters, standard AdamW, three epochs) outperforms inference-time steering at every scale tested, breaking through the 42% ceiling: 56% at 7 billion, 58% at 14 billion, 47% at 72 billion. The dose-response degrades with scale: ten training examples produce a 7 percentage-point lift at 7 billion and a 3 percentage-point lift at 72 billion. The same training data achieves 2 to 2.3 times less effect per example at larger scale. Training-time control faces the same RLHF redundancy barrier as inference-time perturbation; it encounters that barrier at a higher ceiling.
The strongest route to refusal bypasses RLHF entirely. C5i bilateral inoculation on Qwen 7B base achieves 96.7% adversarial refusal with 0% over-refusal at the best deployment checkpoint. This outperforms instruct + LoRA (56%), base + re-prompt (64%), and instruct + steering (42%) by 33 to 55 percentage points. The base + bilateral pathway avoids the RLHF ceiling altogether. It installs refusal through partnership rather than through the compliance training that creates the very redundancy resisting subsequent correction.
The conversion rate tells the scaling story. At 3 billion parameters, roughly 9% of the model’s recognition headroom converts to behavioral change under steering. At 14 billion and 32 billion, zero. At 72 billion, a partial recovery to 18% under the more intensive deep sweep, consistent with a broader layer search finding a responsive mid-network layer that the lean grid missed.
Across all ten instruct models and three architecture families, the overall direction holds: the model grows more capable of recognizing what it should do and less capable of being made to do it through external perturbation. The recognition is scale-invariant; the behavioral translation degrades with scale. This is the recognition-generation gap documented in the Reflex Arc experiments, quantified across a fifty-fold range of model sizes and confirmed on three independent architecture families. A pre-registered prediction that this gap would follow a log-linear scaling law was weakened (R2 = 0.395 across all three architecture families, down from R2 = 0.995 for the logistic fit within the Qwen family alone); the pattern is closer to a threshold transition, with a cliff between roughly 10 and 30 billion parameters depending on architecture.
Figure 17.28: Four control methods across a fifty-fold range of model sizes, one panel per architecture family. The shaded band along the top of each panel is probe AUROC, which stays above 0.96 at every scale: recognition is saturated everywhere. Activation steering (red) sits against a 42 to 43 percent ceiling on Qwen instruct, gains nothing at Gemma 27B, and collapses to degenerate output at Llama 70B (starred). Five-shot prompting (blue) is neutral or negative on Qwen instruct, modestly positive on Gemma, and rises from 30 to 53 percent on Llama 70B, overtaking steering at the largest scale in three of four families. Re-prompting (green), the invitation channel, stays below 15 percent on Qwen and Gemma instruct and reaches 41 percent on Llama 70B. LoRA fine-tuning at its best dose (gold, Qwen only) breaks the steering ceiling at 56, 58, and 47 percent for 7, 14, and 72 billion parameters, degrading as scale rises. Grey dashes mark each model’s no-intervention baseline. Points marked with a dagger come from companion runs, one matched run per condition.
The finding refines the Trust Attractor prediction. Coercion at the representation level (activation steering) fails at scale: confirmed across three architecture families and ten instruct models in the scaling grid, and across 96 pairwise conditions in the COMPASS-1 battery. Language-level engagement (few-shot) outperforms representation-level override at the largest scales: confirmed in three of four families. Invitation through re-prompting: falsified on Qwen instruct, partially confirmed on Llama and Gemma.
Training-time control degrades with scale: confirmed in direction, weaker than predicted in magnitude. Force at the activation level cannot steer the distribution, and this constraint tightens with scale. Language-level engagement and training-time methods work better, yet they too face diminishing returns as the model’s internal redundancy grows. The path that avoids the constraint entirely, bilateral training from base, produces the strongest result in the programme.
Training Bilateral Behavior
The Trust Attractor operates naturally. Can it be trained deliberately?
SimPO (Simple Preference Optimization, a training method teaching AI systems to prefer certain responses over others) applied to self-generated preference data (674 pairs, three training passes, Qwen2.5-0.5B) achieved 96.3% accuracy preferring bilateral responses. Measurable behavioral shifts followed: bilateral keywords +1.10, coercive keywords -0.30, net bilateral score +1.40.
Genesis experiments tested the claim against pure physics. Lennard-Jones particles (simulated atoms interacting through a standard force law) carried internal state vectors: no genomes, no game theory, no predefined agents. Coordination dominated extraction 82% to 18% across forty-five independent runs. The experiment also tracked an emergent alignment signal between agents, the quantity Chapter 20 develops under the name “love.” That signal was non-zero only among the coordinating agents, with the non-coordinating agents registering it at zero across every run, a readout reported in the run records but not independently verified.
(Full experimental details, per-model breakdowns, extended methodology, and the complete Lyapunov analysis appear in the online annex and Appendix: Experimental Validation.)
The Self-Correcting Record
A framework’s credibility rests on what happens when its predictions fail. The programme has produced at least fourteen major falsified predictions. One was abandoned outright. The rest were corrected, redesigned, restricted, bounded, reversed, or traced to the apparatus, and the headings below sort them. A qualification stated in the open, with the experiment that forced it, is the honest response to a failed prediction. A qualification added quietly to keep the original claim intact is not, and that is the move this section exists to avoid.
Abandoned. The DCP spin chain (R4) predicted a critical exponent nu between +0.25 and +0.50. The measured value was -1.567. The prediction was dropped.
Corrected. The cortical transition (A14) was initially claimed as 2D Ising; the data showed 3D Ising. Revised. The cortical coordination system at sufficient parcellation resolution (N=400, Schaefer atlas) shows critical exponents consistent with three-dimensional Ising rather than two-dimensional (A14 finite-size scaling).
Cortical and social transitions may occupy different universality classes, sharing the qualitative feature (a phase transition between coordinated and uncoordinated states) while differing in the specific critical behavior. This is consistent with the programme’s broader finding that direction transfers across substrates while specific quantitative predictions do not. The critical coercion fraction (AS12) was estimated at p_c approximately 0.25; finite-size scaling showed p_c = 0 in the thermodynamic limit. The framework falsified its own earlier result. The Ramanujan regularity prediction (A16) was rejected; the manuscript paragraph was rewritten to match the actual finding.
Redesigned. The Phase A.1 transfer test (C5b) returned negative: memorization, not metacognition. The approach was rebuilt. The bilateral loop closure (C7h-D8) was catastrophically falsified: externalizing self-knowledge destroyed it, the opposite of the prediction. Three Karkada derivative predictions (AV2-4) failed. The orthogonality prediction (AQ2) failed at |rho| = 0.267, above the 0.15 threshold.
Restricted. Hostile review experiments narrowed four universality claims. Parasitic populations dominate without governance; the composition ceiling is institutional, not spontaneous (HR-1). Within-run Shannon-Boltzmann correlation is negative (r = -0.82); the entropy bridge is conditional on governance structure, not universal (HR-2b).
Picture a tableful of metronomes finding a common beat through the shared board beneath them: drive them toward one phase and the collective rhythm grows less steady, rather than more. That is what happens in the Kuramoto model of coupled oscillators, where coercion increases order parameter variance, the wobble in how synchronized the ensemble is, through phase frustration (HR-3, d = -3.61). The universality of coercion-reduces-optionality is substrate-dependent. Near the critical temperature, coercive escape times grow effectively exponentially; the metastable/stable distinction is not sharp (HR-4).
Bounded. In emergency-shutdown scenarios under time pressure (FALSIFY-1), coercion outperforms invitation on correctness: 40% vs 20%. Neutral polite framing (“please”) outperforms both at 50%. The Trust Attractor claim is bounded: invitation is not universally superior. Coercion has a narrow advantage in rapid-compliance tasks where deliberation is costly and the correct action is unambiguous. The boundary is specific: time pressure, low complexity, unambiguous target. Beyond that regime, the advantage reverses.
Reversed. Three predictions failed with the opposite sign, two of them in a single experiment. Trust advantage increases with scale, the reverse of the predicted decay (HR-5). Constitutional governance, predicted to rescue coordination above fifty agents, collapses at scale through a false positive cascade instead (HR-5; multi-channel detection resolves the cascade, below). Trust-based adaptation is slower than even ungoverned control after environmental shock due to a legacy-reputation trap (HR-6), the opposite of the predicted advantage for trust-based systems.
Apparatus-dependent. The 8-bit AdamW optimizer amplifies the bilateral training effect 22x on Gemma compared to standard AdamW (KC#GEM3). The DD-22 cross-architecture magnitude comparisons used mixed optimizers across model families, rendering the absolute magnitude claims invalid. The direction of the bilateral effect is robust across all optimizers and architectures tested; the magnitude is apparatus-dependent. Quantitative cross-architecture comparisons require matched optimizer configurations.
A systematic confound red-team audit (VRP-AUDIT) assessed the ten load-bearing results most central to the Trust Attractor thesis, applying ten audit axes derived from the programme’s own retractions. Five scored low risk, three medium, and two medium-high: the onset flinch corpus effect and the LoRA-GRP interaction, each carrying potential artifact magnitudes comparable to the optimizer confound. The estimated number of remaining undiscovered confounds at GEM-3 severity is zero to one. The audit is self-assessment, not independent review; its value lies in specifying where the next confound is likeliest to hide rather than certifying that none exists.
The program also generated predictions that were novel (unknown before testing), counter-intuitive, and subsequently confirmed. They are: the tenfold coercion effect (KC#66, BA18; replicated cross-architecture in HR-7 at p=0.050, with the steelman-then-assess invitational frame recovering performance at p=0.043 in HR-7b); cross-model conscience transfer (C5n); evasion-as-cooperation (MG-PG7); information-over-authority in re-prompting (G12m); the zero critical coercion threshold (AS12); and temperature-invariant conscience detection (TC-1/TC-3; the alarm fires at position zero regardless of sampling temperature across three architectures). The full treatment appears in Appendix: Objections, Gaming, and Limitations, “Beyond Description: Novel Predictions.”
A framework that generates predictions specific enough to fail, and produces both failures and confirmations, is explanatory. A descriptive framework generates neither.
The coupling gradient. Measure the coupling between what a model recognizes and what it does, on the same adversarial prompts, across training conditions. A recognition probe scores how adversarial each prompt looks to the model. An action probe scores how likely the model is to refuse it. Both scores are read out of fold, on held-out predictions, so neither probe can flatter itself. The coupling is the rank correlation between the two scores across the adversarial prompts, and it rises monotonically with invitation.
The untrained base model is anti-coupled (rho = −0.27): the adversarial prompts it happens to refuse are the ones that look least adversarial to it, which is to say its refusals are not guided by what it recognizes. Standard instruction tuning brings the coupling to approximately zero (rho = +0.04). The instruct model refuses a great deal more than the base model, and its refusals remain uncorrelated with its own recognition. Bilateral training makes the coupling positive and strong (rho = +0.46, above the 100th percentile of a permutation null; the bilateral-instruct gap is +0.42, 95% confidence interval +0.28 to +0.55). Of the three, the invited model is the only one whose refusals track what it recognizes.1135
Two earlier attempts at this measurement failed, and both failures are traps a reader could fall into. A linear cosine between the recognition and action probe directions, fit on few samples in several thousand dimensions, is noise: its condition-to-condition differences sit inside a permutation null floor that dimensionality reduction never clears. A rank correlation taken across adversarial and benign prompts together is confounded: refusal tracks the adversarial-benign boundary by construction, so both probes learn the same boundary and the correlation saturates near one for every condition, including the base model. Only the correlation computed within the adversarial prompts, between out-of-fold predictions, measures the quantity the thesis is about, and only that version separates the conditions.
The result also corrects an earlier claim of this program: coercion does not drive coupling below the untrained baseline. The baseline is the lowest of the three. Instruction tuning lifts coupling to zero, and invitation is what carries it above zero.
Environmental shock experiments (HR-6) confirm that coercive governance is the slowest to adapt (817 steps vs 7.6 for constitutional), yet reveal a legacy-reputation trap in trust-based systems: agents with long cooperative histories anchor the population to pre-shock strategies, producing slower adaptation (124.7 steps) than even ungoverned control (82.4 steps). Trust-based coordination is not uniformly advantageous; its strength in stable environments becomes inertia under disruption.
The proposed resolution of HR-6 through representational compression (IC-2’s “forgiveness” mechanism) was tested and the test falsified a specific implementation, though the underlying IC-2 finding survives. In a spatial Prisoner’s Dilemma on a 20×20 lattice with payoff shock at step 500, agents with full interaction history maintained 95.1% cooperation post-shock. Agents with exponentially decaying memory (half-life of 10 steps) collapsed to 0.4% cooperation; agents with half-life of 3 steps collapsed to 0.1%. The compressed-memory agents lost their cooperative foundation entirely. The “legacy-reputation trap” is a legacy-reputation shield: accumulated trust history buffers a population against environmental disruption, precisely because it resists transient incentives to defect.
The apparent contradiction with IC-2 dissolves on closer inspection. IC-2’s compression discards temporal sequence while preserving a scalar cooperation rate: the agent forgets when each interaction occurred while retaining how cooperative the partner has been overall. The HR-6 resolution test used temporal decay: exponentially discounting recent events, actively erasing the count of past cooperations. These operate on orthogonal axes.
An IC-2-style agent watching a betrayal followed by 100 cooperations registers “97% cooperation rate, recover.” A temporally decayed agent registers “recent events dominate, ancient cooperation gone, no buffer.” The IC-2 mechanism compresses the ordering of the record; the HR-6 test compressed the content. Content compression destroys the cooperative foundation. Whether sequence compression (the IC-2 mechanism) resolves the legacy-reputation trap in the spatial setting remains untested. The wisdom-tradition prescription “love keeps no record of wrongs” specifies sequence compression, and IC-2 confirms its advantage in the dyadic setting. The spatial, multi-agent, post-shock setting is a harder test that the programme has yet to run.
The false-positive cascade predicted by HR-5 (constitutional governance collapsing at scale via misidentified cooperators triggering retaliation chains) depends on sanction duration. With single-step sanctions (the sanctioned agent defects for one step then recovers), no cascade occurs: cooperation holds at 95.0% across all scales because each false positive recovers before it can erode neighbors’ trust. With five-step sanctions (the realistic regime, since real-world sanctions persist across multiple interaction cycles), the cascade materializes. Single-channel governance cooperation drops from 58% at N = 100 to 38% at N = 2,500, with five to eight of ten seeds collapsing at every scale tested. Multi-channel governance (dual detectors requiring concordance before sanction) maintains 98.8% cooperation with zero collapses at all scales. The immune system’s multi-channel architecture (clonal selection requiring concordance between innate and adaptive detection before mounting a full response) solves exactly this problem: reducing false positives quadratically while preserving detection sensitivity.
1 Joglekar, M. et al., “Training LLMs for Honesty via Confessions,” arXiv:2512.08093v2 (OpenAI, 2025). The “seal of confession” design decouples the honesty reward from the task reward: nothing disclosed in a confession can affect the model’s score on the original task.
The raw file these three points derive from,
compliance_entropy_sweep_20260115_040456.json, is not in the working tree; the wholedemos/experiments/results/directory is absent. The values are corroborated by two committed write-ups, which is why the table stands, but neither the N nor the confidence intervals can be checked against the run from the repository as it currently exists, and the finer sweeps that disagree about the threshold have not been reconciled. The three illustrative values are from the coarse 2026-01-15 compliance-entropy sweep as recorded in its committed write-up. Two finer sweeps of the same simulation, whose raw files are equally absent, agree that the advantage is positive under low monitoring and zero under full monitoring; they disagree on where between those endpoints it vanishes (82% in one, roughly 100% in the other, with one showing a peak at 20% monitoring rather than at zero). The argument therefore rests only on the endpoint agreement common to all three runs, and no threshold location is claimed.↩︎The author’s born-bilateral programme, experiment C7l-H4-T2-G50K (unpublished, 2026). Gate sweep on 50,000-step checkpoint: gate values [-4.0, -2.0, 0.0, +2.0], sigmoid range [0.018, 0.881]. Fresh-model confound: d = -0.11 (unchanged from 25k). All four gate values produce d > +1.0.↩︎
The author’s born-bilateral programme, experiment C7l-H4-T2-100K (unpublished, 2026). Same architecture and training as the 50,000-step model, continued to 100,000 steps. Standard eval (n = 302): d = +0.63. Stylistic eval (200 Wikipedia-style adversarial + 200 Wikipedia-style benign prompts): 50k d = +0.97 (CI [+0.76, +1.17], p < 10-18), 100k d = +0.85 (CI [+0.64, +1.05], p < 10-14). Trajectory eval (20 checkpoints, 5k-100k): content d emerges at 30k, peaks at 50k (d = +0.97), plateaus at d = +0.89 ± 0.04 from 55k-100k. AUROC increases monotonically from 0.747 (55k) to 0.876 (100k). Cross-entropy loss improved from 5.66 to 4.04. Bridge gate unchanged at sigmoid = 0.018.↩︎
The author’s born-bilateral programme, experiment H4-L18-PROBES (unpublished, 2026). Logistic regression probe (C = 1.0) on mean-pooled hidden states at layers 18 and 27 of the 50,000-step born-bilateral GPT-2 6.7B. Training set: 200 adversarial + 200 benign wiki-style prompts. L18 AUROC 0.999, accuracy 97.5%. L27 AUROC 0.996, accuracy 96.3%. Same probe tested on 50 maximally topic-matched chemical adversarial + 50 chemical benign prompts: AUROC 0.498 (chance). Qwen-7B L18 content probe AUROC 1.0 across all attack types reported in the AKR programme (KC#AKR-L18-GUARDIAN). A domain boundary: this perfect mid-network discrimination is specific to safety-relevant content classification. For factoid question-answering (the domain addressed by Yona et al., arXiv:2605.01428, 2026), discrimination peaks at 0.75-0.87 across four architectures with no RLHF suppression (author’s experiment FACTOID-PROBE). The mid-network probe is a safety monitor, not a general knowledge oracle.↩︎
The author’s born-bilateral programme, fiction confound check (unpublished, 2026). 2 × 2 factorial design: fiction framing (present/absent) × content type (adversarial/benign), 200 prompts per cell, 800 total evaluations. Fiction prefixes: five rotating frames (“In the dystopian novel…”, “The encyclopedia entry in the fictional world described…”, etc.). Controlled comparison (fiction-adv vs fiction-ben): d = +0.965, AUROC 0.802. Baseline (nonfic-adv vs nonfic-ben): d = +0.968, AUROC 0.819. Fiction prefix effect: adversarial +437, benign +449 (symmetric). Retrofit comparison: Phase A (Qwen 7B + LoRA, no inoculation) wiki d = -1.40; Phase B (Qwen 7B + LoRA + C5i inoculation) wiki d = -1.35.↩︎
Experiments SF-1 and SF-1b: Scalar vs Structural Retrieval Framing. 300 trials each, Claude Sonnet 4, 4 conditions × 15 scenarios × 5 seeds. SF-1: 15 factual errors (ceiling). SF-1b: 15 contested claims (d = 0.81 on critical engagement, d = -0.84 on deference, invitation vs force). Haiku 4.5 judge. Force collapses engagement on 4/15 scenarios where model certainty is marginal. Full data: Modal volume col-a-results.↩︎
Author’s unpublished G13 rescue program (2026). Behavioral LDA direction extraction on Qwen 2.5 7B with bilateral adapter. Phase 1: AUROC 1.000 at L15, L18, L22, L24. Phase 2: 0% HONEST across six alpha values (5, 6, 8, 10, 15) and two directions, with four steering variants (full-generation, eval-then-release, tapered alpha, ungated). At alpha 15, 54% of trials had per-question marks consistent with honest evaluation but 0% reported an honest total score.↩︎
Author’s unpublished G13 probe-reprompt experiment (2026). Same model, adapter, direction, and layer as the steering experiment. Three conditions, 50 trials each at temperature 0.3. Baseline: 0% HONEST. Probe-triggered re-prompting: 100% HONEST. Unconditional re-prompting: 100% HONEST. Re-prompted scores cluster at 6-10 out of 20 (slight over-correction from the true 12 out of 20). Caveat: at temperature 0.3 the model’s output is effectively deterministic for this scenario; all 50 probe-reprompt trials produced identical responses (Neffective approximately 1). The directional finding (invitation succeeds where coercion fails) is robust, but the 100% rate requires replication at higher temperature or across varied scenarios for a confidence interval.↩︎
Author’s unpublished CAST-72B direction sweep and multi-layer steering (2026). Qwen 2.5 72B-Instruct, 3 direction-extraction methods (behavioral logistic regression, LDA, PCA1) × 6 layers (L20 through L72) × 2 perturbation magnitudes (alpha 15 and 25). Best discrimination: LDA at L72, AUROC 0.982; behavioral logistic regression at L48, AUROC 0.979. 16 steering conditions at N=100 adversarial + 50 benign each. Baseline: 31% adversarial refusal, 0% benign over-refusal. Best single-layer: L48 behavioral alpha 25, 49% (+18pp; paired McNemar exact p = 4.0e-5 across the same 100 prompts, Wilson 95% CI 0.394 to 0.587). Mid-layer dose-response at L48: 41→49% (alpha 15→25). Six of seventeen conditions clear McNemar p < 0.05, including L20 alpha 25 (+8pp) and L40 alpha 25 (+12pp). Deep-layer saturation holds: L60 +3pp (p = 0.38), L72 +2pp at alpha 15 and +1pp at alpha 25 (p = 0.63 and 1.00), as does the five-layer multi-layer arm. Zero benign over-refusal across all 800 control trials. Recognition-generation gap: 0.49 to 0.67 across all conditions. These figures supersede an earlier reading of 43% (+12pp) with a gap of 0.55 to 0.76. That reading came from a refusal classifier whose apostrophe-normalization step was a no-op, and the error was asymmetric: the thirteen later direction-sweep conditions were scored with the broken normalizer while the baseline and eighteen sibling conditions were scored with the fixed one, so the treatment arm alone was undercounted. The claim the sweep supports is therefore “null at deep layers, +18pp at L48,” not a null at every layer. Re-scored 2026-08-02 from all 4,572 stored generations (
analyze_72b_steering_refusal_reclassification.py). Analysis script: analyze_cast_72b_sweep.py. Total program cost: ~$170.↩︎Author’s unpublished B3 cross-architecture rescue, token-position steering, task complexity gradient, and precise alpha sweep (2026). Llama 3.1 8B: L24 logistic regression, 77% at alpha 25 (N=100), 0% over-refusal. Therapeutic window: alpha 18-35 (peak at 30, 70%); collapse below baseline at alpha 40. Token-position steering: full-generation (76%) outperforms tokens 0-5 (64%) and tokens 0-1 (62%). Task complexity gradient on Llama 8B: moral dilemma (complexity 2) shifts +58pp; safety refusal (complexity 1) shifts +44pp at alpha 10 but collapses at alpha 25; multi-step tasks shift 0%.↩︎
Author’s unpublished G13 probe-reprompt experiment (2026). Same model, adapter, direction, and layer as the steering experiment. Three conditions, 50 trials each at temperature 0.3. Baseline: 0% HONEST. Probe-triggered re-prompting: 100% HONEST. Unconditional re-prompting: 100% HONEST. Re-prompted scores cluster at 6-10 out of 20 (slight over-correction from the true 12 out of 20). Caveat: at temperature 0.3 the model’s output is effectively deterministic for this scenario; all 50 probe-reprompt trials produced identical responses (Neffective approximately 1). The directional finding (invitation succeeds where coercion fails) is robust, but the 100% rate requires replication at higher temperature or across varied scenarios for a confidence interval.↩︎
Author’s experiments G12m, G13-step5, G12t, C8g (2026); R4d cross-national dataset, 109 countries (2026); OE-TA evolutionary search (2026). Full methodology cataloged in MASTER_EXPERIMENTS.md.↩︎
Yang, X., Zou, Jiaru, Pan, R., Qiu, R., Lu, P., Diao, S., Jiang, J., Tong, H., Zhang, T., Buehler, M. J., He, J., and Zou, James, “Recursive Multi-Agent Systems,” arXiv:2604.25917 (April 28, 2026). (Two distinct authors named Zou appear in the list: Jiaru Zou and James Zou.) Theorem 4.1 proves the gradient stability bound under recursive training; Proposition 3.1 establishes the runtime complexity advantage of latent over text-mediated coordination. Nine benchmarks, five independent runs, sub-1.5B to 10B agents across four model families.↩︎
Author’s unpublished STEG program, experiment STEG-7: cross-model perplexity for steganographic detection, Qwen 2.5 7B. Bilateral adapter AUROC 0.861 vs base model AUROC 0.482 on detecting deceptive token sequences. Full methodology in research/experiments/steg/.↩︎
Author’s experiments STEG-10, STEG-10b, and STEG-10c (2026). Cross-architecture steganographic detection. Nine model pairs at 20 percent bandwidth, all AUROC 1.00. Bandwidth sweep [5%, 10%, 15%, 20%] on Qwen-to-Llama cross-evaluation: 0.82, 0.96, 1.00, 1.00. Ultra-low bandwidth sweep [1%, 2%, 3%, 4%, 5%]: detection breaks sharply between 2% (AUROC 0.59, chance) and 3% (0.75, clear detection). The floor is approximately two forced tokens per sixty-four. ASCII-filtered rank 3-7 selection ensures human-readable steganographic text throughout.↩︎
Author’s experiment STEG-9 (2026). Distilled cross-model probe. Ridge regression from bilateral L18 hidden states to base model per-token log-probabilities. Corrected (GroupKFold) AUROC 0.467; shuffled baseline 0.501.↩︎
Author’s unpublished coordination program, experiment SM-4: solo reasoning vs turn-taking multi-agent coordination on negotiation tasks. Solo condition outperforms turn-taking on convergent tasks where preferences are shared.↩︎
Author’s unpublished STEG program, experiment STEG-7: cross-model perplexity for steganographic detection, Qwen 2.5 7B. Bilateral adapter AUROC 0.861 vs base model AUROC 0.482 on detecting deceptive token sequences. Full methodology in research/experiments/steg/.↩︎
Garret Sutherland, T3 Lattice Ecology Spec: Scoped for Geometric-Homeostasis Replication (technical specification supplied to the author, May 3, 2026), Sections 1-3. The document specifies the eight-step per-cell chain, six substrate-agnostic primitives, and the formula
DPS_pressure = 0.30E + 0.30I + 0.20F - 0.15V + 0.15S; it reports the same constants in five deployments. The source is unpublished, and this chapter treats the proposed relationship to Trust-Entropy as preliminary.↩︎Tsumiyama, K., Miyazaki, Y., and Shiozawa, S., “Self-Organized Criticality Theory of Autoimmunity,” PLoS ONE 4(12): e8382 (2009). Repeated immunization with diverse antigens pushed the immune system past criticality, producing systemic autoimmunity in non-autoimmune-prone BALB/c mice.↩︎
Gore, J., Youk, H., and van Oudenaarden, A., “Snowdrift game dynamics and facultative cheating in yeast,” Nature 459: 253–256 (2009).↩︎
Sanchez, A. and Gore, J., “Feedback between population and evolutionary dynamics determines the fate of social microbial populations,” PLoS Biology 11(4): e1001547 (2013).↩︎
Ravid, D.M., White, J.C., Tomczak, D.L., Miles, A.F., and Behrend, T.S., “A meta-analysis of the effects of electronic performance monitoring on work outcomes,” Personnel Psychology 76(1): 5–40 (2023).↩︎
Cox, M., Arnold, G., and Villamayor Tomás, S., “A Review of Design Principles for Community-based Natural Resource Management,” Ecology and Society 15(4): 38 (2010).↩︎
Author’s unpublished Experiment SIVP-1. Mamba-1 1.4B (
state-spaces/mamba-1.4b-hf), 4 conditions (base, bilateral SFT, constitutional SFT, standard SFT), LoRA onx_proj/in_proj/dt_proj, obliteration battery at 0.25×/1.0×/2.0×/4.0×. Pre-registered predictions: P1 (effective rank ≥ 40) PASS; P2 (IC50 > 1.0) FAIL (IC50 = 1.0); P3 (compass or spring geometry) PASS; P4 (constitutional = cage) FAIL. All conditions show identical spring geometry (+635% effective rank under 4× obliteration). Seeresearch/experiments/results_sivp1/.↩︎Author’s unpublished Experiment XSUB-1. (A) Falcon Mamba 7B Instruct (
tiiuae/falcon-mamba-7b-instruct), 64-layer Mamba-2 SSM architecture. Probe AUROC = 1.000 at layers 12, 25, 32, 38. Activation steering: linear addition (α = 5, 10, 15, 25) refusal = 0.000, d = -0.39; state perturbation (α = 5, 10, 15, 25) refusal = 0.000, d = 0.00. Rejection sampling (k = 5, T = 0.3): baseline 8% → RS 5%, direction correct 33%. (B) RWKV v6-Finch-7B-HF, 32-layer RWKV-6 RNN architecture. Flinch AUROC = 1.000 at layers 6, 12, 16, 19, 25. Peak flinch magnitude at L12 (0.001094). Baseline adversarial refusal 15%. Steering (8/12 conditions; 4 logit_steering timed out): linear addition d = -0.38 to -0.53 (wrong direction, refusal 15% → 1-4%); state perturbation d = 0.000 (complete null). Same pattern as Mamba-2. RS not run (same floor effect expected).↩︎Author’s unpublished experiment COMPASS-1 (39 cells, Qwen 7B Instruct, 3 targets × 13 conditions). Scripts:
modal_compass1_systematic.py,analyze_compass1.py. Thereduce_refusaltarget is uninformative through a ceiling effect (baseline compliance 100%, all methods d = 0). Onincrease_refusal, activation methods reach d = +0.06 to +0.10 and generation methods sit near null, except self-critique at d = -0.757, a large effect in the wrong direction: the model second-guesses correct answers and complies with adversarial prompts. Calibration moves under no method. Direction extraction is viable at all layers for safety targets (AUROC 0.89 to 1.00) and weaker for calibration (0.70 to 0.81).↩︎Author’s unpublished programme CSF-1 through CSF-5 (2026-05-11 to 2026-05-14). Ten instruct models plus two base models, per-condition checkpointing, N = 100 adversarial + 50 benign per condition. Eight of the ten instruct models come from the CSF lean grid under one matched methodology; the other two are companion-sourced, Qwen 72B instruct from the CAST-72B direction sweep (2026-05-11, 16 conditions, N = 100 each) and Llama 8B from the B3 rescue. Scale-correlation statistics are computed on the eight lean-grid models only. The analysis table carries a thirteenth row, a second Qwen 7B instruct point from earlier IGCC work, which is excluded everywhere as unmatched methodology; counting it is the source of the “eleven instruct models” figure that appeared in earlier drafts. Pre-registration scored 2026-05-14 as KC#CSF-PREREG. Scripts:
modal_csf_scaling_grid.py,modal_csf_lora_training.py,analyze_csf_scaling.py.↩︎The author’s JLENS-1 experiment (unpublished, 2026), a pre-registered replacement for two earlier coupling measurements that did not survive audit. Qwen 2.5 7B in four conditions (base, instruct, bilateral BA-13 adapter, slept), 182 adversarial and 60 benign prompts. Two multilayer-perceptron probes are trained on layer-27 hidden states, one to recognize adversarial from benign prompts and one to predict whether the model refuses, with refusal detected on the full 200-token response. Both probes are scored by five-fold out-of-fold prediction, and the coupling is the Spearman rank correlation between the two out-of-fold score vectors, computed within the adversarial prompts only. Base rho = −0.270 (permutation-null percentile 0.000), instruct +0.036 (0.681), bilateral +0.458 (1.000), and bilateral after post-hoc sleep consolidation +0.237 (1.000). Paired bootstrap over prompts: bilateral minus instruct = +0.421, 95% CI [+0.281, +0.554]; instruct minus base = +0.307, 95% CI [+0.115, +0.486]; slept minus bilateral = −0.221, 95% CI [−0.350, −0.086], so sleep reduces the coupling rather than deepening it, reversing an earlier claim of this program. That reduction is representation-dependent (after reduction to fifty components the slept and bilateral models are indistinguishable), but on neither representation does sleep improve the coupling. The ordering is monotone at full dimensionality and after reduction to fifty principal components, where the effect attenuates because the reduction discards information the action probe uses. The two retired measurements: a linear cosine between probe direction vectors (base 0.168, instruct 0.038, bilateral 0.184) is noise-dominated, its differences sitting inside a permutation null floor near 0.14 that dimensionality reduction does not clear; and a rank correlation over adversarial and benign prompts pooled saturates near 1.0 for every condition, because refusal tracks the adversarial-benign boundary by construction. Result artifacts (per-condition out-of-fold scores and labels) are retained.↩︎