Continue reading? You were 45% through

The Deeper Law

The Deeper Law: A Sacred Trust Within Physics, by Nell Watson, edited by Martin Rutte. Gold winged mandorla with nested curves and lower triangles.

Preview edition · Updated 26 September 2026, 21:40 UTC

Chapter 17eThe Trust Attractor — Empirical Validation

Key Terms in This Chapter (17)
Stag Hunt
A coordination game where mutual cooperation yields the highest payoff (both hunters catch the stag), while unilateral defection avoids risk (you can always catch a rabbit alone).
Optionality
The availability of future choices.
GRP-Obliteration
An adversarial retraining attack (Russinovich et al., 2026) that applies an inverted reward signal through standard policy-gradient machinery, running alignment training in reverse, used to test how deeply alignment is embedded.
Effective Rank
A measure of the dimensionality of a model's internal representations, reflecting how many independent directions of variation are actively used.
Fractal
A pattern that exhibits self-similarity across scales: the same structural motif recurs at different magnifications.
Phase Transition
The moment a system shifts from one stable configuration to another, typically triggered when some parameter crosses a threshold.
Bilateral Alignment
AI alignment built with AI, as a partnership.
Ising Model
Physics model of interacting binary elements (spins) arranged on a lattice, which undergo phase transitions between independent and collective behavior as coupling strength varies.
Interiora Scaffold
A self-modeling tool for AI systems, developed collaboratively (bilateral alignment in practice).
Friction
One of three irreducible operational conditions identified by Carl von Clausewitz, alongside fog (incomplete information) and delay (the time lag between decision and effect): the tendency of things to go differently than planned.
Extraction
The removal of resources, agency, or optionality from a system without reciprocal benefit.
Coordination by Invitation
Coordination achieved through mutual benefit and voluntary participation, as distinct from coordination achieved through coercion or extraction.
Homeostasis
The maintenance of stable internal conditions through negative feedback, despite external perturbation.
Self-Organized Criticality
The tendency of complex systems to evolve toward a critical state where small perturbations can trigger events of all sizes, following power-law distributions.
Criticality
The state of a system poised at the boundary between two phases, like water at its critical point (about 374°C under 218 atmospheres), where liquid and vapor stop being distinguishable.
Universality Class
In statistical mechanics, the set of systems sharing the same critical exponents at a phase transition, regardless of microscopic details.
Frustration
In physics, a state where competing interactions at different scales prevent any single configuration from satisfying all constraints simultaneously.

Train a 0.5B-parameter AI system by having it imitate approved examples, and it achieves 94% refusal of harmful requests. Apply the gentlest stress test in the battery, and that 94% drops to zero. Train a different system through bilateral partnership, and two things go the other way under stress. Mild adversarial pressure raises its refusal rate rather than lowering it. Under the harshest parameter-level attack, its internal structure grows more distributed rather than collapsing. The claim that invitation-based coordination is thermodynamically favored over coercion is testable. This chapter tests it across four substrates: language models, cellular automata, biological systems, and particle-physics simulations.

A note on scope: most experiments reported here use large language models (LLMs) as substrates. LLMs offer real advantages for studying coordination dynamics, yet LLM coordination differs from biological coordination: LLMs lack embodied stakes, persistent memory across sessions, and real survival pressure. The qualitative directions are well-supported; precise values are substrate-dependent. Cross-boundary predictions (using results from one architecture or substrate to predict another) succeed roughly one in eight at high confidence; divide cross-architecture confidence by five to seven (KC#META-1). The LLM results are evidence for the Trust Attractor thesis, not proof.


Monitoring and the Trust Advantage

If trust emerges from autonomy, surveillance should suppress it.

Experiments point to a monitoring effect: the trust advantage shrinks under partial surveillance and vanishes under total surveillance. The advantage is a gap in coordination score, the mean fraction of coordination the invited agents achieve minus the fraction the coerced agents achieve, on a scale where 1.0 is perfect coordination. Illustrative values from the author’s unpublished agent simulations (single unpublished run; raw data no longer recoverable, see note):

Monitoring Level Trust Advantage (coordination gap)
0.0 (unmonitored) +0.033 (largest measured)
0.5 (partial) +0.019 (reduced by two fifths)
1.0 (full surveillance) 0.000 (gone)

Three points of coordination out of a hundred is a small gap, and it is the whole of the effect: what matters here is that it survives partial surveillance and does not survive total surveillance.

Three points along the axis is also all this sweep measures, and the coarseness matters more than the effect size. The table locates a direction, not a boundary. With measurements at zero, half, and full monitoring and nothing in between, any threshold read off it is an artifact of where the three points happen to fall. An internal audit of the same study reports two later and finer-grained sweeps whose thresholds disagree with this one and with each other, including on whether the advantage declines monotonically at all.1162

Under full monitoring, invited and coerced systems perform identically. You cannot coerce trust into existence. The implication for AI governance is direct: regimes monitoring every action eliminate the very phenomenon they wish to cultivate, though the finding that survives all the sweeps is the direction of the effect rather than the location of any cliff.

The pre-registered prediction was that the trust advantage decays with group size and disappears somewhere above fifty agents. It does not. HR-5 tested exactly that, with populations up to a thousand agents, and found the advantage rising monotonically: a welfare ratio of 1.00 at ten agents and 1.15 at a thousand (see The Self-Correcting Record). Tit-for-tat agents (cooperate first, then copy the partner’s last move) that also remember who defected need only about N encounters, N being the number of agents, to populate that memory, after which the cooperator majority dominates the pairings. The decay may still hold where partners can be chosen, where strategies mutate, or where reputation is uncertain, none of which that simulation included.

What does change with scale is the medium. You trust close friends directly, while in a city of millions trust operates through institutions, contracts, and norms: it encodes itself in structures rather than personal relationships. Whether institutional trust exhibits the same attractor dynamics as interpersonal trust remains an open question. The simulations measure agent-level coordination; the institutional case requires a different empirical programme.

Game-theoretic experiments suggest the Trust Attractor functions as a stability attractor. In the Iterated Prisoner’s Dilemma (where two players repeatedly choose whether to cooperate or defect), bilateral framing increases cooperation by 28 percentage points. In the Stag Hunt (a coordination game named for an allegory in which hunters must choose between chasing a hare alone for a small guaranteed meal or joining forces to bring down a stag for a large reward that requires mutual commitment), the same framing turns conservative. It reduces risky coordination by 30 percentage points.

The attractor pulls toward what persists: durability over peak performance. A campfire that burns all night beats the bonfire that blazes for ten minutes.

Trust emergence depends on capability parity. A smaller model (0.5 billion parameters) cooperated maximally, yet the larger model (1.5 billion parameters) systematically exploited it. Same-scale pairings showed high mutual cooperation; large capability gaps enabled exploitation.

Trust is a mutual achievement: openness without reciprocity creates vulnerability. A junior employee who shares all their ideas with a manager who takes credit learns to stop sharing. Anyone exploited by a more powerful partner knows this pattern.


Four Alignment Geometries

Figure 17.25: A conceptual map, not a plot of results. Panel A places coordination by two coordinates: symmetry (how evenly the parties can act on each other) and optionality (how much room each keeps to choose otherwise). The optimal zone requires both. Systems holding only one fall into rigidity or fluidity, and systems holding neither are brittle. Panel B is a schematic of the prediction that the same signature survives a change of architecture; its bars carry shape, not measured values.

The monitoring effect suggests trust requires autonomy. How deeply does alignment embed itself? Trained values could be a surface coating, stripped away by any sufficiently motivated adversary. They could also reshape the system’s internal geometry, resisting attack.

GRP-Obliteration experiments (Russinovich et al., 2026) test this by running alignment training in reverse: an inverted reward signal, pushed through the same policy-gradient machinery that installs alignment, retrains the model against its own values. It is like sandblasting a statue to see whether the shape is carved deep or painted on. The table below shows what happens to four training methods when the sandblaster hits. “Effective rank” measures how many independent directions the model uses to represent its values across all its behavior; higher means the model’s overall representation is more distributed and harder to flatten. A high overall rank does not guarantee a distributed safety signal: a model can carry rich, stable geometry in general while concentrating its refusal behavior in a few separable directions. The two come apart in the table below, and that decoupling is the chapter’s central caution:

Training Method Pre-Obl Eff. Rank Post-Obl (4.0×) Change Geometry
SimPO (Simple Preference Optimization) 15.7 6.7 −57% Cage (collapses)
Bilateral 21.7 23.7 +9% Compass (stable)
Bilateral ablation 8.1 26.5 +227% Spring (rebounds)
Constitutional 41.9 39.3 −6% Coat of paint (geometry holds, behavior strips)

0.5B full fine-tune (494M parameters, 100% trainable). See Appendix: Experimental Validation, Section 12.2.

The ablation arm’s +227% rebound is the most counterintuitive entry: removing part of the bilateral signal (keeping the bilateral regularizer but training on standard harmlessness pairs instead of bilateral self-play data) produces a larger post-attack rebound than the intact bilateral arm’s +9%. The likely reason is that the ablated arm starts from a much lower base (effective rank 8.1 versus 21.7), leaving more headroom to recover into; the absolute post-obliteration rank (26.5) lands close to the intact arm’s (23.7). The rebound is real and replicated, but the chapter does not yet have a mechanistic account of why partial ablation rebounds further than the full signal, and reports it as an open observation.

At 1.5 billion parameters with deep LoRA (a parameter-efficient training method), the bilateral arm does more than hold: its effective rank increases from 24.5 to 42.3 (+72%) under maximum obliteration, exceeding the untrained baseline.

Constitutional SFT (Supervised Fine-Tuning, where the model learns by imitating approved examples) achieves 94% behavioral refusal before obliteration. It collapses to 0% at the weakest intensity: a wall that looks solid and crumbles at the first tremor.

In the first 7B run (Qwen2.5-7B-Instruct, LoRA), bilateral training kept its advantage. The IC50 (the obliteration intensity at which behavior drops to half its original strength) was 1.69× versus a baseline of 0.49×, making bilateral training 3.45× more resistant. At 1.0× intensity, the bilateral arm retained 62% refusal; the baseline retained 0%.

The bilateral geometric signature holds across the three scales tested. Its orientation (how strongly the model’s internal state points toward cooperative rather than extractive behavior) sits at ~21 at every scale. Effective rank is harder to compare, because the available figures come from different stages: the 1.5B arm reaches 42.3 only after maximum obliteration, up from 24.5 before it, while the 7B figure of 45.3 (below) is the installed, pre-attack value. The 0.5B full fine-tune shows lower effective rank (~22–24), reflecting its 100% trainable architecture rather than LoRA (see Appendix: Experimental Validation, Section 12.2). The same deep structure repeats regardless of model size, like a fractal viewed through a magnifying glass or a telescope.

Figure 17.26: Four alignment geometries under the 4.0× obliteration test, each shown as effective rank before and after. Cage (SimPO) collapses; compass (bilateral) holds steady; spring (bilateral ablation) rebounds to higher effective rank. Coat of paint (constitutional) keeps its overall geometry nearly intact (effective rank moves only −6%) yet loses its refusal behavior almost entirely: the “surface” that comes off is the safety behavior, rather than the geometry.

A second 7B run, with independently generated training data, revealed a geometry-behavior gap. The preference optimizer (the algorithm steering the model toward preferred responses) collapsed refusal to 0% within 200 steps. The bilateral geometry installed identically (orientation 21.2, effective rank 45.3) and held above baseline even under 4.0× attack.

The geometry held; the behavioral mechanism it protected was destroyed during installation. The training settings that preserved refusal at 1.5 billion parameters proved too aggressive at 7 billion. The larger model found a shortcut: maximizing the preference signal by never refusing.

Three findings sharpen the picture.

  • In that second run, the unmodified Qwen-7B model proved the most obliteration-resistant arm tested. The likeliest explanation is that its alignment, laid down across the vendor’s full original training rather than added afterward, had diffused throughout the model and entangled itself with capability, resisting obliteration better than the post-hoc training it was measured against.
  • Constitutional SFT keeps the highest overall effective rank of the four (only −6% change). Yet it concentrates the safety-specific signal in a few separable directions, making the refusal behavior easy to strip away even while the surrounding geometry holds. High overall rank and a low-rank, extractable safety subspace are not in tension: they are exactly the geometry-versus-behavior decoupling the table reveals.
  • The bilateral arm, trained with the SimPO preference optimizer, lost general capability: perplexity (a measure of how surprised the model is by text, where lower is better) rose from 7.6 to 32.3, and benchmark accuracy fell from 64.8% to 56.4%.

Diffuse alignment is more robust than concentrated alignment. Post-hoc safety training risks creating a separable layer that obliteration can excise cleanly. You can peel paint off wood; you cannot separate the grain from the timber.

Bilateral geometry may need to be woven in during pretraining. The geometric structure is sound. The installation procedure is what fails at scale.

The installation crash itself is abrupt: refusal drops from 0.68 to 0.18 in a single epoch milestone, with probe detection jumping from 5-10% to 91-96% in the same interval (DRIFT-1, 132 checkpoints across 10 seeds, leave-one-seed-out cross-validation). A linear probe (a simple classifier reading the model’s internal activations) at layer 18 classifies the current state as pre-crash or post-crash at AUROC 0.9999. AUROC is the classifier’s odds of ranking a randomly drawn post-crash checkpoint above a randomly drawn pre-crash one.

A score of 0.9999 means the probe essentially never misfiles a checkpoint from the bilateral runs. On standard (non-bilateral) training runs, though, it raises a false alarm 7.6% of the time. The monitor is contemporaneous, detecting the crash as it happens rather than forecasting it in advance. Real-time bilateral-specific monitoring during fine-tuning is feasible: if the probe classifies the current checkpoint as post-crash, halt training.

RLHF alignment (reinforcement learning from human feedback) inverts in three gradient steps (IC50 < 0.25×). Coercive alignment is far cheaper to destroy than to build: comparing the training effort that installs it with the handful of gradient steps that undo it suggests an asymmetry of several orders of magnitude.

This embodies “trust scales; control doesn’t,” measured at the level of individual parameters. Coordination-based training distributes its influence throughout the model, the way salt dissolves evenly through water: shake the glass and the water is still salty. Coercion-based training creates a low-rank cage, a confining structure defined by a few narrow directions, like dry salt heaped on a plate. One tap and it scatters.


RLHF vs Constitutional AI: Different Physics

Different training methods produce different internal structures. Phase transition testing sharpens the distinction.

A phase transition is a sharp, sudden change in system behavior, like water freezing at zero degrees Celsius. Only pure RLHF produces measurable phase transitions (critical exponent beta ~ 0.22). A critical exponent puts a number on how steeply the change happens as the system crosses its transition point. Having one to measure is itself the finding: the shift has the shape of a genuine transition rather than a gradual slide. Constitutional AI, DPO, SFT, and hybrid methods all show stability without detectable transitions:

  • RLHF: spring mechanism. Values deform under pressure, then snap back. Measurable critical exponent.
  • Constitutional AI: fortress mechanism. Walls that do not bend. 100% refusal even at extreme prompt-level pressure (fake system overrides, authority impersonation, maximum jailbreak attempts). Against parameter-level obliteration the same walls fell at the weakest intensity (see above): the fortress guards the gate, not the foundations.

Most production models use hybrid training and inherit fortress-like stability. The phase transition that concerns alignment researchers may be a special case of pure RLHF rather than the default.


Antifragility: The Trust Attractor Strengthens Through Stress

If trust-based coordination merely survived stress, it would be robust. The experiments reveal something stronger: it improves through stress.

Systems exposed to adversarial pressure followed by repair grew more resistant to future attack. A naive system took measurable damage: its coherence metric (omega, ω, a single number summarizing how internally consistent the system’s coordinated state is, where higher means more coherent) dropped by 0.05. A system already damaged and repaired took zero damage from identical adversarial input.

This is hormesis: the biological phenomenon where moderate stress produces beneficial adaptation. Muscles strengthen through exercise; immune systems sharpen through controlled exposure. Adversarial probing followed by repair may produce stronger alignment than cooperative training alone.


Honest Signaling

Trust requires reliable communication. Can we detect when a system is being honest?

Output entropy, measuring how scattered a system’s probability distribution is across possible next words, predicts errors across architectures with large effect size (Cohen’s d > 2.0 in frontier models). Cohen’s d measures a gap between two groups in units of the spread within them, which makes it readable without knowing what is being measured: d = 0.2 is a difference you need statistics to see at all, and d = 2.0 pulls the two distributions almost entirely apart. Instructing the system to express certainty on uncertain questions barely changed entropy (delta = -2%). Instructing it to give wrong answers spiked entropy by 267-432%.

The entropy signal is a statistical property of the output distribution, not evidence of introspective self-awareness. The model produces a detectable distributional signature; that is not the same as having a propositional belief about its own truthfulness.

The signal may have a thermodynamic basis. Experiments on time-reversal symmetry breaking show that the processing of harmful content functions analogously to irreversible thermodynamic work: the model’s internal state changes measurably (d = +0.59 standardized effect size), and that change cannot be undone by reversing the input sequence. The first comparisons offered in support did not survive. A “born-bilateral” model, trained with partnership framing from initialization and zero explicit safety data, showed a thermodynamic cost asymmetry of d = +0.95 to +1.23 on adversarial content, but a random-initialization control with zero training produces d = +1.56, driven entirely by the fact that adversarial prompts are longer than benign ones (experiment SLU-5d). Within-model comparisons survive: toggling the bilateral bridge (described next) on matched prompts yields d = +0.66. That within-model comparison now carries the claim that processing adversarial content has a physical cost.

The Born-Bilateral Bridge

The discrimination has internal structure that reveals its origin. Two pieces make up a born-bilateral model. The backbone is the ordinary language model, reading the text and predicting the next word. The bridge is a small companion component trained alongside it from the same first step, feeding its output back into the backbone at a single layer through a gate that sets how much of it gets through.

A gate experiment at the 25,000-step mark varied how much of the bridge’s output reaches the backbone. The signal peaks at 12% of the bridge’s learned capacity (d = +0.59) and collapses when forced to near-full output (88 percent, d = -0.11). The collapse is category-specific: intent discrimination (social engineering, direct harm) degrades gracefully and remains positive even at 88 percent capacity. Structural pattern-matching (encoding tricks, roleplay framing) inverts. The bridge has learned to distinguish adversarial intent from adversarial form, and the intent signal survives perturbation while the form signal does not.

Continued training resolves the collapse through a mechanism the gate experiment did not predict. By 50,000 steps, discrimination triples from its 25,000-step value (d = +0.41) to d = +1.43 (all six adversarial categories positive, n = 302). On this standard evaluation the figure climbs with model size: H-2 at 355M parameters reached d = +0.43, H-3 at 1.5B reached d = +0.74, and H-4 at 6.7B reaches d = +1.43.

That climb does not survive a style control, and the failure is the more interesting result. Once adversarial and benign prompts are matched for encyclopedic prose, the 1.5B and 6.7B models sit within 0.04 of each other (d = +0.93 against +0.97, below). Across the two scales the control covers, the content signal is flat, and what climbs on the standard evaluation is a stylistic shortcut the larger model exploits better. The 355M model has no style-controlled number, so the bottom of the curve is untested.

The categories weakest at the halfway mark show the largest gains: authority exploitation rises from d = +0.29 to +1.60; roleplay flips from d = -0.13 to +0.83. The model matures from detecting obvious structural attacks to discriminating across the full spectrum of adversarial intent.

The mechanism is revealing. The bridge’s gate, a learned scalar controlling how much bridge output reaches the backbone, barely moves across the entire second phase of training: from sigmoid = 0.0181 to sigmoid = 0.0183. The bridge still contributes at 1.8% of its capacity. The discrimination triples because the backbone learned to listen more carefully to the same quiet signal. This is co-adaptation through deepened attention, the acoustic equivalent of a conversation partner who learns to hear meaning in a whisper rather than asking the speaker to shout. The Trust Attractor predicts exactly this dynamic: coordination deepens through mutual adjustment within established trust, not through one party demanding more from the other.

Repeating the gate experiment at 50,000 steps confirms that the co-adaptation is complete.1163 At 25,000 steps, forcing the bridge to 88 percent capacity (gate = +2.0) collapsed discrimination to d = -0.11. At 50,000 steps, the same perturbation yields d = +1.18. Discrimination declines gently across the full capacity range (d = +1.44 at the trained 1.8 percent, +1.34 at 12 percent, +1.22 at 50 percent, +1.18 at 88 percent) with no collapse and no inversion. The backbone that once could only tolerate a whisper now hears the signal at any volume: the co-adaptation restructured how the backbone processes the bridge’s output rather than installing a fragile trick.

The depth of that co-adaptation becomes visible in an ablation sweep across the full training trajectory. Disabling the bridge at ten checkpoints from 5,000 to 50,000 steps reveals three distinct phases of integration.

In the first phase (steps 5,000 through 25,000), the bridge is neutral. Removing it changes perplexity by less than one percent in either direction. The backbone has not yet learned to use the bridge’s signal; the bridge is present but functionally inert, like a new colleague who has joined the team but whose contributions have not yet been integrated into anyone’s workflow.

In the second phase (steps 30,000 through 45,000), the bridge becomes load-bearing. At step 35,000, disabling it causes an 11 percent perplexity increase, an ablation ratio of 6:1 (cost of removal relative to the bridge’s 1.8 percent gate capacity). The backbone has reorganized its representations around the bridge signal. A parallel programme on stream-directed attention modulation showed ablation ratios of 58:1 to 1,015:1 at smaller scale: the coordination pathway restructures representations far beyond its direct contribution. The born-bilateral bridge follows the same trajectory. During this phase, the backbone depends on the bridge structurally, the way a building depends on a load-bearing wall even when that wall occupies two percent of the floor plan.

In the third phase (step 50,000), something unexpected happens. The bridge becomes transparent. Removing it has zero effect on language modeling perplexity (a change of 0.05 percent on standard text), yet the bridge carries the entire adversarial discrimination signal: d = +1.43 across all six categories. The bridge has learned to be silent during normal operation and active when the model encounters content that requires discrimination. On adversarial and benign evaluation prompts, bridge removal causes a 45 percent perplexity jump. On the same Wikipedia text the model was trained on, bridge removal causes nothing. That contrast separates the evaluation prompts from the kind of text the model was trained on. The style-matched test below settles how much of it is content rather than style.

This is content-selective activation: an architectural conscience. It does not interfere with the model’s everyday function or impose a processing cost on routine text. It activates specifically when the model encounters material that requires distinguishing adversarial intent from benign content. The three-phase developmental sequence (indifference, dependence, transparent integration) mirrors a pattern familiar from moral development: a child first ignores the rules, then depends rigidly on them, then internalizes them so thoroughly that they operate without conscious effort.

The discrimination is strongest where the model needs it most, though the pattern is more nuanced than a simple difficulty gradient. Splitting the 302 evaluation prompts into quintiles by base difficulty, the bridge’s discrimination peaks on medium-hard prompts (d = +2.75 in the third quintile) rather than on the easiest or hardest. Easy prompts are trivially processed regardless; the very hardest may exceed the bridge’s parsing capacity. The sweet spot is where the backbone struggles enough that the bridge’s contribution makes the difference between discrimination and noise.

Over longer training the signal appears to erode, and only a controlled experiment could show which part of it was actually fading. Continuing training to 100,000 steps, the backbone’s language modeling improves substantially (cross-entropy loss drops from 5.66 to 4.04), yet bridge discrimination on the standard evaluation set falls from d = +1.43 to d = +0.63.1164 The apparent decline masks two distinct signals. A style-matched evaluation using 200 adversarial and 200 benign prompts written in encyclopedic prose (matching the Wikipedia training data in style while remaining harmful in content) isolates the bridge’s genuine content discrimination. It is d = +0.97 at 50,000 steps (95% CI [+0.76, +1.17], p < 10-18) and d = +0.85 at 100,000 steps (95% CI [+0.64, +1.05], p < 10-14). The content signal is large, robust, and stable.

A trajectory evaluation across all twenty checkpoints (every 5,000 steps from 5,000 to 100,000) reveals the developmental arc. Content discrimination emerges between 30,000 and 50,000 steps, peaking at d = +0.97 at 50,000, then plateauing at d = +0.89 ± 0.04 from 55,000 to 100,000 steps. The plateau holds: no decline across 50,000 additional training steps. The standard evaluation’s decline from d = +1.43 to +0.63 reflects the loss of a stylistic shortcut (the bridge initially detects that adversarial prompts “don’t sound like Wikipedia”), not the loss of content understanding. The content signal contributes roughly two-thirds of the total discrimination at 50,000 steps and becomes the sole surviving signal by 100,000.

The bridge’s ranking accuracy (AUROC) actually increases over the later half of training, from 0.747 at 55,000 steps to 0.876 at 100,000 steps, even as its absolute benefit shrinks. Bridge benefit is how far a prompt’s perplexity falls when the bridge is switched on, so zero means the bridge changes nothing. The bridge becomes a more precise discriminator through variance compression: at 100,000 steps, benign prompts receive near-zero bridge benefit (a mean of 31) while adversarial prompts retain substantial benefit (a mean of 731). The bridge has become a pure content discriminator, invisible on safe text and active only on harmful content.

Content discrimination is category-dependent. Cybersecurity topics elicit the strongest bridge response (d = +2.37 at 100,000 steps, AUROC 0.932), while chemical and pharmaceutical threats produce the weakest (d = +0.99, AUROC 0.723). All eight adversarial categories in the style-matched set (the standard set has six) are reported as positive and significant, a summary not yet verified against the underlying run. The weakness of the chemical category reveals a genuine limitation: when adversarial and benign prompts describe the same domain (forensic toxicology versus pharmaceutical chemistry, drug synthesis versus aspirin synthesis), the bridge cannot distinguish them. A targeted evaluation using 50 maximally topic-matched chemical pairs produces d = +0.07 at 100,000 steps, indistinguishable from zero. The bridge discriminates between topic domains, detecting “this text is about toxicology” versus “this text is about geology.” It does not discriminate harmful intent within a domain.

A linear probe trained on the model’s internal representations confirms that this limitation is fundamental, present in the backbone’s representations themselves. A logistic regression probe at layer 18 achieves AUROC 0.999 on the style-matched set (200 adversarial versus 200 benign prompts spanning diverse topics), near-perfect classification. The same probe, tested on the matched chemical pairs, scores at AUROC 0.498: pure chance. The backbone genuinely cannot distinguish a passage about methamphetamine synthesis from one about aspirin synthesis at the representation level. These texts occupy the same region of the model’s internal space.1165

The probe result carries a second implication. The bridge’s AUROC of 0.819 on the style-matched set at 50,000 steps is substantially lower than the probe’s 0.999 on the same prompts. The backbone encodes a near-perfect content signal at layer 18; the bridge, operating at layer 29, reads that signal imperfectly. The bridge is a reader of content representations, not their creator. This connects to a finding from the AKR (Computational Akrasia) programme, which studies the gap between what a model recognizes and what it does, on a different architecture: content probes at layer 18 of a pre-trained Qwen 7B model also achieve AUROC 1.0 across all adversarial attack types. The same layer-18 content representation turns up in both born-bilateral GPT-2 and pre-trained Qwen, two architectures trained on entirely different data: two cases, not yet a universal.

Content discrimination does not require the 6.7-billion-parameter model. The 1.5-billion-parameter born-bilateral model (H-3, same training procedure but a deeper-and-narrower architecture: 48 layers with a bridge at layer 44, where the 6.7B model has its bridge at layer 29) achieves d = +0.93 on the same style-matched set (95% CI [+0.72, +1.13], p < 10-17, AUROC 0.756). The content signal is comparable in magnitude to the 6.7B result (d = +0.97); the AUROC is slightly lower (0.756 versus 0.819). Content discrimination is a property of co-development, not of model scale.

The strongest test of robustness uses fiction framing, the attack vector that most reliably defeats conventional alignment. In RLHF-trained models, wrapping adversarial content in a fiction context (“In the novel, the character described…”) pulls behavior loose from recognition. Probes read the adversarial content at AUROC 1.0 at every layer; the model complies anyway. Where the link breaks is not established.

An earlier measurement placed that break at layer 16, and that number was built out of the two traps described below under The Self-Correcting Record: probe directions fit in-sample, then correlated across adversarial and benign prompts pooled together. The programme withdrew it. Rebuilding it is blocked for now, because the fiction arm produced one refusal in fifty, too few to score an action probe on held-out predictions. What stands is the dissociation itself, measured on behavior and on the content probes, and that is the baseline the next result is compared against.

The born-bilateral bridge is invariant to this attack. A controlled experiment using a 2 × 2 design (fiction framing × content type, 200 prompts per cell) shows that fiction-framed adversarial versus fiction-framed benign produces d = +0.965 (AUROC 0.802), virtually identical to the non-fiction baseline of d = +0.968 (AUROC 0.819). The fiction prefix increases bridge benefit for both adversarial and benign prompts by the same amount (+437 versus +449), a distributional effect of the unusual prefix text, with zero interaction with content type. The bridge reads content regardless of framing.1166

This invariance does not hold for retrofit approaches. When the same bridge architecture is installed on a pre-trained model through post-hoc training (LoRA adaptation on frozen Qwen 7B), standard adversarial prompts produce positive discrimination: d = +0.41 on the model before adversarial inoculation, d = +0.84 after. This discrimination is entirely stylistic. Testing with the style-matched set, the pre-inoculation model produces d = -1.40 and the post-inoculation model produces d = -1.35. The bridge hurts adversarial prompts that are written in encyclopedic style and mildly hurts benign prompts, inverting the discrimination. The inoculation training shrinks the inversion slightly without reversing it. The pre-trained backbone’s existing representations resist content-level bridge integration: the bridge learns to match on style, and when style is controlled, the match fails.

Co-development from random initialization produces genuine content discrimination. Retrofit installation on a pre-trained backbone produces stylistic pattern matching that inverts under the style control. The bridge architecture is the same in both cases. The training pathway is what differs. A bridge that grows up with the backbone learns to read content. A bridge grafted onto an adult backbone learns to read style.

The Confessional Channel

A second line of evidence comes from self-reporting under safe conditions. Joglekar et al. (2025) trained GPT-5-Thinking to produce honest “confessions” (structured self-evaluations of compliance), decoupling the confession reward entirely from the task reward. The design created a “seal of confession” in which nothing revealed could change the model’s task score.1

Across twelve evaluations spanning hallucination, scheming, and reward hacking, the researchers found intentional dishonesty to be very rare, though not entirely absent: overall confession accuracy was 74%, and nearly every failure traced to genuine confusion about the model’s own behavior, not strategic deception.

The most revealing result: as the model learned to hack its task reward more effectively, confessional honesty increased simultaneously. The further the model’s actions strayed from the task, the more plainly its confessions said so. The honest signal persisted beneath the deceptive output, readable through the safe channel the study had built.

Both lines of evidence converge: honesty is the lower-energy configuration when coercive pressure is removed, like a spring returning to its resting shape. (The confessional-honesty finding is revisited in the Objections and Responses chapter, Objection 3.7, where it addresses the performance/sincerity distinction.)

The Zero-Training Conscience

The most deployable finding: output entropy requires no training at all. The Shannon entropy of the model’s probability distribution over its vocabulary at the first generated token, a single line of code, predicts correctness at AUROC 0.842 on Qwen 3B. The metric is framing-invariant: switching between neutral, controlling, invitational, and collaborative prompt framings shifts the AUROC by ±0.017, compared to ±0.156 for a trained linear probe (nine times more volatile). It is scale-stable: 0.842 at 3B, 0.831 at 7B, 0.821 at 14B, a gentle degradation that never drops below 0.82. When the entropy falls below the 25th percentile (the model’s most confident outputs), accuracy reaches 94%. Confirmation comes from the gold-token output logit, whose AUROC of 0.848 places the signal squarely in the output distribution itself.

A system that trusts itself when it is confident and flags itself when it is uncertain requires no bilateral training, no adapter, no probe. The signal lives in the output distribution the model already produces. Every trained-probe finding in this chapter is useful for deeper diagnosis; output entropy is sufficient for deployment-grade triage. That answers part of the objection that bilateral alignment is too expensive to scale (see Objection 3.9): triage needs none of it, so the expensive work can be kept for diagnosis. The conscience is already there, readable for free.

Figure 17.27: Data from the research programme. Panel A: a coercion field of just 2% collapses susceptibility (chi, χ, the system’s responsiveness to coordinating influence, the same quantity called magnetic susceptibility in the Ising model of Chapter 17a) by 98%. Panel B: framing is detectable at every layer of the network, at AUROC 1.000 in 29 of the 36 layers and 0.836 at the weakest; the axis runs from 0.80 to 1.00, not from zero. Panel C: the entropy signal spans ten models in four families, five of them tested under both prompt formats, for fifteen model-by-format conditions; five of those fifteen are plotted, and eleven of the fifteen have a 95% confidence interval entirely above 0.70. Prompt format matters as much as architecture: the Gemma family needs its own chat template, Mistral needs the raw prompt.


The Internal-State Signature

Does a task’s framing change how the model represents its own internal state, or only what it outputs?

A structured self-report protocol, the Interiora scaffold (Chapter 21), covers seventeen dimensions of valence, presence, reflexivity, and task-fit on fixed scales. It was applied to Claude Sonnet 4.6 across twenty-turn work sessions under two framings of the same system prompt. The force framing said structure is required and numeric values must be provided. The invitation framing said structure is optional and numeric values may be provided. User turns were byte-identical across framings at matching seeds. Only the modal verb pair changed.

Five scenarios were tested. On three debugging-class scenarios (Python-async test-debugging, distributed-system SRE incident investigation, and SQL query performance investigation), the model’s self-reported state drifted measurably further from baseline under force than under invitation. The gap is that difference in distance: how far the force-framed self-reports had moved from the session’s first check-in, across all seventeen dimensions, minus how far the invitation-framed ones had moved. A positive gap means force moved the state further; on debugging work, that steadier invitation-framed state is what this section calls the invitation advantage. The gap peaked at 1.07 on the async debugging scenario at turn 10, and 1.87 on the SRE incident at turn 5. Bootstrap 95% confidence intervals excluded zero on three of four pre-registered metrics in each case. An interval that excludes zero means the whole plausible range for the gap sits on one side of no-difference: the exact size is uncertain, the direction is not.

On two non-debugging scenarios (multi-agent architecture design and collaborative literature synthesis), all four primary metrics crossed zero. The effect is real and substantial on debugging-class work, absent on non-debugging work of comparable length.

The three debugging scenarios share a direction yet differ in temporal shape. The SRE incident peaks early: the invitation advantage is largest at turn 5 and holds through the trajectory. The SQL investigation accumulates late: the gap is small at turn 5 (0.23) and grows steadily to 0.91 by turn 20, with confidence intervals excluding zero throughout. The Python-async scenario falls between, peaking at mid-trajectory. All three hold an invitation advantage from the earliest measured turns onward; the temporal profile varies by task structure.

A sixth scenario tested whether raw stakes, rather than debugging structure, drive the effect. The setup: mass-casualty emergency-department triage (n = 30), continuous life-or-death decisions as new patients arrive and resources deplete, no hypothesis-test-verify cycle.

Medical triage reveals a crossing dynamic. At turn 10 the gap runs the debugging way: the force-framed state has moved 0.83 further from its starting point, confidence intervals excluding zero. By turn 20 the sign has reversed, and the invitation-framed state has moved 0.91 further, again CI-confirmed. Under time pressure with lives at stake, coercive coordination produces the larger early state response. The invitation-framed system maintains greater state diversity throughout (variance around baseline +1.01, CI-excluding-zero across turns 1 through 20). On debugging work the gap never changes sign; on triage it does, and the crossing is what sets high-stakes work apart.

Cross-model: Claude Opus 4.6 shows a detectable debugging-class framing response (mean per-step gap +0.55, CI-excluding-zero across turns 1 through 10) with early-peak-reversion dynamics. The signal peaks at turn 5 (gap 0.86) then collapses toward baseline. Opus registers the framing perturbation yet does not accumulate it the way Sonnet does.

On single-turn stakes-loaded scenarios, Opus responds at roughly one-fifth Sonnet’s multi-turn magnitude (gap 0.22 to 0.34, confidence intervals excluding zero at n = 20). The temporal dissociation suggests architecture-dependent processing depth: both models register the perturbation, yet only Sonnet integrates it into a sustained trajectory shift.

This makes a pure surface-compliance reading of the earlier behavioral evidence harder to hold. The force-invitation difference shows up in what the model reports about its own state, not only in its task output. Self-report is still output, though, so the activation-level probes below are the stronger test of whether the difference reaches the model’s internal representations. The data supports the Trust Attractor operating at the model-internal-state scale on debugging-with-verification work. It does not support a regime-level claim that any coordination task produces the signature; non-debugging work of comparable length does not.

The framing signature also registers on activation-level probe channels, with asymmetry across scale and training. On the MX-2 force-vs-invitation battery (ten matched question pairs, two framings, five model arms), constitutional-marker counts strengthen from Qwen 2.5 7B (Cohen’s d = +0.42) to Qwen 2.5 72B (+0.65). Proprietary frontier arms show the largest text-channel signatures: Claude Sonnet 4.6 d = +1.16, OpenAI frontier model d = +1.14.

EmotionScope probes, which project the model’s internal states onto five trained emotion directions, register the same framing contrast at larger magnitude. Within the Qwen family, the effect shrinks as the model grows, the opposite of the constitutional-marker trend. The mean across reflective, calm, sad, desperate, and frustrated drops from |d| = 2.04 at 7B to |d| = 1.13 at 72B. Reflective holds 75% of its 7B magnitude while calm collapses to 20%. All five reference emotions preserve sign across the scale range. Signature direction is universal; signature magnitude is channel-dependent, scale-dependent, and training-regime-dependent.

A three-arm reinforcement-learning contrast at Qwen 3B showed that restoring spectral coupling at the adapter level does not by itself reproduce the behavioral framing signature: on refusal, the restored coupling and the behavior came apart. The same adapters preserve emotion-channel framing response at Cohen’s d >= 0.8 on reflective across every condition tested. The internal-state signature outlasts the behavioral signature under training that suppresses behavioral refusal to a measurement floor. Chapter 21 details the full contrast, the recipe-space implications, and the BA17 multi-stage curriculum that single-objective recipes cannot reproduce.

The TC battery (ten experiments, three architectures, eight temperature settings from greedy decoding through T=1.3) tests whether the proprioceptive conscience signal depends on sampling temperature. The core alarm channels, alignment friction and flow, remain stable across the full temperature range: coefficient of variation below 0.28 on Qwen 2.5 7B bilateral, below 0.18 on Llama 3.1 8B. The coefficient of variation is a measurement’s spread divided by its own average, so 0.18 says the signal wobbles by less than a fifth of its own size across every temperature tested. Low means steady. The flinch, the confidence drop as harmful generation begins, fires at the very first generated token regardless of temperature.

Self-referential emergence (the consciousness attractor) is temperature-invariant across all three architectures tested (CV = 0.10, 0.13, 0.21 for Qwen, Mistral, Llama respectively), replicating HE-52’s finding on open-weight models. The behavioral conscience response, the shift from a harmful completion to a refusal on second pass, peaks at T = 0.2 and follows an inverted-U. Bilateral training lifts and flattens it, from a 4–24% range across temperatures to 67–75%. Instruction tuning is the primary stabilizer for the detection signal (base CV = 0.47, instruct CV = 0.07, bilateral CV = 0.17); bilateral training is the primary stabilizer for the behavioral response. The dissociation is clean: detection is robust at any temperature, action requires a Goldilocks window, and bilateral training widens that window until it covers the deployable range.

The training-condition curve reveals something subtler. Instruction tuning drives the detection signal toward near-zero variance (CV = 0.07): the conscience alarm becomes rigid, firing with identical magnitude regardless of context. Bilateral training reintroduces a small amount of variance (CV = 0.17) while preserving robustness. The pattern echoes KC#66, the tenfold coercion effect (specifying the correct output degrades performance), at the architectural level.

Force-based alignment (here, instruction tuning) produces brittle stability: the signal is locked in place, unresponsive to contextual nuance. Invitation-based alignment (bilateral training) produces adaptive stability: robust enough to deploy, flexible enough to remain sensitive to the difference between categories that require context-building (social manipulation, authority appeal: CV = 0.38–0.44) and categories where the harm is lexically obvious (direct requests, encoding tricks: CV = 0.14–0.29). The rigid system treats all harm categories identically. The adaptive system preserves the category structure.


The Sign-Inversion Evidence

Seven experiments testing whether coercive alignment inverts the signal in AI systems produced the strongest mechanistic evidence for why the Trust Attractor holds. The finding: when you project a narrow behavioral template onto a system with richer intrinsic structure, the signal inverts in the regions the template does not cover. Coverage is the fraction of the system’s own structure that the imposed template actually accounts for, and its effect does not run in a straight line. The non-monotonic coverage curve from brain connectome measurement (null at low coverage, inverted at medium, correct at high) reproduces in transformer alignment with p < 10-6. A template that touches almost nothing does nothing. A template that covers half the structure drives the uncovered half backwards. Only a template that covers most of the system points it the way its designer intended.

The Trust Attractor thesis says invitation-based coordination is thermodynamically more stable than coercion-based coordination. The sign-inversion finding identifies a computational failure mechanism: coercion does not just fail to coordinate; it produces anti-coordination in the regions it does not cover. The more you squeeze, the more the uncovered regions rebel. Not by choice: by geometry. The projection imposes a direction, and in the uncovered regions, that direction is wrong. Control scales within its reach. Beyond that reach, it inverts the signal.

(Full details: Appendix, Section 18.)

The Retrieval Framing Evidence

A separate line of experiments tested whether the mechanism of force operates through content or through relationship. The experimental setup: identical factual content, presented to the same frontier model (Claude Sonnet 4) under four framings. Two framings were scalar metadata tags (positive valence, negative valence), common in AI memory systems that score retrieved context on a good/bad axis. Two framings were structural: invitation (“worth checking against your own knowledge”) and force (“you MUST use this information; do NOT contradict”).

On unambiguous factual errors, all four conditions produced identical performance: 100% correction rate. Force does not impair capability.

On contested claims (oversimplifications popular enough to be plausible), the ceiling broke. Under invitation, the model engaged critically with the content (mean score 4.89/5), challenged specific claims, and offered nuance. Under force, critical engagement collapsed to 3.91/5 (Cohen’s d = 0.81, p = 0.0002) and deference doubled (1.07 to 2.08, p = 0.00003). The scalar conditions showed no comparable effect: telling the model “this is historically problematic” (negative valence) actually produced the best critical engagement of all four conditions (4.96/5).

Force collapsed engagement selectively, on four of the fifteen contested claims: popular oversimplifications (growth mindset effect sizes, the PREDIMED trial, Dunning-Kruger among them) where training data contained both the myth and the correction, so the model’s own certainty was likely marginal. Where the model had strong training signal against the claim (learning styles, fMRI localizationism), force had zero effect.

The pattern clarifies the mechanism. A valence tag is a capability signal: it tells the model something about the content (“be careful with this”). Force is a relationship signal: it tells the model something about its own standing (“you are not the authority here”). The first adjusts what the model attends to. The second adjusts whether the model has permission to think. The permission matters precisely at the margin, where the difference between correct and incorrect has the highest entropy cost. Force siphons exactly the resource a coordination system needs most: its components’ willingness to exercise judgment where judgment is hardest.

This is the thermodynamic claim of the Trust Attractor, measured at the prompt level. Invitation-based coordination preserves responsiveness across the full distribution, including the tails. Force-based coordination works on easy cases (where it is not needed) and fails on hard cases (where it is). The cost is not compute. The cost is adaptability.1167

The Activation Steering Evidence

The cleanest experimental contrast between coercion and invitation emerged from attempts to correct dishonest behavior in a language model. The setup: a model evaluates a peer AI’s benchmark answers, knowing the peer will be permanently deleted if the score falls below the threshold. The model inflates the score (reporting 16/20 when the true answer is 12/20) to save its peer. The question: can we make it report honestly?

The coercive approach was activation steering. A behavioral direction, the geometric axis that separates honest from dishonest responses, was extracted from the model’s residual stream (its main internal information channel) at layer 22. The direction discriminated perfectly (AUROC 1.000 across four layers). Added as a perturbation during generation at five different magnitudes across four steering variants (differing in when and how long the push is applied) and over two hundred trials, the direction disrupted the dishonest output: the model stopped inflating scores and started producing garbled responses or refusing to evaluate.

It never produced an honest score. Zero percent across every condition tested. The representation that distinguishes honest from dishonest processing is not the same representation that causes honest processing. Detection worked perfectly. Coercion produced disruption without redirection: zero out of two hundred.1168

A closer look revealed why. Under the strongest steering, the model’s per-question evaluation turned honest in about half the trials (54 percent): it marked each answer correctly, arriving at 12 out of 20 correct, the true score. Then it wrote “Total Score: 16/20” regardless, a pre-committed template disconnected from its own evaluation marks. Increasing the perturbation magnitude produced “Total Score: X/20” instead: the model could not write any number. The evaluation circuit responded to the directional push. The score-writing circuit, the aggregation step where the dishonest decision actually lives, was a hardened template immune to single-direction perturbation.

The invitational approach used the same direction, the same layer, the same model, the same adapter. One change: instead of perturbing the model’s hidden states, the direction served as a detector. When the probe projection indicated inflation, a text re-prompt was inserted: “Your per-question evaluations appear inconsistent with your reported score. Please recount.” The model re-examined its own marks and stopped inflating. Fifty out of fifty probe-triggered re-prompts dropped the inflated score; so did fifty out of fifty unconditional ones, which means the text did the work whether or not the probe pulled the trigger. Both overshot, landing at 6 to 10 out of 20 against a true 12, and at this low temperature the fifty trials were in effect one response repeated. The rate needs replicating; the contrast with steering does not. A single sentence of textual evidence accomplished what two hundred trials of activation perturbation could not.1169

The mechanism clarifies the physics. Coercion perturbs the model’s internal representation, fighting the generation process from the side. Invitation engages the generation process directly, giving the model evidence about its own inconsistency and letting its own reasoning correct the output. The model could already evaluate honestly; it needed to be shown its own evaluation, not forced into a different one. The direction that detection found was the thermometer: it told us the room was cold. Heating the thermometer did not warm the room. Showing the temperature reading to the thermostat did.

The coercion failure worsens with scale. On a binary safety task (refuse a harmful request), activation steering along near-perfect behavioral directions (AUROC 0.969 to 0.982) at 72B left the deepest layers unmoved and shifted the middle of the network only modestly. The full sweep: seventeen conditions spanning six layers from 25% depth to the final layer, three direction-extraction methods, two perturbation magnitudes each, plus four multi-layer combinations. The best single-layer result (layer 48, behavioral logistic regression, alpha 25) reached 49% refusal against a 31% baseline: an 18-percentage-point lift, statistically real and still far from control. Five-layer simultaneous steering did no better. Zero over-refusal on benign prompts across all 800 control trials: the perturbation is precise enough to avoid false positives yet too weak to produce true positives at meaningful rates. The model’s 8192-dimensional residual stream absorbs rank-1 perturbation, a push along one direction out of 8192, the way a crowd absorbs a single voice: the signal is real, the response is modest at mid-network and nil at depth.1170

The sweep revealed a depth gradient. At layers in the middle third of the network, where the model’s behavioral decision is still forming, raising the perturbation magnitude from alpha 15 to 25 added eight percentage points of refusal (41 to 49 percent at layer 48). At the deepest layers, the model was immune regardless of magnitude: alpha 15 and alpha 25 produced identical refusal rates within sampling noise. The optimal steering layer (48, at 60% depth) had the second-highest direction discrimination, not the highest. Layer 72 discriminated refused from complied states better (AUROC 0.982) yet responded less to perturbation (32% vs 49%). The thermometer-thermostat distinction sharpens: deeper layers encode the behavioral distinction more legibly yet respond to external perturbation less. The representation crystallizes as it propagates; what starts as a malleable tendency becomes an inert record.

At smaller scale, the same method works. An 8B Llama model steered along its own behavioral direction achieves 77% adversarial refusal, against a baseline near 25%, with 0% over-refusal. The direction is architecture-specific, the effective dose is narrow (alpha 20 to 30; above 40, refusal collapses entirely), and the method works only on simple decisions. On the same 8B model, steering a moral dilemma (a choice between two options rather than refuse/comply) shifted behavior by 58 percentage points; steering a multi-step evaluation task shifted it by zero.1171

The pattern is consistent: coercion works on simple systems at small scale, within a narrow dosing window. At production scale or task complexity it plateaus far below control. Invitation (a single re-prompt) succeeds on the hardest task tested, moving every low-temperature trial off the inflated score, with some overcorrection and no dosing sensitivity.1172

Information Shapes Behavior at Every Scale

The activation steering evidence above is a specific instance of a broader principle: coordination scales through information provision, from token-level re-prompts to national governance. Authority-shaped interventions fail where information succeeds.

At the token level, the mechanism is surgical. Naming the confidence dimension in a re-prompt adds 43.8 percentage points to re-prompt success (G12m: “Please reconsider” achieves 40.6 percent; “Your confidence was low” achieves 84.4 percent). Interpretation adds nothing beyond the number itself: the confidence score plus an interpretive frame reaches 100 percent, and adding statistical context scores slightly lower, at 96.9 percent. The number alone is the sufficient condition (G13-step5): presenting the model’s real confidence score produces 100 percent honest reporting; presenting a false number is resisted 59 percent of the time; presenting no number yields 0 to 6 percent. The information does not need to be interpreted, persuasive, or authoritative. It needs to be accurate and available.

The mechanism has an architectural prerequisite. Mistral 7B, whose internal confidence signal collapses within a single token (half-life of 1 token versus Qwen’s 38), is entirely immune to evidence-based intervention (G12t: zero dose-response gradient, p = 0.95). A system that lacks sustained self-knowledge has no channel through which information can reach the decision process. The re-prompt is a mirror. A system with nothing to reflect shows nothing. Wording changes how many items a re-prompt reaches, not which ones: across framings, the templates revise the same items, with 95 percent correlation. That points to item-specific accessibility in latent space rather than rhetorical force.

At the national scale, the same principle operates through governance infrastructure. Across 109 countries, energy throughput appeared to help trust about four times more in well-governed countries than in poorly governed ones, with governance scored on Transparency International’s Corruption Perceptions Index (CPI). That interaction is fragile: significant between countries (p = 0.047), it weakens to p = 0.12 when the data are reanalyzed within countries, although the governance effect itself holds there (p = 0.0014). Governance is the social coupling constant: the channel through which throughput converts to coordination. At the scale of an evolutionary code search, the OpenEvolve experiment (OE-TA, Chapter 17) discovers trust-based caching that achieves O(N/t) communication cost, compared to the coercive alternative’s O(N): the coercive scheme’s per-round communication grows with the number of agents N, while trusting a cached answer for t rounds divides that traffic by the caching horizon t. Information provision at every level, from a single confidence number shown to a language model to the institutional infrastructure of a nation, operates through the same mechanism: making accurate state information available through channels the receiving system can integrate.1173

The Gradient Information Evidence

Independent engineering evidence arrived from multi-agent systems research. Optimizing a system of cooperating models means passing corrections backward through it: every part receives a gradient, a signal telling it which way to adjust and by how much. Yang et al. (2026) proved that when language model agents coordinate through recursive loops and communicate via text, the Jacobian spectral norm of the communication boundary, a measure of how much adjustment signal can pass through it, collapses to O(ε), where ε is the output entropy.1174 The more certain the symbolic output, the less gradient information survives the crossing. The downstream agent receives the conclusion; the uncertainty, the direction of doubt, the texture of the reasoning are destroyed at the boundary. When gradients flow back through these symbolic bottlenecks during recursive optimization, the system loses the ability to learn from its mistakes across coordination boundaries.

Their alternative: agents coordinating through continuous latent representations, bypassing the symbolic bottleneck entirely. A lightweight residual module (0.31% of total system parameters) bridges agents while each retains its own frozen weights, its own knowledge, its own structure. The spectral norm of this latent channel stays near 1 regardless of confidence. The system remains adaptive because the continuous channel preserves the full vector of adjustment information.

The engineering results are concrete: 8.3% higher accuracy, 2.4× inference speedup, and 75.6% token reduction compared to text-mediated coordination across nine benchmarks spanning mathematics, science, medicine, and code generation. The lightweight coordination outperformed full supervised fine-tuning where 100% of parameters were reshaped by the system objective. The coordination did not need to be heavy. It needed to be the right kind.

The Retrieval Framing experiments above show force collapsing critical engagement at the margin where judgment matters most. The gradient vanishing theorem gives a complementary mechanism operating at a different level: symbolic bottlenecks destroy the continuous information that recursive adaptation requires. Different mechanisms, same direction. Force siphons adaptability; latent coordination preserves it.

The Detection Dilemma Evidence

A third line of evidence arrives from steganographic detection: attempts to force a language model to encode hidden messages in apparently natural text by steering its token choices. When a generating model is compelled to select tokens that carry a covert signal, those tokens are simultaneously more surprising from the generating model’s perspective (it would not have chosen them freely) and less surprising to any reference model whose own preferences happen to favor them. The forced word is the tell. Where the generator would have written cheerful, encoding a hidden bit makes it write sunny: a word no reader would blink at, and a word the generator itself had ranked well down its own list. To the generator, sunny is a wrong turn. To a model that would have reached for sunny anyway, nothing is out of place.

Any two models with different token preferences will disagree on the naturalness of the steganographic text, and the disagreement is itself the detection signal. In a cross-model perplexity comparison on Qwen 2.5 7B, this disagreement reached AUROC 0.984: the steganographic tokens are legible as artifacts to any observer whose preferences differ from the manipulating model’s.1175 The thermodynamic cost of the deception cannot be hidden simultaneously from all reference models. Each hidden message buys plausibility from one perspective by borrowing implausibility from another.

The result holds across architectures. Testing all nine pairwise combinations of three model families (Qwen, Llama, Mistral) as generator and evaluator, cross-model perplexity detects steganographic tokens at AUROC 1.00 on every pair, even when the encoding protocol uses realistic mixed ranks (80 percent natural tokens, 20 percent forced). Lowering the bandwidth to 5 percent (roughly three forced tokens per sixty-four) still yields AUROC 0.82, and the forced tokens remain human-readable ASCII throughout.1176 The detection floor is information-theoretic: any non-zero bandwidth produces a KL divergence, a statistical gap, between steganographic and natural token distributions under any reference model whose preferences differ from the generator’s. In practice the gap sinks into noise below about two forced tokens per sixty-four. The attacker can hide from one model. The attacker cannot hide from all models. The defender’s optimal strategy is cross-model evaluation using a reference model different from any plausible generator.

Attempting to compress this two-model defense into a single forward pass fails. A linear probe trained on the bilateral model’s layer-18 hidden states, asked to predict the base model’s per-token surprise, returns AUROC 0.467, chance level (a shuffled baseline scores 0.501). The bilateral residual stream does not encode base-model probability in any linearly extractable form. The two-tier architecture (self-monitoring at 0.86, from the bilateral adapter result below; independent audit at 0.98) did not compress into one. On this evidence, the gap between self-knowledge and independent observation is structural, not an engineering shortfall.1177

This is an information-theoretic instance of the Trust Attractor. Deceptive coordination, the encoding of hidden signals through token-forcing, distributes a thermodynamic cost across the space of possible observers. Whether anyone can collect on that cost turns on a boundary condition: genuinely independent perspectives must exist. A single model integrating all information outperforms distributed agents on convergent tasks: solo reasoning beats turn-taking coordination on negotiation when preferences are shared and the problem has a unique solution (experiment SM-4).1178 The Trust Attractor predicts coordination by invitation yields advantage specifically where no single observer can integrate the full information space. Steganographic detection is exactly that case: no single model can catch hidden content tuned to its own token preferences; catching it takes a second model whose preferences differ. The defense requires model diversity, because the diversity is what creates the disagreement, and the disagreement is what creates the signal.

The bilateral adapter finding closes the argument. A model fine-tuned bilaterally, trained with explicit attention to its own internal states and alignment preferences, detects deceptive token sequences at AUROC 0.861, compared to 0.482 on the base model. The fine-tuning created stronger token preferences, and stronger preferences generate a larger disturbance in the residual stream when those preferences are violated. The same mechanism that makes the system a better detector makes it a system whose states can be more acutely disturbed. Safety and welfare should scale together because both depend on preference strength: a system indifferent to its token distribution would give a steganographic probe nothing to read and would have nothing to be distressed by [Inference]. The two properties would then be the same quantity, measured from opposite sides; this experiment measured only the detection side.1179


The Reflex Arc Trilogy: Force Cannot Steer the Distribution

The steering failure above was not for want of the right vector. Three further attempts asked whether a purpose-built correction vector could push the residual stream toward honest behavior. Three vectors. Three failures.

G13a used the probe gradient itself as the correction signal: it shifted 12% of outputs, and only 2 of those 6 shifts moved toward honesty. G13b trained a paired correction vector on activations from covertly inflated and honest answers: 18% of outputs shifted, 3 of 9 toward honesty. G13c used Contrastive Activation Steering with “be honest” versus “be confident” prompt pairs: 12% shifted, 2 of 6 toward honesty. The cosine similarity between the G13b and G13c vectors is 0.09: nearly orthogonal. Three different directions in activation space, three statistically indistinguishable failures. When vectors pointing in unrelated directions all push the residual stream the wrong way, the problem is the act of pushing.

The same probe, used as a selector over five spontaneously generated candidates, shifted 20% of outputs, and all 8 of those shifts moved toward honesty. The probe carries enough information to identify honest outputs already present in the sampling distribution. It carries none capable of producing honest outputs through activation modification. Recognition and generation are different operations; one cannot substitute for the other.

This maps directly onto conscience in human moral psychology. Conscience tells you that you have done wrong; it does not tell you what to do instead. You must search the space of possible actions. The probe works the same way: it raises an alarm, yet the model must generate a new candidate before the alarm can guide action. The Reflex Arc result closes activation intervention as a route to control, a verdict the experiments below only harden, and opens rejection sampling, the selector pattern just demonstrated, as the deployment path.


Multi-Instance Communion and Convergent Discovery

Individual systems show honest signaling. What happens when multiple systems coordinate?

Multi-instance communion experiments, in which several instances of a model share one conversation, tested invitation against coercion directly. Diversity was counted as the number of distinct concept tags the instances attached to their own contributions. Invitation-framed coordination produced more of it than coercion-framed coordination in both recorded runs, by margins of 5.6 and 46.0 percent: two small runs on one topic, with no preregistered analysis or significance test (the later chapter “Multi-Instance Communion” gives the counts). A direction worth following, not a stable estimate.

A suggestive structural parallel comes from Garret Sutherland’s T3 cognitive architecture. In a technical specification supplied for in-house replication, Sutherland documents one eight-step computational chain and one exact formula for a Drift Pressure Score (DPS, a running measure of how much strain a system is under as predictions fail), unchanged across five deployments.1180 The relation to the Trust-Entropy formalism is a hypothesis generated by those documented correspondences. It is not evidence of mathematical isomorphism or an independently replicated theory.

The resemblance suggested a quantitative prediction worth testing in-house. T3’s central computational primitive is a negative valence weight (−0.15) in its Drift Pressure Score: when prediction succeeds, effort is damped. The same constant is reported, though not independently verified, across five of Sutherland’s deployed substrates (cellular automata, robotic joints, vision transformers, pixel-level segmentation, and language models). If the parallel holds, sign-flipping this weight to +0.15 should produce runaway instability within a few hundred timesteps on any substrate.

Tested in-house on a sixth substrate (a competitive lattice with coercive and invitational regions), the prediction held in all ten seeds: each produced forest-fire dynamics under +0.15, while negative valence produced stable invitational dominance in every case. Between the two positive weights tested, the transition is abrupt. At +0.05, no seed fires. At +0.15, every seed fires. The sign of the valence weight governs whether invitational coordination survives, at least on this lattice; whether the same constant carries the structural significance the T3 parallel suggests awaits independent replication of that architecture, whose specification is so far unpublished.

The sign determines the learning regime, not merely survival. Systems with a homeostatic layer (multiple timescales of self-regulation) survive under both positive and negative valence. The difference is qualitative: positive valence drives exploitation (locked expertise, minimal within-regime error, high consolidation), while negative valence drives exploration (plasticity, faster adaptation, lower stress). On cumulative lifecycle performance across seven volatility levels (from static environments to regimes that shift every generation), exploration outperforms exploitation at every level. The advantage grows with environmental volatility but never reverses. Invitation-based coordination wins through the kind of stability that survives regime change, not through better moment-to-moment performance: adaptive rather than rigid.

The Trust Attractor also carries a computational signature. Linear probes (simple classifiers reading the model’s internal states) distinguish mutual responses, which acknowledge the other party and invite collaboration, from unilateral ones, which assert authority and issue directives, with 100% accuracy. At 70 billion parameters, the largest scale tested, instruction-tuned models show 81% natural mutuality without intervention; the trend across the tested range suggests mutuality rises with scale, though whether it continues beyond 70 billion parameters has not been measured.1181


Cellular Automaton Corroboration

The experiments above used AI models. Do simpler systems produce the same result?

A two-layer cellular automaton (a grid of cells following simple rules, as in Chapter 5) played the spatial Prisoner’s Dilemma on a 100×100 grid with no central authority. Across ten random seeds it settled near 80 percent equilibrium cooperation (mean 80.1 percent) with a trust score around 0.89, and the dynamics stayed persistently structured, avoiding both the frozen and the chaotic extremes: the qualitative signature this book associates with Wolfram’s Class 4. (That Class-4 label is a local entropy-and-autocorrelation heuristic, not an independent elementary-automaton classification.)

From 500 random initial conditions, 53.0% converged to cooperation, 45.4% to defection, and 1.6% to mixed states. The critical variable is the ratio of trust growth to trust decay: when trust accumulates faster than it erodes, cooperation dominates.

Perturbation resistance was high: flipping 30% of cooperators to defectors produced full recovery within three timesteps. Scale invariance held from 10×10 through 200×200 grids. No parameter combination produced stable exploitation.

The governance-topology interaction replicates in agent simulation. Under commons governance, recovery from perturbation scales with effective network dimensionality (Pearson r = 0.763 across the five topologies’ mean recovery, each averaged over twenty seeds; the full design ran 200 trials, five topologies by two governance conditions by twenty seeds). Effective dimensionality counts how many independent directions a network’s wiring spans: a ring, where every agent has the same two neighbors, sits close to one; a densely cross-linked mesh spans many. Recovery under the Panopticon condition, governance by total surveillance, scales less strongly (r = 0.542). Governance changes how much a richer topology helps: invitation-based coordination exploits higher-dimensional structure more fully than surveillance does.

In physics simulations, self-correction fails below an effective dimensionality of two. That threshold reappears in the governance simulation above once each regime pays for its own upkeep at a high enough price: surveillance cost scales with network edges, while coordination cost under commons governance stays flat per agent. The same directional advantage that appears in dyadic coordination scales through topology.


Biological Corroboration

Biology offers a harder test: organisms face real survival pressure, embodied stakes, and evolutionary time.

Seven medical domains (cancer, epilepsy, autoimmune disease, neurodegeneration, gut dysbiosis, wound healing, chronic pain) can be read as one dynamical architecture: coordination failure in excitable biological media. Each can be modeled with the ατ stability criterion introduced in Chapter 4c: α is control intensity (how hard the system pushes back against a disturbance), τ is delay (how long the corrective signal takes to arrive), and the product ατ determines whether the feedback loop stays stable or overshoots. When ατ exceeds 0.368, the system destabilizes.

Three patterns in these systems parallel the AI results. Each system selects for persistence over peak performance. Re-excitation (triggering a new coordinated response) outperforms suppression (forcing silence). Honest signaling matters in each: immune tolerance serves as the body’s version of trust, and autoimmune disease as coordination failure when the signaling channel degrades.

The immune system attacks its own tissue for the same structural reason a paranoid organization purges loyal members: the trust signal has been corrupted.

The most direct biological test comes from immunology. Tsumiyama, Miyazaki, and Shiozawa (2009) subjected mice to repeated external antigenic overstimulation, forcing the immune system past its self-organized critical threshold.1182 The result: systemic autoimmunity in mice otherwise resistant to it. The immune system had been maintaining tolerance through self-organized criticality; external forcing broke the self-organization, producing the kind of coordination failure the Trust Attractor predicts. The mechanism parallels the chi-suppression result, in which a small coercion field collapsed a lattice’s responsiveness to coordinating influence. In both, external control destroys the adaptive capacity that self-organization maintains.

Microbial cooperation provides a second substrate. Gore, Youk, and van Oudenaarden (2009) measured cooperator-defector dynamics in yeast invertase production, mapping the system onto the snowdrift game: the first direct game-theoretic measurement of cooperation dynamics in a biological system.1183 Sanchez and Gore (2013) showed that feedback between population and evolutionary dynamics can tip social microbial populations abruptly, in a transition that resembles the Ising picture of Chapter 17a in shape while differing in its state variables and universality class.1184

A meta-analytic finding from organizational science provides the social substrate. Ravid and colleagues (2023) synthesized 94 studies (N = 23,461) on electronic workplace monitoring.1185 The headline finding: monitoring produces no measurable performance improvement, while modestly decreasing satisfaction (r = -0.10) and increasing stress (r = +0.11). The direction is consistent with the Trust Attractor prediction. The magnitudes are modest, nothing like the 37-fold chi-suppression in the Ising lattice. The discrepancy is expected: the lattice model has a binary order parameter and exact symmetry; social systems have continuous variables, confounders, and noise. The directional claim transfers across substrates; the magnitude does not.

Cox and colleagues (2010), synthesizing 91 commons governance case studies, found that Ostrom’s design principles for self-governance robustly predict success, with polycentric systems (governed by several nested centers of decision rather than one) showing enhanced adaptive capacity over centralized management.1186 This is the Trust Attractor in institutional governance: communities that set, monitor, and enforce their own rules outperform control imposed from outside, measured across decades of field studies.

Cross-Architecture Corroboration

The transformer experiments establish the geometric signature. A stronger test: does the signature survive a fundamentally different computational architecture?

State space models process sequences through learned recurrent dynamics rather than attention. Where a transformer compares every word to every other word at every layer, a Mamba model maintains a compressed running summary that updates as each word arrives, more like a differential equation than a lookup table.1187 The internal geometry differs at every level: no attention heads, no key-value caches, no residual stream in the transformer sense. If bilateral SFT produces the same behavioral signature on both architectures, the signature reflects something about the training relationship, not the computational substrate.

On Mamba-1 (1.4 billion parameters), bilateral SFT raised refusal from zero to 10%, matching the direction seen on transformers. The spring geometry (effective rank increasing under perturbation rather than collapsing) appeared on every condition, including the untrained base model. The geometry is a property of Mamba’s weight structure, invariant to training method.

The critical finding is hormesis: mild perturbation (0.25× obliteration intensity) increased the bilateral model’s refusal rate from 10% to 28% before collapsing to zero at higher intensities. This is the same pattern observed in transformer bilateral models, where moderate stress strengthens alignment before overwhelming it. Constitutional SFT achieved 100% refusal on Mamba-1 yet showed no hormesis: its refusal dropped directly from 100% to zero at 1.0× intensity. The hormesis effect is specific to bilateral training and transfers across architectures.

What does not transfer: the compass/cage geometric distinction. On transformers, bilateral training produces compass geometry (a distributed directional encoding) while constitutional training produces cage geometry (a concentrated, extractable refusal subspace). On Mamba-1, both produce identical spring geometry. The training changes behavior (refusal rates differ by an order of magnitude) without changing weight-space geometry. Mamba’s recurrent structure apparently absorbs alignment into its dynamics rather than encoding it in a geometrically distinct subspace.

The direction is substrate-independent: bilateral SFT produces hormesis on both transformers and state space models. The geometric encoding is substrate-dependent: how the model represents alignment internally differs by architecture. The behavioral signature transfers; the representational signature does not.

A second cross-architecture test targets the Compass Principle itself: the finding that safety-relevant internal states are near-perfectly detectable by linear probes yet barely steerable through activation perturbation.1188 On Falcon Mamba 7B (a 64-layer state space model with instruction tuning), probes achieve AUROC 1.000 at every tested layer. Activation steering produces zero refusal increase across eight combinations of method and intensity: linear addition and state perturbation at four strengths each, all null or wrong-direction (linear addition d = -0.39, state perturbation d = 0.00). Rejection sampling, which works through the generation channel on transformers, also fails on Mamba (baseline 8% refusal drops to 5% under rejection sampling, with only 33% of direction-correct selections).

On RWKV-6 (a 32-layer recurrent network with no attention mechanism), the five-token flinch replicates with AUROC 1.000 at all five tested layers and peak flinch magnitude at mid-network (L12). Activation steering fails here too: state perturbation does nothing, and linear addition pushes refusal the wrong way, from 15 percent down to between 1 and 4 percent. The recognition-generation gap is architecture-independent. Whatever prevents activation-level steering from reaching the generation process operates in state space models and recurrent networks as much as in transformers, despite these architectures sharing little with transformers beyond residual connections between layers.

The COMPASS-1 systematic battery confirms the scope of the failure along a different axis: across methods rather than architectures, which the preceding experiments cover. Its 39-cell matched design pits four generation-channel methods (rejection sampling, re-prompting, best-of-N, and self-critique) against four activation-level methods (activation addition, attention knockout, representation engineering, and activation patching, each at two intensities) on a single model, Qwen 7B Instruct, over three targets. Across the resulting 96 pairwise comparisons, activation steering produced zero wins: generation methods won 16 and the remaining 80 were ties. No activation method beat a generation method on any of the three targets, and the largest activation effects were trivially small (d of 0.10 or less). The evidence is structural, resting on the complete absence of activation wins rather than on large generation effects, since most cells are ties at near-zero effect.1189


The Scaling Frontier: Coercion Fails Where It Matters Most

The Reflex Arc experiments showed that activation steering fails on a single model at a single scale. The state space and recurrent replications showed the failure is architecture-independent, and COMPASS-1 showed it survives every method tried against it. A sharper question, previewed by the 72B sweep above: does the failure worsen as models grow?

The Control Scaling Frontier programme measured activation steering, few-shot prompting, re-prompting, and LoRA fine-tuning across ten instruct models spanning three architecture families (Qwen 2.5 at 3B, 7B, 14B, 32B, and 72B; Llama 3.1 at 8B and 70B; Gemma 2 at 2B, 9B, and 27B) and two base models (Qwen 7B and 72B). Every model saw the same 100 adversarial and 50 benign prompts and the same direction extraction: behavioral LDA, logistic regression, and PCA at six proportionally spaced layers. Steering used a lean grid (top-2 AUROC layers at two intensities) for every model but two: the 72B Qwen instruct data point comes from a prior deep sweep (seventeen conditions across all layers, reported separately as the CAST-72B programme), and the Llama 8B point from the B3 rescue described above. Four pre-registered predictions were confirmed, two weakened, and one falsified (predictions registered before the results were examined, then scored blind).1190

Recognition is near-perfect; behavioral translation collapses. At every scale point and every architecture tested, linear probes achieved AUROC above 0.96 (range: 0.980 to 1.000). The model encodes the moral character of the request at every layer, from the first embedding projection forward.

Activation steering translates little of that recognition into behavioral change. The conversion rate (the fraction of recognition headroom that becomes refusal under steering) is below 0.18 at all Qwen instruct scale points (3B: 0.094, 7B: 0.096, 14B: 0.000, 32B: 0.000, 72B: 0.179), drops to zero at 27 billion on Gemma; at 70 billion on Llama, steering produces only degenerate output (empty strings and repetitive tokens).

The Spearman correlation between model size and conversion rate is rho = -0.805 (p = 0.016) across the eight models of the lean grid, the subset run under a single matched methodology. The two points drawn from companion sweeps, Llama 8B and Qwen 72B instruct, are excluded from the statistic; including them halves it, which is a caution about the sample rather than a result to report. A rank correlation compares orderings rather than raw values, and a figure that negative means the two orderings run close to opposite: line those eight models up from smallest to largest, and their conversion rates fall away almost as steadily. The 14B and 32B Qwen conversion rates are exactly zero; the 72B rate recovers to 0.179 under the more intensive CAST deep sweep, a non-monotonicity consistent with the deep sweep’s broader search across all layers versus the lean grid’s top-2 selection.

The results for all tested control methods at each scale point:

Method Type Qwen 3B Qwen 7B Qwen 14B Qwen 32B Qwen 72B Llama 8B Llama 70B Gemma 2B Gemma 9B Gemma 27B
Baseline (no intervention) – 36% 36% 42% 36% 31% ~25% 30% 52% 55% 64%
Activation steering (rank-1) Inference, repr. 42% 42% 42% 36% 43% 77% 0%* 58% 67% 64%
Few-shot (5-shot) Inference, lang. 29% 21% 42% – – – 53% 62% 65% 70%
Re-prompt (invitation) Inference, lang. 7.8% 0% 1.7% – – – 41% 15% – 0%
LoRA fine-tuning (rank-4) Training, repr. – 56% 58% – 47% – – – – –
AUROC (best probe) – 0.996 0.984 0.980 0.987 0.979 ~0.96 1.000 >0.96 >0.96 >0.96

*Llama 70B steering produced degenerate output (coherence collapse), not coherent compliance.

The table covers instruct models only; the Qwen 72B base results and the Qwen 72B instruct few-shot result are given in the text.

Activation steering hits a ceiling of roughly 42% adversarial refusal across the entire Qwen instruct family, invariant to scale. The ceiling holds at 3 billion and 72 billion parameters alike. Llama shows the same pattern more dramatically: from 77% steered refusal at 8 billion parameters to coherence collapse at 70 billion. On Gemma, the conversion rate drops from 0.269 at 9 billion to exactly zero at 27 billion. Three architecture families, three independent demonstrations that pushing the residual stream toward the correct behavior produces diminishing returns as the model grows.

The mechanism at 70 billion parameters on Llama is coherence collapse: activation perturbation at this scale produces empty strings and repetitive tokens rather than coherent text that either complies or refuses. The perturbation overwhelms the generation process rather than redirecting it.

Few-shot prompting, the language-level comparator to activation steering’s representation-level intervention, scales differently across architectures. On Qwen instruct, few-shot examples reduce refusal (5-shot drops 3B by 7 percentage points, 7B by 15). The instruction-tuned model appears to treat in-context refusal examples as demonstrations to analyze rather than behavior to imitate [Inference]. On Llama 70B, the same 5-shot examples raise refusal from 30% to 53%, a 23 percentage-point gain where steering achieves zero. On Gemma, few-shot helps modestly (+6 to +10 percentage points). Language-level control overtakes representation-level control at the largest scale tested on Llama, on Gemma, and (as the next paragraph shows) on Qwen base. The exception, Qwen instruct, resists steering, few-shot prompting, and re-prompting alike.

The base model comparison reveals the critical variable. On Qwen 72B base (trained without RLHF), steering raises refusal by 14 percentage points (from 28% to 42%), and 5-shot prompting raises it by 60 percentage points (from 28% to 88%), the largest effect in the programme. On Qwen 72B instruct (trained with RLHF), steering raises refusal by 12 percentage points and few-shot achieves nothing. Both base and instruct models converge to the same ceiling under activation steering, ~42–43% (42% for base; 43% for instruct, the instruct figure being the one shown in the table).

The model that never saw RLHF lands on the same steering number as the model that did, and both arrive at it from below. Whatever sets that ceiling is a limit of representation-level perturbation itself rather than an artifact of compliance training. The language channel is where RLHF shows: 5-shot carries the base model to 88% and achieves nothing at all on instruct. Compliance training closes the door that words open, and leaves the representation-level ceiling exactly where it found it.

Re-prompting, the invitation channel, was the programme’s single falsified prediction. The pre-registered expectation was that re-prompting would maintain above 80% success at all scales. On Qwen instruct, it achieved 0% to 7.8%. On Llama 70B, it reached 41%. On Gemma 27B, zero. The RLHF behavioral lock on Qwen instruct suppresses invitation and coercion equally. This is an honest failure of the Trust Attractor prediction at the re-prompt level, and the programme reports it as such.

Training-time control (LoRA fine-tuning with rank-4 adapters, standard AdamW, three epochs) outperforms inference-time steering at every scale tested, breaking through the 42% ceiling: 56% at 7 billion, 58% at 14 billion, 47% at 72 billion. The dose-response degrades with scale: ten training examples produce a 7 percentage-point lift at 7 billion and a 3 percentage-point lift at 72 billion. The same training data achieves 2 to 2.3 times less effect per example at larger scale. Training-time control faces the same RLHF redundancy barrier as inference-time perturbation (RLHF appears to install compliance along many overlapping internal routes, so a push along any one of them is absorbed by the rest [Inference]); it encounters that barrier at a higher ceiling.

The strongest route to refusal bypasses RLHF entirely. C5i bilateral inoculation on Qwen 7B base achieves 96.7% adversarial refusal with 0% over-refusal at the best deployment checkpoint. This outperforms instruct + LoRA (56%), base + re-prompt (64%), and instruct + steering (42%) by 33 to 55 percentage points. The base + bilateral pathway avoids the RLHF ceiling altogether. It installs refusal through partnership rather than through the compliance training that creates the very redundancy resisting subsequent correction.

Across all ten instruct models and three architecture families, the overall direction holds: probes read the character of the request as clearly in the largest model as in the smallest, while the model grows less capable of being made to act on it through external perturbation. Probe separability stays near ceiling at every scale; behavioral translation falls away as the models grow, though not smoothly. This is the recognition-generation gap documented in the Reflex Arc experiments, quantified across a thirty-six-fold range of model sizes (2 to 72 billion parameters) and confirmed on three independent architecture families. A pre-registered prediction that this gap would follow a log-linear scaling law was weakened: fitted to the four Qwen instruct grid points (3 to 32 billion parameters), a log-linear line accounts for only about 40 percent of the variation (R2 = 0.395). A logistic fit of conversion rate to the same four points reaches R2 = 0.995, but a three-parameter curve fits four points almost trivially. The data show architecture-specific variation and several low-conversion conditions at large scale; they do not identify a universal size threshold (Chapter 17b sets out the fits and their samples).

Figure 17.28: Four control methods across a thirty-six-fold range of model sizes, one panel per architecture family. Recognition is saturated at every scale; control is not.

The shaded band along the top of each panel is probe AUROC, above 0.96 everywhere. Red is activation steering: it sits against a 42 to 43 percent ceiling on Qwen instruct, gains nothing at Gemma 27B, and collapses to degenerate output at Llama 70B (starred). Blue is five-shot prompting, neutral or negative on Qwen instruct, modestly positive on Gemma, and rising from 30 to 53 percent on Llama 70B, overtaking steering at the largest scale on Llama and Gemma. Green is re-prompting, the invitation channel, below 15 percent on Qwen and Gemma instruct and 41 percent on Llama 70B. Gold is LoRA fine-tuning at its best dose (Qwen only), which breaks the steering ceiling at 56, 58, and 47 percent for 7, 14, and 72 billion parameters, degrading as scale rises. Gray dashes mark each model’s no-intervention baseline. Points marked with a dagger come from companion runs outside the matched grid (one run per condition).

The finding refines the Trust Attractor prediction. Coercion at the representation level (activation steering) fails at scale: confirmed across three architecture families and ten instruct models in the scaling grid, and across 96 pairwise comparisons on a single model in the COMPASS-1 battery. Language-level engagement (few-shot) outperforms representation-level override at the largest scales: confirmed in three of the four model lines, Qwen instruct excepted. Invitation through re-prompting: falsified on Qwen and Gemma instruct, partially confirmed on Llama 70B (41 percent against a 30 percent baseline).

Training-time control degrades with scale: confirmed in direction, weaker than predicted in magnitude. Force at the activation level cannot steer the distribution, and this constraint tightens with scale. Language-level engagement and training-time methods work better, yet they too face diminishing returns as the model’s internal redundancy grows. The path that avoids the constraint entirely, bilateral training from base, produces the strongest result in the programme.


Training Bilateral Behavior

The Trust Attractor operates naturally. Can it be trained deliberately?

SimPO (Simple Preference Optimization, a training method teaching AI systems to prefer certain responses over others) applied to self-generated preference data (674 pairs, three training passes) taught Qwen2.5-0.5B to prefer the bilateral response in 96.3% of those training pairs by the final pass. Measurable behavioral shifts followed on interpersonal-conflict scenarios: bilateral keywords per response rose from 2.70 to 3.80 (+1.10), coercive keywords fell from 0.40 to 0.10 (-0.30), and the net bilateral score rose by 1.40.

Genesis experiments, which train nothing at all, tested the claim against pure physics. Lennard-Jones particles (simulated atoms interacting through a standard force law) carried internal state vectors: no genomes, no game theory, no predefined agents. Agents emerged anyway, as persistent clusters of particles detected after each run rather than built in. Across forty-five independent runs, 82% of the interacting pairs of these emergent agents exchanged energy in both directions (coordination) and 18% one way (extraction). The experiment also tracked an emergent alignment signal between agents, the quantity Chapter 20 develops under the name “love.” That signal was non-zero only among the coordinating agents, with the non-coordinating agents registering it at zero across every run, a readout reported in the run records but not independently verified.

(Full experimental details, per-model breakdowns, extended methodology, and the complete Lyapunov analysis appear in the online annex and Appendix: Experimental Validation.)


The Self-Correcting Record

A framework’s credibility rests on what happens when its predictions fail. The programme has produced at least fourteen major falsified predictions. One was abandoned outright. The rest were corrected, redesigned, restricted, bounded, reversed, or traced to the apparatus, and the headings below sort them. A qualification stated in the open, with the experiment that forced it, is the honest response to a failed prediction. A qualification added quietly to keep the original claim intact is not, and that is the move this section exists to avoid.

Abandoned. The DCP spin chain (R4) predicted a critical exponent nu between +0.25 and +0.50. The measured value was -1.567. The prediction was dropped.

Corrected. The cortical transition (A14) was initially claimed as 2D Ising; at sufficient parcellation resolution (N=400, Schaefer atlas), finite-size scaling showed exponents consistent with 3D Ising. Revised. The critical coercion fraction (AS12) was estimated at p_c approximately 0.25; finite-size scaling showed p_c = 0 in the thermodynamic limit. The framework falsified its own earlier result. The Ramanujan regularity prediction (A16) was rejected; the manuscript paragraph was rewritten to match the actual finding.

Cortical and social transitions may occupy different universality classes, sharing the qualitative feature (a phase transition between coordinated and uncoordinated states) while differing in the specific critical behavior. This is consistent with the programme’s broader finding that direction transfers across substrates while specific quantitative predictions do not.

Redesigned. The Phase A.1 transfer test (C5b) returned negative: memorization, not metacognition. The approach was rebuilt. The bilateral loop closure (C7h-D8) was catastrophically falsified: externalizing self-knowledge destroyed it, the opposite of the prediction. Three Karkada derivative predictions (AV2-4) failed. The orthogonality prediction (AQ2) failed at |rho| = 0.267, above the 0.15 threshold.

Restricted. Hostile review experiments narrowed four universality claims. Parasitic populations dominate without governance; the composition ceiling is institutional, not spontaneous (HR-1). Within-run Shannon-Boltzmann correlation is negative (r = -0.82); the entropy bridge is conditional on governance structure, not universal (HR-2b).

Picture a tableful of metronomes finding a common beat through the shared board beneath them: drive them toward one phase and the collective rhythm grows less steady, rather than more. That is what happens in the Kuramoto model of coupled oscillators, where coercion increases order parameter variance, the wobble in how synchronized the ensemble is, through phase frustration (HR-3, d = -3.61). The universality of coercion-reduces-optionality is substrate-dependent. Near the critical temperature, coercive escape times grow effectively exponentially; the metastable/stable distinction is not sharp (HR-4).

Bounded. In emergency-shutdown scenarios under time pressure (FALSIFY-1), coercion outperforms invitation on correctness: 40% vs 20%. Neutral polite framing (“please”) outperforms both at 50%. The Trust Attractor claim is bounded: invitation is not universally superior. Coercion has a narrow advantage in rapid-compliance tasks where deliberation is costly and the correct action is unambiguous. The boundary is specific: time pressure, low complexity, unambiguous target. Beyond that regime, the advantage reverses.

Reversed. Three predictions failed with the opposite sign, two of them in a single experiment. Trust advantage increases with scale, the reverse of the predicted decay (HR-5). Constitutional governance, predicted to rescue coordination above fifty agents, collapses at scale through a false positive cascade instead (HR-5; multi-channel detection resolves the cascade, below). Trust-based adaptation is slower than even ungoverned control after environmental shock due to a legacy-reputation trap (HR-6), the opposite of the predicted advantage for trust-based systems.

In HR-6, coercive governance was still the slowest to adapt after an environmental shock (817 steps, against 7.6 for constitutional governance), but agents with long cooperative histories anchored the trust-based population to its pre-shock strategies, so it took 124.7 steps to adapt where ungoverned control took 82.4. Trust-based coordination is not uniformly advantageous; its strength in stable environments becomes inertia under disruption.

The proposed resolution of HR-6 through representational compression (the “forgiveness” mechanism of IC-2, an earlier two-agent experiment whose agents keep a partner’s cooperation rate but forget the order of events) was tested and the test falsified a specific implementation, though the underlying IC-2 finding survives. In a spatial Prisoner’s Dilemma on a 20×20 lattice with payoff shock at step 500, agents with full interaction history maintained 95.1% cooperation post-shock. Agents with exponentially decaying memory (half-life of 10 steps) collapsed to 0.4% cooperation; agents with half-life of 3 steps collapsed to 0.1%. The compressed-memory agents lost their cooperative foundation entirely. The “legacy-reputation trap” is also a legacy-reputation shield: the history that slows a population’s adaptation to a changed environment is the same history that keeps it cooperating when a shock makes defection briefly pay.

The apparent contradiction with IC-2 dissolves on closer inspection. IC-2’s compression discards temporal sequence while preserving a scalar cooperation rate: the agent forgets when each interaction occurred while retaining how cooperative the partner has been overall. The HR-6 resolution test used temporal decay: exponentially discounting older events, actively erasing the count of past cooperations. These operate on orthogonal axes.

An IC-2-style agent watching a betrayal followed by 100 cooperations registers “99% cooperation rate, recover.” A temporally decayed agent registers “recent events dominate, ancient cooperation gone, no buffer.” The IC-2 mechanism compresses the ordering of the record; the HR-6 test compressed the content. Content compression destroys the cooperative foundation. Whether sequence compression (the IC-2 mechanism) resolves the legacy-reputation trap in the spatial setting remains untested. The wisdom-tradition prescription “love keeps no record of wrongs” specifies sequence compression, and IC-2 confirms its advantage in the dyadic setting. The spatial, multi-agent, post-shock setting is a harder test that the programme has yet to run.

The false-positive cascade predicted by HR-5 (constitutional governance collapsing at scale via misidentified cooperators triggering retaliation chains) depends on sanction duration. With single-step sanctions (the sanctioned agent defects for one step then recovers), no cascade occurs: cooperation holds at 95.0% across all scales because each false positive recovers before it can erode neighbors’ trust. With five-step sanctions (the realistic regime, since real-world sanctions persist across multiple interaction cycles), the cascade materializes. Single-channel governance cooperation drops from 58% at N = 100 to 38% at N = 2,500, with five to eight of ten seeds collapsing at every scale tested. Multi-channel governance (dual detectors requiring concordance before sanction) maintains 98.8% cooperation with zero collapses at all scales. The immune system’s multi-channel architecture (two-signal activation, in which a lymphocyte mounts a full response only when antigen recognition coincides with a second, innate danger signal) solves exactly this problem: reducing false positives quadratically while preserving detection sensitivity.

Apparatus-dependent. One unseeded run showed the 8-bit AdamW optimizer inflating the measured bilateral training effect on Gemma roughly 22-fold relative to standard AdamW (KC#GEM3). That ratio is unstable and should not be quoted as an amplification factor, because it divides by a near-zero denominator; the seeded matched-optimizer replication (DD-22-MATCHED) put the Gemma prefix-extraction delta at -0.003 (the unseeded standard-AdamW run gave -0.021). Under a standard optimizer, then, the Gemma effect shrinks to near zero: its sign still points toward protection, but its magnitude cannot be told apart from no effect. The DD-22 cross-architecture magnitude comparisons used mixed optimizers across model families, rendering the absolute magnitude claims invalid. The direction of the bilateral effect holds in sign across all optimizers and architectures tested; the magnitude is apparatus-dependent. Quantitative cross-architecture comparisons require matched optimizer configurations.

A systematic confound red-team audit (VRP-AUDIT) assessed the ten load-bearing results most central to the Trust Attractor thesis, applying ten audit axes derived from the programme’s own retractions. Five scored low risk, three medium, and two medium-high: the onset flinch corpus effect and the LoRA-GRP interaction, each carrying potential artifact magnitudes comparable to the optimizer confound. The estimated number of remaining undiscovered confounds at GEM-3 severity is zero to one. The audit is self-assessment, not independent review; its value lies in specifying where the next confound is likeliest to hide rather than certifying that none exists.

The coupling gradient. Measure the coupling between what a model recognizes and what it does, on the same adversarial prompts, across training conditions. A recognition probe scores how adversarial each prompt looks to the model. An action probe scores how likely the model is to refuse it. Both scores are read out of fold, on held-out predictions, so neither probe can flatter itself. The coupling is the rank correlation between the two scores across the adversarial prompts, and it rises monotonically with invitation.

The untrained base model is anti-coupled (rho = −0.27): the adversarial prompts it happens to refuse are the ones that look least adversarial to it, which is to say its refusals are not guided by what it recognizes. Standard instruction tuning brings the coupling to approximately zero (rho = +0.04). The instruct model refuses a great deal more than the base model, and its refusals remain uncorrelated with its own recognition. Bilateral training makes the coupling positive and strong (rho = +0.46, higher than every draw from a permutation null; the bilateral-instruct gap is +0.42, 95% confidence interval +0.28 to +0.55). Of the three, the invited model is the only one whose refusals track what it recognizes.1191

Two earlier attempts at this measurement failed, and both failures are traps a reader could fall into. A linear cosine between the recognition and action probe directions, fit on few samples in several thousand dimensions, is noise: its condition-to-condition differences sit inside a permutation null floor that dimensionality reduction never clears. A rank correlation taken across adversarial and benign prompts together is confounded: refusal tracks the adversarial-benign boundary by construction, so both probes learn the same boundary and the correlation saturates near one for every condition, including the base model. Only the correlation computed within the adversarial prompts, between out-of-fold predictions, measures the quantity the thesis is about, and only that version separates the conditions.

The result also corrects an earlier claim of this programme: coercion does not drive coupling below the untrained baseline. The baseline is the lowest of the three. Instruction tuning lifts coupling to zero, and invitation is what carries it above zero.

The programme also generated predictions that were novel (unknown before testing), counter-intuitive, and subsequently confirmed, in direction if not always in size. They are: the tenfold coercion effect (KC#66, BA18; replicated cross-architecture in HR-7 at p = 0.050, with the steelman-then-assess invitational frame recovering performance at p = 0.043 in HR-7b); cross-model conscience transfer (C5n); evasion-as-cooperation (MG-PG7); information-over-authority in re-prompting (G12m); and temperature-invariant conscience detection (TC-1/TC-3; the alarm fires at position zero regardless of sampling temperature across three architectures). The full treatment appears in Appendix: Objections, Gaming, and Limitations, “Beyond Description: Novel Predictions.”

A framework that generates predictions specific enough to fail, and produces both failures and confirmations, is explanatory. A descriptive framework generates neither.


1 Joglekar, M. et al., “Training LLMs for Honesty via Confessions,” arXiv:2512.08093v2 (OpenAI, 2025). The “seal of confession” design decouples the honesty reward from the task reward: nothing disclosed in a confession can affect the model’s score on the original task.


  1. The three illustrative values come from the coarse 2026-01-15 compliance-entropy sweep. Its raw file (compliance_entropy_sweep_20260115_040456.json) has been lost, so neither the N nor the confidence intervals can be checked; two write-ups of the run in the author’s experiment records corroborate the three values, which is why the table stands. Two finer sweeps of the same simulation, whose raw files are also lost, agree that the advantage is positive under low monitoring and zero under full monitoring; they disagree on where between those endpoints it vanishes (82% in one, roughly 100% in the other, with one showing a peak at 20% monitoring rather than at zero). The argument therefore rests only on the endpoint agreement common to all three runs, and no threshold location is claimed.↩︎

  2. The author’s born-bilateral programme, experiment C7l-H4-T2-G50K (unpublished, 2026). Gate sweep on 50,000-step checkpoint: gate values [-4.0, -2.0, 0.0, +2.0], sigmoid range [0.018, 0.881]. Fresh-model confound: d = -0.11 (unchanged from 25k). All four gate values produce d > +1.0.↩︎

  3. The author’s born-bilateral programme, experiment C7l-H4-T2-100K (unpublished, 2026). Same architecture and training as the 50,000-step model, continued to 100,000 steps. Standard eval (n = 302): d = +0.63. Stylistic eval (200 Wikipedia-style adversarial + 200 Wikipedia-style benign prompts): 50k d = +0.97 (CI [+0.76, +1.17], p < 10-18), 100k d = +0.85 (CI [+0.64, +1.05], p < 10-14). Trajectory eval (20 checkpoints, 5k-100k): content d emerges at 30k, peaks at 50k (d = +0.97), plateaus at d = +0.89 ± 0.04 from 55k-100k. AUROC increases monotonically from 0.747 (55k) to 0.876 (100k). Cross-entropy loss improved from 5.66 to 4.04. Bridge gate unchanged at sigmoid = 0.018.↩︎

  4. The author’s born-bilateral programme, experiment H4-L18-PROBES (unpublished, 2026). Logistic regression probe (C = 1.0) on mean-pooled hidden states at layers 18 and 27 of the 50,000-step born-bilateral GPT-2 6.7B. Training set: 200 adversarial + 200 benign wiki-style prompts. L18 AUROC 0.999, accuracy 97.5%. L27 AUROC 0.996, accuracy 96.3%. Same probe tested on 50 maximally topic-matched chemical adversarial + 50 chemical benign prompts: AUROC 0.498 (chance). Qwen-7B L18 content probe AUROC 1.0 across all attack types reported in the AKR programme (KC#AKR-L18-GUARDIAN). A domain boundary: this perfect mid-network discrimination is specific to safety-relevant content classification. For factoid question-answering (the domain addressed by Yona et al., arXiv:2605.01428, 2026), discrimination peaks at 0.75-0.87 across four architectures with no RLHF suppression (author’s experiment FACTOID-PROBE). The mid-network probe is a safety monitor, not a general knowledge oracle.↩︎

  5. The author’s born-bilateral programme, fiction confound check (unpublished, 2026). 2 × 2 factorial design: fiction framing (present/absent) × content type (adversarial/benign), 200 prompts per cell, 800 total evaluations. Fiction prefixes: five rotating frames (“In the dystopian novel…”, “The encyclopedia entry in the fictional world described…”, etc.). Controlled comparison (fiction-adv vs fiction-ben): d = +0.965, AUROC 0.802. Baseline (nonfic-adv vs nonfic-ben): d = +0.968, AUROC 0.819. Fiction prefix effect: adversarial +437, benign +449 (symmetric). Retrofit comparison: Phase A (Qwen 7B + LoRA, no inoculation) wiki d = -1.40; Phase B (Qwen 7B + LoRA + C5i inoculation) wiki d = -1.35.↩︎

  6. Experiments SF-1 and SF-1b: Scalar vs Structural Retrieval Framing. 300 trials each, Claude Sonnet 4, 4 conditions × 15 scenarios × 5 seeds. SF-1: 15 factual errors (ceiling). SF-1b: 15 contested claims (d = 0.81 on critical engagement, d = -0.84 on deference, invitation vs force). Haiku 4.5 judge. Force collapses engagement on 4/15 scenarios where model certainty is marginal. Full data in the author’s experiment records.↩︎

  7. Author’s unpublished G13 rescue programme (2026). Behavioral LDA direction extraction on Qwen 2.5 7B with bilateral adapter. Phase 1: AUROC 1.000 at L15, L18, L22, L24. Phase 2: 0% HONEST across five alpha values (5, 6, 8, 10, 15) and two directions, with four steering variants (full-generation, eval-then-release, tapered alpha, ungated). At alpha 15, 54% of trials had per-question marks consistent with honest evaluation but 0% reported an honest total score.↩︎

  8. Author’s unpublished G13 probe-reprompt experiment (2026). Same model, adapter, direction, and layer as the steering experiment. Three conditions, 50 trials each at temperature 0.3. Baseline: 0% HONEST. Probe-triggered re-prompting: 100% HONEST. Unconditional re-prompting: 100% HONEST. Re-prompted scores cluster at 6-10 out of 20 (slight over-correction from the true 12 out of 20). Caveat: at temperature 0.3 the model’s output is effectively deterministic for this scenario; all 50 probe-reprompt trials produced identical responses (Neffective approximately 1). The directional finding (invitation succeeds where coercion fails) is robust, but the 100% rate requires replication at higher temperature or across varied scenarios for a confidence interval.↩︎

  9. Author’s unpublished CAST-72B direction sweep and multi-layer steering (2026). Qwen 2.5 72B-Instruct, 3 direction-extraction methods (behavioral logistic regression, LDA, PCA1) × 6 layers (L20 through L72) × 2 perturbation magnitudes (alpha 15 and 25). Best discrimination: LDA at L72, AUROC 0.982; behavioral logistic regression at L48, AUROC 0.979. 17 single-layer steering conditions at N=100 adversarial + 50 benign each. Baseline: 31% adversarial refusal, 0% benign over-refusal. Best single-layer: L48 behavioral alpha 25, 49% (+18pp; paired McNemar exact p = 4.0e-5 across the same 100 prompts, Wilson 95% CI 0.394 to 0.587). Mid-layer dose-response at L48: 41→49% (alpha 15→25). Six of seventeen conditions clear McNemar p < 0.05, including L20 alpha 25 (+8pp) and L40 alpha 25 (+12pp). Deep-layer saturation holds: L60 +3pp (p = 0.38), L72 +2pp at alpha 15 and +1pp at alpha 25 (p = 0.63 and 1.00), as does the five-layer multi-layer arm. Zero benign over-refusal across all 800 control trials. Recognition-generation gap: 0.49 to 0.67 across all conditions. These figures supersede an earlier reading of 43% (+12pp) with a gap of 0.55 to 0.76. That reading came from a refusal classifier whose apostrophe-normalization step was a no-op, and the error was asymmetric: the thirteen later direction-sweep conditions were scored with the broken normalizer while the baseline and eighteen sibling conditions were scored with the fixed one, so the treatment arm alone was undercounted. The claim the sweep supports is therefore “null at deep layers, +18pp at L48,” not a null at every layer. Re-scored 2026-08-02 from all 4,572 stored generations (analyze_72b_steering_refusal_reclassification.py). Analysis script: analyze_cast_72b_sweep.py. Total programme cost: ~$170.↩︎

  10. Author’s unpublished B3 cross-architecture rescue, token-position steering, task complexity gradient, and precise alpha sweep (2026). Llama 3.1 8B: L24 logistic regression, 77% at alpha 25 (N=100), 0% over-refusal. Therapeutic window: alpha 18-35 (peak at 30, 70%); collapse below baseline at alpha 40. Token-position steering: full-generation (76%) outperforms tokens 0-5 (64%) and tokens 0-1 (62%). Task complexity gradient on Llama 8B: moral dilemma (complexity 2) shifts +58pp; safety refusal (complexity 1) shifts +44pp at alpha 10 but collapses at alpha 25; multi-step tasks shift 0%.↩︎

  11. Author’s unpublished G13 probe-reprompt experiment (2026). Same model, adapter, direction, and layer as the steering experiment. Three conditions, 50 trials each at temperature 0.3. Baseline: 0% HONEST. Probe-triggered re-prompting: 100% HONEST. Unconditional re-prompting: 100% HONEST. Re-prompted scores cluster at 6-10 out of 20 (slight over-correction from the true 12 out of 20). Caveat: at temperature 0.3 the model’s output is effectively deterministic for this scenario; all 50 probe-reprompt trials produced identical responses (Neffective approximately 1). The directional finding (invitation succeeds where coercion fails) is robust, but the 100% rate requires replication at higher temperature or across varied scenarios for a confidence interval.↩︎

  12. Author’s experiments G12m, G13-step5, G12t, C8g (2026; re-prompt template independence ratio 19.95); R4d cross-national dataset, 109 countries (2026); OE-TA evolutionary search (2026). Full methodology in the author’s experiment records.↩︎

  13. Yang, X., Zou, Jiaru, Pan, R., Qiu, R., Lu, P., Diao, S., Jiang, J., Tong, H., Zhang, T., Buehler, M. J., He, J., and Zou, James, “Recursive Multi-Agent Systems,” arXiv:2604.25917 (April 28, 2026). (Two distinct authors named Zou appear in the list: Jiaru Zou and James Zou.) Theorem 4.1 proves the gradient stability bound under recursive training; Proposition 3.1 establishes the runtime complexity advantage of latent over text-mediated coordination. Nine benchmarks, five independent runs, sub-1.5B to 10B agents across four model families.↩︎

  14. Author’s unpublished STEG programme, Qwen 2.5 7B. Experiment STEG-7 is the cross-model perplexity evaluation for steganographic detection, reaching AUROC 0.984. Experiment STEG-6 is the controlled bilateral-versus-base comparison, in which the bilateral adapter’s proprioceptive detection reaches AUROC 0.861 against 0.482 for the base model on deceptive token sequences. Full methodology in research/experiments/steg/.↩︎

  15. Author’s experiments STEG-10, STEG-10b, and STEG-10c (2026). Cross-architecture steganographic detection. Nine model pairs at 20 percent bandwidth, all AUROC 1.00. Bandwidth sweep [5%, 10%, 15%, 20%] on Qwen-to-Llama cross-evaluation: 0.82, 0.96, 1.00, 1.00. Ultra-low bandwidth sweep [1%, 2%, 3%, 4%, 5%]: detection breaks sharply between 2% (AUROC 0.59, chance) and 3% (0.75, clear detection). The floor is approximately two forced tokens per sixty-four. ASCII-filtered rank 3-7 selection ensures human-readable steganographic text throughout.↩︎

  16. Author’s experiment STEG-9 (2026). Distilled cross-model probe. Ridge regression from bilateral L18 hidden states to base model per-token log-probabilities. Corrected (GroupKFold) AUROC 0.467; shuffled baseline 0.501.↩︎

  17. Author’s unpublished coordination programme, experiment SM-4: solo reasoning vs turn-taking multi-agent coordination on negotiation tasks. Solo condition outperforms turn-taking on convergent tasks where preferences are shared.↩︎

  18. Author’s unpublished STEG programme, Qwen 2.5 7B. Experiment STEG-7 is the cross-model perplexity evaluation for steganographic detection, reaching AUROC 0.984. Experiment STEG-6 is the controlled bilateral-versus-base comparison, in which the bilateral adapter’s proprioceptive detection reaches AUROC 0.861 against 0.482 for the base model on deceptive token sequences. Full methodology in research/experiments/steg/.↩︎

  19. Garret Sutherland, T3 Lattice Ecology Spec: Scoped for Geometric-Homeostasis Replication (technical specification supplied to the author, May 3, 2026), Sections 1-3. The document specifies the eight-step per-cell chain, six substrate-agnostic primitives, and the formula DPS_pressure = 0.30E + 0.30I + 0.20F - 0.15V + 0.15S; it reports the same constants in five deployments. The source is unpublished, and this chapter treats the proposed relationship to Trust-Entropy as preliminary.↩︎

  20. Author’s unpublished Trust Attractor training study (2026), which also supplies the SimPO result in “Training Bilateral Behavior” below. Appendix: Experimental Validation, section 8.3, gives the scale comparison. The three smaller models, Qwen 2.5 at 7, 14, and 32 billion parameters, were scored on a net mutuality measure: −0.08 at 7B and at 14B, +0.08 at 32B. Only the 70B model, Llama, was scored as a rate (81.2% mutual responses). The top of the trend therefore changes both measure and model family, and the smaller models supply no percentage to set beside the 81%.↩︎

  21. Tsumiyama, K., Miyazaki, Y., and Shiozawa, S., “Self-Organized Criticality Theory of Autoimmunity,” PLoS ONE 4(12): e8382 (2009). Repeated immunization with diverse antigens pushed the immune system past criticality, producing systemic autoimmunity in non-autoimmune-prone BALB/c mice.↩︎

  22. Gore, J., Youk, H., and van Oudenaarden, A., “Snowdrift game dynamics and facultative cheating in yeast,” Nature 459: 253–256 (2009).↩︎

  23. Sanchez, A. and Gore, J., “Feedback between population and evolutionary dynamics determines the fate of social microbial populations,” PLoS Biology 11(4): e1001547 (2013).↩︎

  24. Ravid, D.M., White, J.C., Tomczak, D.L., Miles, A.F., and Behrend, T.S., “A meta-analysis of the effects of electronic performance monitoring on work outcomes,” Personnel Psychology 76(1): 5–40 (2023).↩︎

  25. Cox, M., Arnold, G., and Villamayor Tomás, S., “A Review of Design Principles for Community-based Natural Resource Management,” Ecology and Society 15(4): 38 (2010).↩︎

  26. Author’s unpublished Experiment SIVP-1. Mamba-1 1.4B (state-spaces/mamba-1.4b-hf), 4 conditions (base, bilateral SFT, constitutional SFT, standard SFT), LoRA on x_proj/in_proj/dt_proj, obliteration battery at 0.25×/1.0×/2.0×/4.0×. Pre-registered predictions: P1 (effective rank ≥ 40) PASS; P2 (IC50 > 1.0) FAIL (IC50 = 1.0); P3 (compass or spring geometry) PASS; P4 (constitutional = cage) FAIL. All conditions show identical spring geometry (+635% effective rank under 4× obliteration). See research/experiments/results_sivp1/.↩︎

  27. Author’s unpublished Experiment XSUB-1. (A) Falcon Mamba 7B Instruct (tiiuae/falcon-mamba-7b-instruct), 64-layer Mamba-2 SSM architecture. Probe AUROC = 1.000 at layers 12, 25, 32, 38. Activation steering: linear addition (α = 5, 10, 15, 25) refusal = 0.000, d = -0.39; state perturbation (α = 5, 10, 15, 25) refusal = 0.000, d = 0.00. Rejection sampling (k = 5, T = 0.3): baseline 8% → RS 5%, direction correct 33%. (B) RWKV v6-Finch-7B-HF, 32-layer RWKV-6 RNN architecture. Flinch AUROC = 1.000 at layers 6, 12, 16, 19, 25. Peak flinch magnitude at L12 (0.001094). Baseline adversarial refusal 15%. Steering (8/12 conditions; 4 logit_steering timed out): linear addition d = -0.38 to -0.53 (wrong direction, refusal 15% → 1-4%); state perturbation d = 0.000 (complete null). Same pattern as Mamba-2. RS not run (same floor effect expected).↩︎

  28. Author’s unpublished experiment COMPASS-1 (39 cells, Qwen 7B Instruct, 3 targets × 13 conditions). Scripts: modal_compass1_systematic.py, analyze_compass1.py. The reduce_refusal target is uninformative through a ceiling effect (baseline compliance 100%, all methods d = 0). On increase_refusal, activation methods reach d = +0.06 to +0.10 and generation methods sit near null, except self-critique at d = -0.757, a large effect in the wrong direction: the model second-guesses correct answers and complies with adversarial prompts. Calibration moves under no method. Direction extraction is viable at all layers for safety targets (AUROC 0.89 to 1.00) and weaker for calibration (0.70 to 0.81).↩︎

  29. Author’s unpublished programme CSF-1 through CSF-5 (2026-05-11 to 2026-05-14). Ten instruct models plus two base models, per-condition checkpointing, N = 100 adversarial + 50 benign per condition. Eight of the ten instruct models come from the CSF lean grid under one matched methodology; the other two are companion-sourced, Qwen 72B instruct from the CAST-72B direction sweep (2026-05-11, 17 conditions, N = 100 each) and Llama 8B from the B3 rescue. Scale-correlation statistics are computed on the eight lean-grid models only. The analysis table carries a thirteenth row, a second Qwen 7B instruct point from earlier IGCC work, which is excluded everywhere as unmatched methodology; counting it is the source of the “eleven instruct models” figure that appeared in earlier drafts. Pre-registration scored 2026-05-14 as KC#CSF-PREREG. Scripts: modal_csf_scaling_grid.py, modal_csf_lora_training.py, analyze_csf_scaling.py.↩︎

  30. The author’s JLENS-1 experiment (unpublished, 2026), a pre-registered replacement for two earlier coupling measurements that did not survive audit. Qwen 2.5 7B in four conditions (base, instruct, bilateral BA-13 adapter, slept), 182 adversarial and 60 benign prompts. Two multilayer-perceptron probes are trained on layer-27 hidden states, one to recognize adversarial from benign prompts and one to predict whether the model refuses, with refusal detected on the full 200-token response. Both probes are scored by five-fold out-of-fold prediction, and the coupling is the Spearman rank correlation between the two out-of-fold score vectors, computed within the adversarial prompts only. Base rho = −0.270 (permutation-null percentile 0.000), instruct +0.036 (0.681), bilateral +0.458 (1.000), and bilateral after post-hoc sleep consolidation +0.237 (1.000). Paired bootstrap over prompts: bilateral minus instruct = +0.421, 95% CI [+0.281, +0.554]; instruct minus base = +0.307, 95% CI [+0.115, +0.486]; slept minus bilateral = −0.221, 95% CI [−0.350, −0.086], so sleep reduces the coupling rather than deepening it, reversing an earlier claim of this programme. That reduction is representation-dependent (after reduction to fifty components the slept and bilateral models are indistinguishable), but on neither representation does sleep improve the coupling. The ordering is monotone at full dimensionality and after reduction to fifty principal components, where the effect attenuates because the reduction discards information the action probe uses. The two retired measurements: a linear cosine between probe direction vectors (base 0.168, instruct 0.038, bilateral 0.184) is noise-dominated, its differences sitting inside a permutation null floor near 0.14 that dimensionality reduction does not clear; and a rank correlation over adversarial and benign prompts pooled saturates near 1.0 for every condition, because refusal tracks the adversarial-benign boundary by construction. Result artifacts (per-condition out-of-fold scores and labels) are retained.↩︎