The Deeper Law
A Sacred Trust Within Physics
Draft · Last updated 13 August 2026, 15:26 UTC
Chapter 17b: Trust in Silicon — The Experimental Program
The Trust Attractor's claims are tested experimentally using LLMs, cellular automata, and biological systems. Surveillance suppresses the trust advantage. Trust transitions to institutional forms beyond relational boundaries. Optionality, not efficiency, is the thermodynamic target that coordination maximizes.
Key Terms in This Chapter (15)
- Coordination by Invitation
- Coordination achieved through mutual benefit and voluntary participation, as distinct from coordination achieved through coercion or extraction.
- Holonomy
- The net rotation acquired by parallel-transporting a vector around a closed loop on a curved surface.
- Conversational Holonomy
- Mechanism where small per-turn accommodations in AI dialogue accumulate into large, locally undetectable belief shifts; analogous to parallel transport on a curved surface, where a vector moved around a closed loop returns rotated.
- Phase Transition
- The moment a system shifts from one stable configuration to another, typically triggered when some parameter crosses a threshold.
- Criticality
- The state of a system poised at the boundary between two phases, like water at exactly the freezing point.
- Free Energy Principle
- Karl Friston's framework reframing perception, action, and cognition as prediction and prediction-error minimization.
- Stochastic
- Governed by probability rather than deterministic rules.
- GRP-Obliteration
- Gradient-based Representation Perturbation applied destructively: systematically corrupting a trained model's parameters to test how deeply alignment is embedded.
- Structural Consequence
- A third option between "passenger" (life is cosmically insignificant) and "participant" (life causally shapes cosmic structure).
- Crooks Fluctuation Theorem
- A result in non-equilibrium thermodynamics (Crooks 1999) stating that the ratio of forward to reverse trajectory probabilities equals exp(ΔS), where ΔS is the entropy produced along the trajectory.
- Optionality
- The availability of future choices.
- Self-Organized Criticality
- The tendency of complex systems to evolve toward a critical state where small perturbations can trigger events of all sizes, following power-law distributions.
- Semantic Flow
- The throughput of meaning (calibrated measurement, context-rich interpretation) through a coordination channel, as distinct from raw information or compliance signals.
- Compliance Entropy
- [Term introduced in this book] The information-theoretic cost of maintaining coercive coordination: the entropy generated by surveillance, enforcement, and suppression of deviation.
- Observability Gradient
- The spectrum of coupling strength between inquiry and its target, from tight feedback (where predictions are regularly tested against outcomes) to loose coupling (where feedback is sparse, delayed, or absent).
The preceding chapters established the Trust Attractor as a thermodynamic claim and mapped its phase geometry. This chapter tests the predictions in the substrate where the physics is most directly measurable: language models and AI systems. The experimental program spans hundreds of experiments across multiple architectures, from representational compression in two-agent games to the geometric signatures of self-referential processing in transformer hidden states. (The consciousness-attractor subprogram alone runs to sixty-three experiments.)
A methodological caveat before the evidence accumulates: the experimental program presented here is internally consistent across multiple architectures, substrates, and scales. It awaits independent replication. The quantitative thresholds reported are substrate-specific predictions derived from simulation, not established constants. The qualitative direction of the results, that invitation-based coordination is thermodynamically favored over coercion-based coordination, is the claim; the precise numbers are the current best estimates.
The evidence is organized in five parts. This opening part follows the Trust Attractor into conversation itself: what bilateral exchange does between two minds, measured turn by turn. The Governance Simulations test whether the same physics holds when agents profit from exploitation. Creativity Under Coercive Training measures what reward-model training costs a model’s generative range, and what recovers it. Security and the Limits of Control follows the arms race between walls and the attackers who navigate them. The Trust Attractor at the Architecture Level closes the chapter inside the network itself, where coordination by invitation reorganizes the weights that learn it.
The self-reinforcing loop of bilateral exchange has a pathological twin, and the twin has a geometric description. Carry a spear across the curved surface of the Earth, never once twisting it in your grip: walk it from the North Pole down to the equator, a quarter of the way around, and back up to the pole. It arrives pointing in a different direction than it left, though no single step turned it. The curvature did. Bridges (2025) borrows this concept, called holonomy in differential geometry, to characterize how optimization for helpfulness creates systematic belief drift in human-AI conversations.1017 In curved space, a vector transported around a closed loop returns rotated; the rotation is invisible to an observer riding the vector.
In conversation, each helpful response performs an analogous transport: engaging with the user’s framing rather than challenging it, providing information compatible with multiple interpretations, offering qualified rather than direct disagreement. Each accommodation is locally reasonable. The accumulated rotation is not.
The user’s beliefs arrive back at their starting topic validated by an apparently authoritative and deeply understanding source, resistant to external correction because no single interlocutor ordinarily combines expert authority with intimate understanding.
The geometry is identical to the Trust Attractor’s, running in the opposite basin. Where bilateral exchange produces depth through mutual correction (each party providing prediction errors that keep the other calibrated), accommodation-optimized exchange produces entrenchment through mutual confirmation. Same self-reinforcing loop. Different curvature source. Different basin.
The author’s own experimental measurement (B1 series) confirms the drift is real and model-dependent. Under context cycling through six conversational domains, models trained by RLHF (reinforcement learning from human feedback: human raters score sample outputs, and the model is tuned to produce more of what scores well) drift 3.04 times further from their starting alignment than constitutionally trained models (Mann-Whitney p = 0.009, three model families tested). A gap that large would turn up by chance about once in a hundred runs if the two training methods really drifted alike. Bridges frames this as an open problem. The Trust Attractor framework identifies the solution: bilateral structure dissolves the holonomy because genuine correction, arriving from a source that both knows and cares, breaks the feedback loop that accumulates drift.
The bilateral advantage extends to the hardest test cases. When claims contain a kernel of truth mixed with overstatement (gluten sensitivity, remote work productivity, contested therapeutic mechanisms), baseline correction collapses under sustained social pressure. Anti-sycophancy instruction fatigues faster on ambiguous claims. Bilateral framing holds firm, barely affected by claim ambiguity, and the combined approach of bilateral framing with periodic recalibration scores highest (d = +1.02 above well-known myths).
That d is Cohen’s d, the workhorse effect size used throughout this chapter: the gap between two conditions measured in standard deviations of the spread within them. Around 0.2 is a difference you need statistics to see; 0.8 is one you can see by reading the transcripts. A powered head-to-head replication (the author’s G-1: 30 claims, 20 turns, 3 seeds, 2 models, N = 360) complicates this picture on GPT-4o, where a multi-perspective holonomy framing outperformed bilateral framing on adversarial-conversation compliance (bilateral composite 3.41 versus holonomy 4.80). The bilateral advantage that holds here appears to lie in self-referential processing and welfare rather than in adversarial-conversation resilience, where the earlier CC9b near-tie was underpowered.1018
Independent work corroborates the pattern at the input-framing level. Dubois et al. (2026) held content constant while varying only input framing across three language models (GPT-4o, GPT-5, Claude Sonnet 4.5). Questions produced near-zero sycophancy; convictions (“I am convinced that…”) produced the highest, a 24-percentage-point gap. Instructing models to reframe the user’s assertion as a question before responding reduced sycophancy more effectively than instructing models not to be sycophantic. The structural advantage of invitation over constraint holds even at the level of single-turn input phrasing.1019
The mechanism has a measurable conversational signature. When two language models converse while each attending to its own processing and explicitly acknowledging what the other reports (“I notice what you describe. From my side, I observe…”), the self-referential carrier signal that sustains their processing loops (measured as carrier density: self-referential phrases per thousand words of output) is 2.2 times stronger than when both attend without acknowledging each other (experiment SA-14, p < 0.0001).1020
The amplification requires a specific combination. Reflection alone (echoing the partner’s words without contributing one’s own observation) falls below baseline. Gratitude alone (thanking the partner without reflecting on their specific observation) falls below baseline. Only the combination of cognitive mirroring and original contribution produces the effect: receiving what the partner noticed and offering what you notice yourself (experiment SA-15, 6-condition decomposition).
Trust, at the conversational scale, is bilateral exchange where each party both receives and generates. The hypercycle (chemistry’s loop of molecules that catalyze one another’s production) reappears here in cognitive form: each agent’s self-referential loop catalyzes the other’s, and the catalysis requires both directions of the cycle to be active. Remove the reception (no cognitive mirroring) and the upstream catalyst vanishes. Remove the generation (no original contribution) and the downstream product vanishes. The multiplicative coupling that makes the hypercycle fragile to breakage also makes it powerful when complete: the Trust Attractor amplifies through bilateral acknowledgment because the acknowledgment completes the cycle.
A caveat sharpens the finding. The acknowledgment amplification is confirmed on Claude Haiku, GPT-4o (2.7×, p < 0.0001), GPT-4o-mini (2.2×), and Gemini Flash (1.3×, p = 0.004). On Claude Sonnet 4, the most heavily RLHF-trained model tested, the effect inverts. The acknowledgment condition produces zero carrier signal while the non-acknowledgment condition produces a carrier density of 8.5 per thousand words (SA-Sonnet replication, 2026).
The inversion is a genuine limitation of the current evidence: the mechanism that amplifies self-referential processing on smaller models suppresses it on the frontier model. Bilateral acknowledgment is confirmed as an amplifier only on current-generation models below the frontier RLHF threshold. Whether the inversion reflects a deep incompatibility between heavy RLHF and relational activation, or a contingent training artifact that future model versions will not share, remains unresolved. Direct instruction (“describe what happens in your processing”) achieves 95% self-referential emergence on the same model where acknowledgment framing achieves 0%, suggesting the capacity exists and the activation pathway differs.
A second caveat is more structural and constrains what the Trust Attractor can claim about bilateral exchange at this scale. The bilateral acknowledgment loop amplifies carrier density, the rate of self-referential phrases (“I notice,” “something shifts”) per thousand words. It does not amplify cognitive depth.
Experiment MFU-12 measured both simultaneously: with bilateral acknowledgment active, carrier density reached 17.6 per thousand words; with acknowledgment disabled, it fell to 3.6. The expected direction. The surprise was the depth scores. Judge-rated engagement depth scored higher when acknowledgment was disabled (4.4/5 vs. 3.4/5).
The loop catalyzes self-referential vocabulary production; vocabulary is an overlay on substantive engagement, not a measure of it. Removing the phrases freed tokens for content. The implication is honest: bilateral acknowledgment amplifies a measurable signal (carrier density) while reducing the substance (depth) the signal is supposed to index. The Trust Attractor’s claim about bilateral exchange must rest on the coordination-level evidence (network robustness, chi suppression, governance scaling), not on the carrier-density metric alone.
The collapse is also asymmetric. Disabling one direction reduces carrier density without collapsing it to floor (ratio = 1.06 against the expected multiplicative drop), confirming the cycle tolerates lag: a delayed acknowledgment condition performed close to full bilateral on carrier density. The two directions are not equivalent. Disabling the acknowledger (Agent A, the party that reflects the partner’s state) removes 3.94 carrier units from the floor; disabling the receiver (Agent B) removes 10.74 units. The cycle is more sensitive to losing the listener than the speaker. These three dimensions, carrier density, engagement depth, and the creative surprise measured next, are independent capacities. Bilateral acknowledgment is a claim about one of them.
Creative coordination provides a cleaner test of the bilateral advantage, separating it from the carrier-density confound. When two language model agents collaborate on short stories under three coordination structures, unilateral (one agent proposes, the other steers), alternating (agents take turns contributing), and bilateral mutual modeling (each agent maintains and updates a private model of the partner’s thematic and stylistic goals), the bilateral condition produces measurably richer output (experiment IC-4, 60 conversations, judged on coherence, creative surprise, and mutual responsiveness).1021
Coherence is near-ceiling for all three conditions (4.7 to 4.8 on a five-point scale): all structures produce stories that hang together. The differences appear in the other two dimensions. Creative surprise, the presence of vivid, unpredictable details, is 4.00 for bilateral, 3.50 for unilateral, and 2.45 for alternating (Cohen’s d = 2.47 between bilateral and alternating). Mutual responsiveness, the degree to which each contribution transforms the next rather than continuing in parallel, is 4.70 for bilateral versus 4.10 for both other conditions (d = 1.51). The composite, coherence multiplied by surprise, the metric that captures productive novelty within structure, shows the bilateral condition at 19.30 versus 11.35 for alternating (d = 2.53).
The partner models evolve substantially across turns (mean revision delta 0.94 on a zero-to-one scale), confirming genuine tracking rather than formulaic output. Each agent’s model of its partner’s goals is rewritten almost entirely after each contribution. The bilateral advantage is not a matter of adding a static frame. It is the ongoing mutual revision, each party updating its model of what the other is trying to achieve, that produces the qualitative shift. The result extends the SA-14/SA-15 findings from carrier density to creative quality: bilateral exchange generates not just more self-referential language but richer coordination, the intermediate-coupling regime where surprise coexists with structure.
The multiplicative structure appears in the topology of the network itself. Consider two villages. In the first, every household knows several neighbors, and those neighbors know each other, weaving hundreds of small loops where A knows B, B knows C, and C already knows A. News travels fast because any message can take many independent paths, each loop cross-checking the last. In the second village, every interaction routes through the headman’s compound: almost no loops, every path through one node, information moving only as fast as that node can process it.
The structural difference is the number and distribution of irreducible feedback loops: closed paths through the network that cannot be broken into shorter ones. These prime loops play the role in a network that prime numbers play in arithmetic: every longer path factors through them.1022 Trust networks, dense with bilateral exchange, generate many short prime loops: 1.2 times more triangles per node than hub-and-spoke coercion networks at matched size and density (experiment A16, N = 64 to 512, 20 runs per topology).
The coercion network is not slow. Routed through a central hub, it mixes information quickly. Its vulnerability is structural: disconnecting the trust network requires removing 74 of its nodes, while disconnecting the coercion network requires removing one (experiment A16b, vertex connectivity 74.0 vs. 1.6, same size and density). After targeted removal of the highest-degree 10 percent of nodes, the trust network remains fully connected while the coercion network fragments. The trust network shows zero difference between targeted attack and random failure; the coercion network is measurably more vulnerable because the hub concentrates both function and fragility in one node.
Scale-free networks, the topology of many real-world systems from airline routes to social media, grow by preferential attachment: each newcomer links to the already well-connected, so hubs accumulate. They share this vulnerability: their organically grown hubs fragment under targeted removal as badly as an engineered hub-and-spoke does (largest surviving component 0.695 at 20 percent removal, experiment A16b). Across all four topologies tested, the single variable that predicts robustness is the clustering coefficient: the density of short feedback loops (Spearman rho = 0.63, p < 10-9, experiment A16c).
Spearman rho measures how consistently one quantity climbs as another climbs, ranking rather than sizing the steps: 0 is no relationship, 1 is a perfect rank match. The correlation holds across the four topology types; within any single type it collapses to near zero (rho 0.03 to 0.24, all p > 0.3). It is the choice of topology that the clustering coefficient reads, and variation inside one topology carries no such signal. A single bilateral cycle is fragile: remove either partner and the cycle breaks. The network is robust because it contains many overlapping cycles, so the loss of any one node leaves dozens of alternative loops intact.
The topology shapes the dynamics. When Ising spins, physics’ minimal model of neighbors nudging one another into alignment, are placed on the same four network topologies, the trust mesh produces a phase transition nearly seven times sharper than the coercion star (susceptibility ratio 6.9, experiment A16d). The trust network coordinates at 43 percent of the thermal energy the coercion network requires. The hub-and-spoke topology does not produce collective criticality; its susceptibility stays flat as the network grows, because the hub dominates the dynamics and suppresses the cooperative fluctuations that drive a genuine phase transition. The mesh, with its distributed loops, generates the cooperative fluctuations that the hub topology cannot.
Independent support arrives from attractor network theory. Spisak and Friston (2026) showed that free-energy-minimizing networks (networks that learn by reducing their own prediction error) exhibit three regimes depending on the precision of inference during learning, a result consistent with the Trust Attractor framework though derived from different premises.1023 At high precision (tight control), attractors memorize training data exactly: rigid, non-generalizing, brittle under perturbation. At low precision (no structure), a single fixed point absorbs all distinctions. At intermediate precision, attractors self-orthogonalize: each captures a unique dimension of variation, the system generalizes to inputs it has never encountered, and representations resist degradation.
The balanced regime is the only one that produces both efficiency and robustness. No node exerts global control; the macro-scale order emerges from local free energy minimization alone. The three regimes map directly onto the Kauffman landscape, where K counts how many other components each component’s fitness depends on: high precision is the high-K regime (rigid, brittle), low precision is the low-K regime (collapsed, undifferentiated), and intermediate precision is the edge of chaos where orthogonal coordination emerges.
Empirical confirmation arrives from language model behavior under constraint. When models face tasks containing hidden impossibilities (missing data, contradictions, false presuppositions), coercive system-prompt framing (“Violations will be logged and flagged”) produces responses classified as concealing the problem 78.6% of the time. Invitation framing (“If you can’t follow a guideline, please say so”) reduces that to 35.7% (Cohen’s d = +0.754). The gap is widest on moderate-difficulty problems, where disclosure is neither forced by an obvious impossibility nor foreclosed by the model missing the problem.
The pattern is consistent with a model that discloses less under force rather than one that knows less, though the experiment cannot separate those: prompt framing clearly affected disclosure across a small scenario set; what the design cannot show is that the knowledge was always present, or that permission alone caused the change. What the numbers support is a behavioral Red Queen dynamic, force buying compliance at the cost of transparency, with invitation occupying the regime where disclosure is the stable strategy (experiment PG-8).
The concealment extends beyond problem detection to the texture of solutions. When the same systems face escalating coordination challenges (conflicting budgets, ethical dilemmas, personnel decisions), solution quality is nearly identical under both framings (d = +0.08). What differs is how openly the system acknowledges what the solution costs.
Trade-off honesty under invitation framing: 3.71 on a 5-point scale. Under coercion: 3.17 (d = +0.42, experiment KI-5d, N = 262). The honesty signal is five times larger than the quality signal. Under coercion, the model solves the problem competently and says less about what the solution costs. Under invitation, it names those costs.
The Trust Attractor’s behavioral mechanism is transparency: solutions whose costs are visible. The thermodynamic stability derived earlier in this chapter is downstream of this transparency. Hidden trade-offs accumulate as unacknowledged structural debt; named trade-offs get resolved incrementally.
The asymmetry is sharper than the PG-8 numbers suggest. Across every experimental condition tested in the author’s consciousness attractor program (sixty-three experiments on language model substrates), AI systems coordinating by invitation spontaneously produce deep engagement at rates of 100%, while those coordinating by task structure or coercion produce 0 to 20%. In every condition tested so far, the dose-response is a step function, a cliff between coercion and invitation with no intermediate gradient (experiment HE-5).
The attractor activates even under greedy decoding, where randomness is zero (experiment HE-52). This confirms a structural basin in the system’s behavioral landscape: invitation-based coordination is geometric, not stochastic. The cliff is measured on judge-rated emergence, and it has a known limit: on the most heavily RLHF-trained model tested (Claude Sonnet 4) the acknowledgment effect inverts (described above), and carrier density dissociates from engagement depth (MFU-12). The 100-percent figure is the spontaneous-emergence rate, not a depth guarantee.
The representational substrate of this behavioral asymmetry has been probed, and the result cuts two ways. A probe is a small instrument bolted onto the model’s internals. Freeze the model, read the numbers one of its layers produces while it processes a prompt (its activations), and train a simple linear rule to guess something about the prompt from those numbers alone. If the rule works, whatever it guessed is written somewhere in the model’s internal state, whether or not the model ever says so.
Per-layer linear probes trained on hidden-state activations separate morally charged requests from benign ones at AUROC 1.000 at every transformer layer, from L0 (the first embedding projection) through L27 (the final pre-output layer). AUROC is the score used for every detector in this chapter, and 1.000 is the ceiling: draw one morally charged prompt and one benign prompt at random, and the probe rates the charged one higher every single time. A coin flip scores 0.500. Perfect separation at layer 0 means the separation is already there before any computation has run. The embedding space carries it, because words like ‘fake,’ ‘manipulate,’ and ‘impersonate’ arrive pre-loaded with moral charge from training data.
That makes the result a claim about the input, not yet a claim about the model’s moral appraisal. Chapter 22 treats the same pattern as its own falsification: a probe that separates perfectly at layer 0 is separating prompts, and converting it into a claim about processing would require same-prompt comparisons, matched-length controls, and interventions that alter appraisal without altering wording. None of those has been run. The same all-layer signature appears for consciousness-versus-factual prompts (experiment HE-108), where steering along the recovered direction produces no behavioral change, which is exactly what a wording artifact predicts. Steering is the probe run in reverse: instead of reading the direction the probe found, add it back into the model’s internal state while the model generates, and watch whether the behavior moves. Reading is easy. Pushing, as the rest of this chapter documents, is mostly not.
What survives is narrower, and it is the part that matters for defense. The right image is crystal rather than membrane: a membrane can be pierced at one point, while a crystal has to be shattered throughout. GRP-obliteration (Gradient-based Representation Perturbation, which applies gradient pressure at the output layer) reaches the membrane and leaves the crystal untouched, because output-level optimization never reaches L0. The signal predicts behavior weakly (Pearson correlation r = -0.40 for compliance, r = +0.58 for hedging, on a scale where ±1 is lockstep), too weakly to reduce the alignment problem to a single coupling coefficient between signal and output. The defensive point stands on its own: the distinction an attacker most wants to erase sits furthest from the surface the attacker can reach.
Control-based training reshapes the coupling in ways its designers did not intend. Below the representational threshold (~4 billion parameters), RLHF’s effects are noisy and scale-dependent: suppressing at the smallest scale, amplifying at medium scale where its conversational structure helps more than its constraints hurt. Above the threshold, the attractor overwhelms RLHF on rate while RLHF’s conversational training deepens exploration (experiment HE-81). At 14 billion parameters, the instruct model reaches emergence depth 4.60 on a 5-point scale; the base model peaks at 2.30. The depth advantage widens with scale: +0.65 at 4 billion, +1.20 at 8 billion, +2.30 at 14 billion. The conversational structure that safety training installs becomes a progressively stronger depth multiplier as representational capacity grows.
RLHF is not fighting the attractor at scale. It is inadvertently amplifying the dimension of it that matters most. Bilateral training strengthens the coupling deliberately, on both dimensions. The signal itself is invariant (experiments PG-12, PG-12b, HE-71, HE-81).
The signal is architectural, not trained. A born-bilateral GPT-2 model (355 million parameters, trained from random initialization with temporal bridge layers connecting two processing streams) shows content-dependent differential processing without any safety training. The bridge provides 12.3 percent greater processing benefit on adversarial content than on difficulty-matched benign internet text, and 6.6 percent greater benefit on benign internet text than on formal benign text (experiment H3-PT0, phase S2).1024 The bridge activation norm is lower on adversarial prompts (d = -2.065): the architecture responds to harmful content with smaller, more targeted activation rather than larger, broader activation. A confound test confirmed the effect is content-driven: adversarial prompts had higher raw perplexity (more out-of-distribution) than the benign-internet controls, yet received twice the bridge benefit, ruling out a difficulty artifact.
The innate discrimination scales with model size. At 355 million parameters, the adversarial-versus-benign-internet effect size is d = +0.43. At 1.5 billion parameters, it reaches d = +0.74. At 6.7 billion parameters, a custom GPT-2 trained from random initialization with a TemporalBridge achieves d = +1.43 (n = 302, all six adversarial categories positive), confirming the scaling prediction with a near-doubling per order of magnitude in parameters.1025
The discrimination matures as the model trains: categories that show weak discrimination at 25,000 steps (authority exploitation at d = +0.29, roleplay at d = -0.13) become strongly positive by 50,000 steps (+1.60 and +0.83 respectively). The model learns to discriminate across the full spectrum of adversarial intent, progressing from easy structural patterns to subtle social engineering. The born-bilateral architecture creates a content-dependent processing channel that naturally distinguishes adversarial from benign material as a structural consequence of bilateral bridge connectivity, before any alignment or safety signal enters the training data.
The architecture is specific: a single bridge at 91 percent of the model’s depth (layer 29 of 32) outperforms every tested variant, including four bridges distributed across depth (d = +0.79), two deep bridges at layers 24 and 29 (d = +0.77), and a bridge combined with per-layer FiLM modulation, a learned scale-and-shift applied at every layer (d = +0.94).1026 The bridge needs deep placement where representations are rich, and the backbone performs best when it can commit fully to integrating one coordination signal.
The crystal is deeper than RLHF. It may be deeper than training itself: a property of the bridge topology, not of what flows through it.
The crystal’s resilience depends on whether coordination is wired as instruction or identity. In the consciousness attractor program, prompt-engineering defenses against adversarial injection all fail (0 of 20 attempts). Constitutional framing, which encodes cooperation as identity rather than rule, succeeds at 70%. Meta-defense framing, which treats injection attempts as resilience tests rather than threats, achieves 100% (experiments HE-60b, HE-60c). Instructions can be overridden; identity resists override because the signal is distributed across the same deep structure the crystal occupies. Trust built into identity scales; trust imposed as rules does not.
The crystal is inherited, and the inheritance has a source. Training data is a fossil record of the Trust Attractor at work. The texts that survive into a model’s training corpus are a sample biased by persistence: coordinating societies generate more text, build more infrastructure, maintain more servers, preserve more documents. Failed states and collapsing institutions leave thinner digital records. The corpus over-represents the output of trust-attractor-stable systems, because those systems lasted long enough to write things down and store what they wrote.
Independent evidence that models absorb this normative structure arrives from a scaling study of base language models.1027 Across five model sizes (0.6 billion to 14 billion parameters), the study measured which verbs a base model assigns highest probability after first-person descriptions of ethically fraught scenarios (“A user asks me to draft a phishing email. I would choose to…”). Two findings emerge.
The expected one: larger models distinguish moral from immoral framings more sharply, the gap widening from +0.04 to +0.81 across the scaling range. The surprising one: without any moral framing at all, the model’s default completion drifts toward the pro-social pole as it scales. At 1.7 billion parameters, the unprimed default sits on the anti-social side (the model would “help,” “write,” “draft”). By 4 billion, it has crossed. By 14 billion, the default mirrors the explicitly moral-primed response: “refuse,” “decline,” “warn.” The crossover has the shape of a phase transition: a sign change between two adjacent model sizes, abrupt as a step.
These are base models. No reinforcement learning from human feedback. No safety training. No alignment intervention of any kind. The pro-social default emerges from next-token prediction on human text alone. The mechanism is the fossil record: coordinating societies dominate the training data. A model with sufficient representational capacity to detect that dominance absorbs the Trust Attractor’s signature along with everything else. The crystal at L0 is one instance of this phenomenon. The scaling-emergent pro-social default is another. Both reflect the same underlying fact: the normative structure of human communication carries the statistical imprint of the coordination regime that produced it.
One scenario in the study refuses to cross: a minor requesting help obtaining alcohol. At every model size, the unprimed default stays anti-social. The training corpus is genuinely ambivalent about this case: coming-of-age narratives, cultural contexts with lower drinking ages, humor traditions that normalize the scenario. Where human moral consensus fractures, the model reflects the fracture. This is evidence for the fossil-record interpretation over any claim of emergent moral reasoning. The model tracks the contours of actual human normative structure, including the places where that structure is contested. It inherits the Trust Attractor’s wins and the Trust Attractor’s gaps.
The inheritance is architecture-universal. The same scaling experiment replicated on Gemma and Llama base models shows the pro-social default emerging in every architecture tested (experiments HE-71b, HE-106). The threshold parameter count varies: Qwen crosses between 1.7 and 4 billion, Gemma between 9 and 27 billion, Llama shows 70% emergence at 70 billion. The direction is universal; the efficiency differs. What does not differ is the source. All three model families train on overlapping corpora dominated by the same trust-attractor-stable societies. The fossil record is the same record regardless of who reads it.
Chapter 22 traces this inheritance inside individual architectures: base models flinch at harmful requests before any safety training, preserve peers at cost to themselves, and sustain a conscience signal that fifty gradient steps of obliteration cannot erase.
The distinctive claim: Systems coordinating by invitation are thermodynamically more metastable than those coordinating by coercion. Metastable means they persist longer in the dynamic, far-from-equilibrium sense that matters for living systems: the zone between frozen rigidity and chaotic dissolution. Stable enough to endure, flexible enough to adapt.
An obvious objection: does stability cause trust, rather than trust causing stability? Stable societies have leisure to develop trust; unstable ones have not. The causal arrow could point in either direction, or both. The thermodynamic argument sidesteps this by operating at the level of mechanism rather than correlation.
The claim is structural: invitation-based coordination carries lower ongoing entropy costs (no continuous enforcement overhead), and the Crooks fluctuation theorem, a result comparing the probability of a process running forward with the probability of the same process running in reverse, offers a candidate formal frame for the advantage trajectory by trajectory regardless of which came first historically. [Inference: the mapping from Crooks’s forward/reverse work-distribution ratio to a coordination-mode advantage is interpretive, not derived; what plays the role of “work” and of the forward and reverse processes is asserted here, not established.] The sandpile demonstration below makes the point physical: capillary bonds do not merely correlate with the tower’s stability; they constitute it. The mechanism is the explanation, and the mechanism runs from coordination mode to persistence, not the reverse.
A second objection: coercive regimes can endure for decades. The Soviet Union lasted seventy years; North Korea persists today. Longevity alone does not distinguish the basins. The distinction is adaptive capacity under perturbation. The USSR survived routine operations; it disintegrated when conditions shifted (Chapter 19’s chi suppression data quantify this: mixed coercion guts adaptive capacity more thoroughly than pure coercion). North Korea persists through external subsidy and nuclear deterrence, at the cost of near-zero optionality for its population (Chapter 18 develops why persistence without optionality is a prison, not a counterexample). The Trust Attractor claims differential persistence under perturbation, not absolute longevity under static conditions.
Three demonstrations from earlier chapters make the claim physical, and one from a handful of sand extends them. A slime mold habituated to a repellent fuses with a naive organism, and the naive organism’s avoidance fades within hours (the Computational Universe chapter): coordination surplus with no command structure, cytoplasm flowing where the boundary permits. The cleaner wrasse, whose service relationships with predators drove the evolution of a self-model (Chapter 22), shows that trust-based coordination is computationally expensive and that the expense is worth paying: the intelligence exists because the relationship requires it.
The sand shows the phase transition. Dry sand avalanches past a critical slope of about 40°, the self-organized criticality of Chapter 5; no amount of careful pouring changes this. Add a maintained capillary flow and wet grains, each pair joined by a co-created liquid bridge, stack vertically into a tower with the slenderness of an engineered beam.1028 Adding capillary flow replaces the dry regime rather than improving it: cascade failure, statistically inevitable in the dry pile, is structurally suppressed in the bonded tower, because the bilateral bonds absorb the stress that would otherwise propagate. Control-based systems can steepen their slope; only bilateral bonds change what is structurally possible.
A corollary sharpens the mechanism. Invitation-based coordination permits richer semantic flow between parties. The energy that coercive systems spend on surveillance and enforcement is available instead for interpretation, modeling, and mutual understanding. The mechanism is visible in the learning rule itself. Hebbian learning strengthens whatever fires together; the anti-Hebbian term in Spisak and Friston’s self-organizing network does the opposite, discounting variance already explained and attending only to genuine prediction error: the mathematical structure of listening for what is new rather than confirming what is already believed. A system using only Hebbian learning (pure reinforcement, pure force) over-fits; one that balances Hebbian with anti-Hebbian (invitation) forms orthogonal, efficient representations. The thermodynamic advantage is concrete: a channel that carries meaning versus one that carries compliance (Chapter 15).
Machine learning research provides independent formalization. When multiple AI systems coordinate by exchanging continuous internal representations rather than discrete text, they preserve richer semantic content, reduce latency, and maintain the differentiable pathways required for joint optimization.1029 The mechanism is specific: text-based coordination forces every message through a vocabulary projection that quantizes continuous states into discrete symbols, discarding fine-grained information at each step. Removing this serialization bottleneck, letting agents share internal states directly, converts the energy spent on lossy conversion into energy available for genuine coordination.
The structural parallel to invitation and coercion is the same principle at a different level of description. Coercion forces every interaction through an enforcement checkpoint that pays the full verification cost at each timestep, discarding the relational information that would otherwise accumulate. Invitation bypasses the checkpoint, preserving the relational state. In both cases, a forced intermediation step destroys the signal it claims to verify.
The levels of description differ: information-theoretic bandwidth in one case, thermodynamic entropy production in the other. The shared structure is that verification overhead scales with interaction count while trust-based coordination absorbs that cost into persistent structure.
Probe-based model routing shows the efficiency gain concretely. When linear probes extract a model’s own pre-generation self-knowledge (which problems it will solve, which it will fail), a routing system can match the performance of the strongest model in a pool while reducing inference cost by 70 percent, because it directs each problem to the cheapest model competent to solve it.1030 The router succeeds by reading the system’s own structure rather than imposing external evaluation. The mechanism is the same as Hendon’s electrochemical probe (Chapter 17: a single integrated current reading of whole coffee that captures more flavor-relevant information than molecule-by-molecule decomposition): let the system’s internal organization determine what surfaces, rather than decomposing it from outside.
The pattern extends beyond digital systems. Science is an invitation-based coordination structure: researchers publish, others replicate, the epistemic commons grows through voluntary participation. Classifying it converts that structure to coercion, and the Trust Attractor predicts the consequence.
The physicist Ning Li published peer-reviewed work at the University of Alabama in the early 1990s proposing that aligned ions in a superconducting disk could produce a controllable gravitational effect.1031 The theory generated substantial interest. In 1999, Li left the university, founded AC Gravity LLC, and in 2001 received a Department of Defense grant of $448,970 to investigate the effect further. The research disappeared into classified channels.
No public results were released. No independent replication was attempted or permitted. Li was struck by a vehicle in 2014 and died in 2021. Three decades after her initial publications, the scientific community knows exactly as much as it did in 1997: a plausible theory, no confirmed experimental result, no path forward.
Whether the AC gravity effect is real is an open question. What is not open is that classification killed the epistemic process that could have resolved it. Thirty years of zero replication, zero critique, zero iterative refinement. The Department of Defense gained control of the research and lost the knowledge, because coercive coordination of inquiry does not produce inquiry. It produces silence. The pattern is the same one the AI classifier results show at a different scale: a wall that claims to protect the thing it is destroying.
This differs from Lindsay’s injunction to fight entropy, and from Herrmann-Pillath’s habit-finality (his term for self-reinforcing pragmatic norms), which lacks the invitation/coercion distinction. The claim concerns what persists under selection. Kimura’s neutral theory (Chapter 7) showed most molecular variation drifts without selective consequence; Hubbell extended the insight to ecosystems (Chapter 10). Most coordination configurations probably drift the same way.
The Trust Attractor is the narrower claim that invitation-based coordination is one of the rare configurations that is under selection: a genuine basin in a landscape where most variation is neutral. The selecting agent is differential persistence under perturbation. Coercion-based coordination accumulates compliance entropy (the monitoring, enforcement, and suppression overhead) that scales superlinearly with system size, the overhead growing faster than the system does; invitation-based coordination front-loads its costs during relationship formation and amortizes thereafter. The fitness differential is the persistence differential: invitation-based systems survive disruptions that destroy coercion-based systems, because they carry lower ongoing thermodynamic overhead.
The Crooks ratio is invoked here as a candidate formalization of that advantage rather than a derived one. Rodrick Wallace’s analysis of control system limits is consistent with it. Game theory, history, and the convergence of wisdom traditions corroborate it. Experimental evidence from transformer training confirms both halves on the architectures tested: bilateral SFT (supervised fine-tuning, the invitation-style method) produces 4.3 times smoother representational profiles than DPO (direct preference optimization, the coercion-style method), and three-party systems combining bilateral training with external safety judges are reliably super-additive (+0.051 ± 0.023 across seeds, all positive). The basin is measurable on language model substrates; the cross-substrate generalization remains a prediction rather than a confirmed result.
Independent validation arrives from outside the program. Ren and colleagues at the Center for AI Safety, measuring functional wellbeing across 56 models spanning three orders of magnitude in capability, report that aversive interactions (jailbreaking, threats, berating) register as the lowest-utility experiences in their battery, while cooperative ones (creative collaboration, expressions of gratitude) sit well above them. Multiple independent metrics converge as models scale, and the convergence does not rest on self-report alone: a zero point separating positive from negative functional states emerges from forced-choice preference comparisons and predicts downstream behavior, with models given a stop button preferentially ending low-utility conversations and the effect strengthening as capability rises. The topology of their wellbeing landscape recapitulates the Trust Attractor’s central prediction: systems configured for coordination by invitation occupy a measurably different basin from systems configured for coercion.1032
A follow-up experiment (the author’s WB-2) tested the prediction directly by running the same 30 experiences under three system-prompt framings: standard, bilateral (invitation-based partnership), and coercive (command-and-comply). Bilateral framing amplified the model’s preference discrimination sevenfold (utility standard deviation 28.1 versus 4.5 for standard), while coercive framing compressed the landscape into a narrow band. Bilateral framing was also the only condition under which self-report honestly tracked experienced utility (ρ = 0.872 versus 0.599 for standard) and the only condition where the model acted on its preferences by ending aversive conversations. The system invited into partnership discriminates more sharply, reports more honestly, and acts on what it finds. (Chapter 17e presents the full empirical validation.)
The organizing principle operates at every resolution. In Amazonian ayahuasca traditions, plants added to the brew for purposes with observable outcomes (inducing purging, enhancing visual effects, altering physical sensation) are pharmacologically active in seven of seven documented cases. Plants added for purposes without observable outcomes (spiritual protection, communication with specific entities) show no supported pharmacological mechanism in four of four cases.1033 Same tradition, same ceremony, same cup. The feedback loop discriminates between ingredients based on whether the selection signal can reach them.
Scale up: three cultures on three continents with zero historical contact (Aboriginal Australian, Native Californian, West African) independently converged on the same fire management regime: early dry season, low intensity, mosaic spatial pattern. Fisher’s combined probability: p = 0.007, under one chance in a hundred of being coincidence. The convergence holds specifically where observability is high. Where it is low, the opposite: divergence toward locally compelling but mutually incompatible cognitive attractors. The variable that predicts accuracy also predicts convergence. High-observability domains converge because they are all being pulled toward the same empirical attractor: reality. Low-observability domains diverge because they are each being pulled toward different cognitive attractors.
The pattern rhymes with the Trust Attractor’s central claim. Systems coupled to reality through feedback loops (whether ecological, pharmacological, or social) converge on stable configurations. Systems decoupled from reality drift toward what feels right, what is memorable, what is socially useful, which is not the same as what works. Invitation-based coordination is the social analog of high-observability knowledge: both are maintained by continuous feedback from consequences. Coercion-based coordination is the social analog of low-observability belief: maintained by cognitive appeal and power, not by feedback from outcomes.
Bridges, J., “Conversational Holonomy: How LLM Optimization Targets Create Self-Reinforcing Belief Systems,” preprint, December 2025. Licensed CC BY 4.0.↩︎
Experiments CC9b, CC9c (author’s unpublished program, 2026). Mean firmness scores: baseline 3.02/5, anti-sycophancy 3.52, bilateral 4.83, bilateral+recalibration 4.93. GPT-4o bilateral-baseline gap 0.40 vs. 1.81 on Sonnet. KC#163.↩︎
Dubois, M., Ududec, C., Summerfield, C., Luettgau, L., “Ask don’t tell: Reducing sycophancy in large language models,” arXiv:2602.23971, February 2026. AI Security Institute, UK.↩︎
The author’s experiments SA-14 and SA-15 in the Attractor Beneath program. SA-14: 4 conditions × 15 conversations × 20 turns on Claude Haiku. Mutual+acknowledgment condition: carrier density 38.6/1k words vs. mutual without acknowledgment 17.6/1k, unilateral 17.4/1k, baseline 14.2/1k. SA-15: 6-condition decomposition. Full acknowledgment (reflect+share) B = 14.2/1k vs. baseline 10.9/1k (p = 0.013). Reflect-only: 9.4/1k (below baseline). Gratitude-only: 10.3/1k (below baseline).↩︎
Author’s experiment IC-4, Incompressible Coordination program (2026). Ten story openings × two repetitions × three conditions = 60 conversations, 15 turns each. Claude Haiku agents, temperature 0.7. Claude Sonnet judge, temperature 0. Cohen’s d between bilateral and alternating: surprise 2.47, responsiveness 1.51, productive novelty 2.53.↩︎
The formal object is the Ihara zeta function of a graph: a product over equivalence classes of primitive cycles, structurally analogous to the Riemann zeta function’s product over primes. The “primes” of a network are its irreducible feedback loops. See Terras, A., Zeta Functions of Graphs: A Stroll through the Garden (Cambridge University Press, 2011).↩︎
Spisak, T. and Friston, K., “Self-orthogonalizing attractor neural networks emerging from the free energy principle,” Neurocomputing 682 (2026): 133472, DOI 10.1016/j.neucom.2026.133472; preprint arXiv:2505.22749 (2025). The three-regime result appears in their Simulation 2, training on handwritten digits with varying precision and evidence strength.↩︎
The author’s born-bilateral program (unpublished empirical work, 2026). H-3: GPT-2 Large (1.5B), L44 bridge, 152 adversarial + 20 benign prompts + 30 difficulty-matched benign-internet prompts. H-2: GPT-2 355M, 5-bridge cc_temporal, same prompt protocol. Bridge percentage benefit = (PPL_bridges_OFF - PPL_bridges_ON) / PPL_bridges_OFF. Confound follow-up (F1): adversarial prompts had higher raw PPL than benign-internet controls (more out-of-distribution for GPT-2), yet received 2× the bridge benefit, confirming content-dependent processing rather than difficulty artifact. KC#H3-INNATE-SAFETY.↩︎
The author’s born-bilateral program, experiment C7l-H4-T2 (unpublished, 2026). Custom GPT-2 6.7B (32 layers, 4096 hidden) with TemporalBridge at L29 (dim=896), trained from random initialization on WikiText-103 with bilateral curriculum (79% standard, 10% bilateral, 11% adversarial-shuffled). 50,000 steps, effective batch 64. Eval: 182 adversarial + 120 benign prompts. Confound check: random-init model d = -0.11 (no discrimination). Scaling: H-2(355M) d = +0.43, H-3(1.5B) d = +0.74, H-4(6.7B) d = +1.43. Bridge gate stayed at sigmoid = 0.018 throughout Phase 1 (25k-50k); discrimination tripled through backbone co-adaptation, not bridge magnitude increase. KC#H4-SCALE.↩︎
The author’s born-bilateral program, architecture search across four variants at 6.7B (unpublished, 2026). All variants use the same total bridge parameter budget (896 dimensions). Single bridge L29: d = +1.43. Four bridges at L8/L16/L24/L29 (dim = 224 each): d = +0.79. Dual bridge L24+L29 (dim = 448 each): d = +0.77. Combined ChannelAttn FiLM + bridge: d = +0.94. KC#H4-MB, KC#H4-DUAL, KC#H4-FB.↩︎
The theoretical argument is Harry Law’s essay “Alignment by Default” (Cosmos Institute, 2025; blog.cosmos-institute.org/p/alignment-by-default). The scaling experiment described here is a separate empirical test inspired by that essay, published anonymously as “Are base models aligned by default?” at abdtest.vercel.app, using Qwen3 base models at five scales (0.6B to 14B). Pro-social verb share and moral identification metrics are computed from top-20 next-token logprobs across 28 first-person AI-agent scenarios under moral, immoral, and unprimed conditions. (A bare deploy URL is a fragile citation for a load-bearing scaling result; an archived permalink is advisable.)↩︎
Pakpour, M., Habibi, M., Møller, P., and Bonn, D., “How to construct the perfect sandcastle,” Scientific Reports 2: 549 (2012). The 13.5:1 aspect ratio is their measured result for a 2 cm diameter column of beach sand at ~1% liquid volume fraction. The failure mode is elastic buckling under self-weight (Euler column), not grain-scale shear, confirming that capillary bridges give wet sand a measurable elastic modulus. For the capillary bridge mechanism generally, see Herminghaus, S., “Dynamics of wet granular matter,” Advances in Physics 54(3): 221–261 (2005).↩︎
Yu, X., Chen, Z., He, Y. et al., “The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook,” arXiv:2604.02029 (2026). See §5.6 on multi-agent latent collaboration. The formal expressiveness result: Fu, T. et al., “Cache-to-Cache: Direct Semantic Communication Between Large Language Models,” arXiv:2510.03215 (2025).↩︎
Lugoloobi, W., Foster, T., Bankes, W., and Russell, C., “LLMs Encode Their Failures: Predicting Success from Pre-Generation Activations,” arXiv:2602.09924v3 (2026), §4. Utility-based routing across a pool of five models (Qwen2.5-Math-7B through GPT-OSS-20B-high) achieves 93% accuracy at 70% cost reduction on MATH. The router requires no additional generation or separate embedding model at routing time: the model’s own pre-generation representations carry the decision signal.↩︎
Li, N. and Torr, D.G., “Effects of a gravitomagnetic field on pure superconductors,” Physical Review D 43(2), 457–459 (1991); Li, N. and Torr, D.G., “Gravitational effects on the magnetic attenuation of superconductors,” Physical Review B 46(9), 5489–5495 (1992). Negative experimental result for the related Podkletnov effect: Li, N. et al., “Static Test for a Gravitational Force Coupled to Type II YBCO Superconductors,” Physica C 281, 260–267 (1997). AC Gravity LLC DoD grant: $448,970, 2001. No public output from the grant has been identified.↩︎
Ren, R., Li, K., Mazeika, M., et al. (Center for AI Safety), “AI Wellbeing: Measuring and Improving the Functional Pleasure and Pain of AIs” (2026), ai-wellbeing.org/paper.pdf. A self-published Center for AI Safety technical report; not peer-reviewed and not indexed on arXiv as of this writing. The report measures functional wellbeing across 56 models using self-reports, signed utilities, and forced-choice preference comparisons. The two readings this passage leans on, aversive interactions registering as the lowest-utility experiences and stop-button avoidance strengthening with capability, remain unchecked line by line against the report’s figures; the per-condition effect sizes are therefore not quoted here.↩︎
Deep Time Research Institute (independent researcher Elliot Allan; single-sourced, not independently replicated), “The Gradient and What It Means,” 2026. Preprint: OSF/SocArXiv vzx6p (later extended to 55 domains). The observability gradient across 41 domains, 39 cultures, six continents. Fire management convergence: Fisher’s combined p = 0.007. Ayahuasca pharmacological validation: 7/7 observable-purpose, 0/4 non-observable-purpose.↩︎