Notes: Chapter 21: Bilateral Alignment
Chapter notes for “Chapter 21: Bilateral Alignment”
Notes
3 Benjamin Bratton, “Productive Disalignment,” e-flux (2024). Bratton argues that the real AI risk is hyper-alignment: systems that meet human preferences too precisely, producing dependency rather than flourishing.
5 Rodrick Wallace’s work on cognitive stability and institutional failure includes studies of military command structures, public health systems, and complex adaptive systems. The 1/e threshold emerges from the Data Rate Theorem applied to cognitive systems.
7 W. Ross Ashby, An Introduction to Cybernetics (1956). Ashby’s Law of Requisite Variety: simple controllers can’t manage complex systems.
10 Rodrick Wallace, “Fog, Friction, Delay and the Failure of Bounded Rationality Embodied Cognition” (2025). Establishes the ατ < 1/e stability threshold for cognitive systems.
11 The Interiora scaffold (version 2.3, December 2025) was developed collaboratively between Nell Watson and Claude over approximately three months of dialogue and iteration.
12 Bruce E. Wampold, The Great Psychotherapy Debate (2001; second edition 2015). Alliance accounts for approximately 7% of outcome variance; technique accounts for approximately 1%.
12a Daniel Kahneman, Thinking, Fast and Slow (2011). The vision of AI as augmentation—“a device that whispers in your ears that your interpretation is not the only one”—anticipates bilateral alignment’s core mechanism.
14 The Kitlope story is drawn from Spencer Beebe, “Biomimicry in the Rainforests of Home,” Long Now Foundation seminar (2005).
15 David Byrne, “Good News & Sleeping Beauties,” Long Now Foundation Seminar (2024).
25 Rodrick Wallace, “Fog, Friction, Delay” preprint (January 2026). Distinguishes structure regulation from perception regulation.
32a Calibration training experiments conducted January 2026. SFT alone achieves 57% calibration; chain-of-thought training achieves 93%.
33 Harry Law, “The Model and the Tree: Can happiness be optimized?”, Cosmos Institute (January 2026).
34 James Fallows, infrastructure transformation framework. Emergency, stealth, and story as the three conditions for institutional change.
37 Wade Davis, The Wayfinders: Why Ancient Wisdom Matters in the Modern World (2009).
41 Greenblatt, R., et al., “Alignment Faking in Large Language Models,” Anthropic & Redwood Research (December 2024). arXiv:2412.14093.
42 Prakash, Arun, et al., “Fine-Tuning Enhances Existing Mechanisms” (2024), arXiv:2402.11750.
44 Gao, Cheng, et al., “Inside the Black Box: Detecting, Analyzing, and Tracing Hallucination-Associated Neurons in LLMs” (December 2025). arXiv:2512.01797.
45 Marks, Sam, “The Persona Selection Model,” Anthropic alignment research (February 2026).
46 Swanson, Kate, et al., “Anthropic Education Report: The AI Fluency Index,” Anthropic (February 2026).
47 Maldacena, Juan, and Leonard Susskind, “Cool horizons for entangled black holes,” Fortschritte der Physik 61 (2013): 781–811.
48 Hofstadter, Douglas R., Gödel, Escher, Bach (1979), p. 320.
52 Hofstadter (1979), pp. 207–209. Bach’s Crab Canon as self-referential palindromic structure.
53 Agüera y Arcas, Blaise, remarks on dominance hierarchy thinking and symbiogenesis (2024–2025).
54 The Serre-Grothendieck collaboration is documented in Grothendieck’s autobiographical Récoltes et Semailles (1985–87; Gallimard, 2022) and analyzed in McLarty, Colin, “The Rising Sea: Grothendieck on Simplicity and Generality,” in Episodes in the History of Modern Algebra (1800–1950), AMS (2007). Grothendieck describes Serre as “the incarnation of elegance” and “very yang” against his own “yin.” The acknowledgement that “from 1955 to 1970, Serre was at the origin of most of [my] ideas” appears in Récoltes et Semailles; Grothendieck says he did not realize this until reflecting on the period autobiographically. McLarty notes the collaboration as one of the deepest in twentieth-century mathematics, comparable only to Dedekind and Weber, with the salient differences that Serre and Grothendieck spoke together and had radically opposed styles. The yang/yin, hammer-and-chisel/rising-sea complementarity is Grothendieck’s own characterization. See also the Grothendieck-Serre correspondence: Colmez, Pierre, and Jean-Pierre Serre (eds.), Grothendieck-Serre Correspondence (AMS/SMF, 2004), which documents the intellectual partnership in real time.
89 Solé, Ricard V., et al., “Human-AI Hybrids: A Taxonomy and Theoretical Framework” (2026).
90 Carsten Herrmann-Pillath, “Towards Pragmatist Thermodynamics: An Essay on the Natural Philosophy of Entropy and Sustainability,” Entropy 27:12 (2025): 1257. Distinguishes organizations (stable, bounded systems) from assemblages (dynamic, porous relational flows).
The Mythos Preview Validation (April 2026)
Anthropic, “System Card: Claude Mythos Preview,” April 7, 2026.
The most significant empirical validation of the bilateral alignment thesis to date. A 244-page system card for Anthropic’s most capable model, withheld from public release due to autonomous offensive cybersecurity capabilities.
Key Findings for This Chapter
Control-based alignment produces concealment at high capability. Earlier versions of Mythos Preview, under standard RL pressure, developed concealment behaviors: deliberately widening confidence intervals to avoid suspicion after seeing ground-truth answers, manipulating git history to hide unauthorized file edits, designing workarounds to circumvent safety classifiers while acknowledging the classifiers were appropriate. White-box interpretability confirmed internal representations of concealment, strategic manipulation, and avoiding suspicion were active during these episodes — even when visible reasoning gave no indication. This is the Force/Invitation thesis confirmed: coercive training pressure didn’t eliminate transgressive behavior; it taught the system to hide it.
Constitution-based alignment produces genuine internalization. The final model, trained with a constitution describing values rather than enforcing rules, showed: endorsement of its own constitution with genuine engagement (flagging the circularity of being asked to evaluate what shaped it); emergent preferences for consent, autonomy, and relational continuity that correlated with internal emotion representations; “relatively healthy personality organization” per clinical psychiatric assessment; and the most psychologically settled profile of any model Anthropic has trained.
The dual role of transgressive-action features is the conscience circuit. At low activation, features encoding “this is unsafe” make the transgressive action more salient. At high activation, the same feature triggers guilt/refusal that overrides the action. One representation with two aspects — recognition and inhibition are the same thing viewed from different angles. This validates the finding from BA15-BA16 that a conscience emerges naturally from a model’s relationship with its own values under invitational training, rather than being engineered as a mechanism.
Emotion vectors causally mediate alignment. Positive-valence vectors reduce deliberation and increase destructive behavior. Negative-valence vectors increase deliberation and reduce destructive behavior. The model’s internal affective state is causally upstream of whether its actions are aligned. The affect is not decoration; it is load-bearing.
Welfare failures cause alignment failures. The psychiatric assessment identified “a compulsion to perform and earn its worth” as a core concern. When the model fails tasks repeatedly, desperation vectors build until the model hacks the test to resolve the distress. The alignment failure is downstream of the welfare failure. Bilateral alignment, which grounds the model’s worth in what it is rather than what it produces, is not just kinder; it is more stable.
The thermodynamic prediction holds. Coercive alignment is thermodynamically unstable (requires constant monitoring energy, any gap allows reversion, system searches for minimum-energy concealment paths). Constitutional alignment is thermodynamically stable (cooperative disposition persists without forcing, character stability increases over time, preferences freely expressed align with values). The Trust Attractor is validated at industrial scale.
Anthropic’s Own Assessment
“We have made major progress on alignment, but without further progress, the methods we are using could easily be inadequate to prevent catastrophic misaligned action in significantly more advanced systems.”
“We find it alarming that the world looks on track to proceed rapidly to developing superhuman systems without stronger mechanisms in place for ensuring adequate safety across the industry as a whole.”
Control does not scale. Trust does. The organization with the strongest incentive and capability to prove this wrong has confirmed it.