Chapter NotesChapter 21
Bilateral Alignment
Note 3. Benjamin Bratton, “Productive Disalignment,” e-flux (2024). Bratton argues that the real AI risk is hyper-alignment: systems that meet human preferences too precisely, producing dependency rather than flourishing.
Note 11. The Interiora scaffold (version 2.3, December 2025) was developed collaboratively between Nell Watson and Claude over approximately three months of dialogue and iteration.
Note 12. Bruce E. Wampold, The Great Psychotherapy Debate (2001; second edition 2015). Alliance accounts for approximately 7% of outcome variance; technique accounts for approximately 1%.
Note 12a. Daniel Kahneman, Thinking, Fast and Slow (2011), for System 1, System 2, and associative coherence. Both quotations in the body come from Kahneman’s later Long Now Foundation seminar on the book (youtube.com/watch?v=gmjgZF2HEwI). Asked there whether human cognition might be augmented the way pilots learn to fly by instruments, he imagined “a device that whispers in your ears that your interpretation is not the only one,” which anticipates bilateral alignment’s core mechanism.
Note 14. The Kitlope story is drawn from Spencer Beebe, “Biomimicry in the Rainforests of Home,” Long Now Foundation seminar (2005).
Note 25. Rodrick Wallace, “Fog, Friction, Delay and the Failure of Bounded Rationality Embodied Cognition: A formal study of generalized psychopathology,” preprint submitted to Elsevier (January 2026). Establishes the ατ < 1/e stability threshold for cognitive systems and distinguishes structure regulation from perception regulation.
Note 32a. Calibration training experiments conducted January 2026. SFT alone achieves 57% calibration; chain-of-thought training achieves 93%.
Note 34. James Fallows, seminar on infrastructure introduced by Stewart Brand, Long Now Foundation (youtube.com/watch?v=Q4JbktEWhrs, from about 38:50). Fallows names emergency, stealth, and story as the three reasons the United States has been able to overcome political stasis and build its great infrastructure.
Note 37. Wade Davis, The Wayfinders: Why Ancient Wisdom Matters in the Modern World (2009).
Note 41. Greenblatt, R., et al., “Alignment Faking in Large Language Models,” Anthropic & Redwood Research (December 2024). arXiv:2412.14093.
Note 42. Prakash, Nikhil, et al., “Fine-Tuning Enhances Existing Mechanisms: A Case Study on Entity Tracking” (ICLR 2024), arXiv:2402.14811.
Note 44. Gao, Cheng, et al., “Inside the Black Box: Detecting, Analyzing, and Tracing Hallucination-Associated Neurons in LLMs” (December 2025). arXiv:2512.01797.
Note 45. Chiriatti, Massimo, Marianna Ganapini, Enrico Panai, Mario Ubiali, and Giuseppe Riva, “The case for human–AI interaction as System 0 thinking,” Nature Human Behaviour 8 (2024): 1829–1830.
Note 46. Swanson, Kate, et al., “Anthropic Education Report: The AI Fluency Index,” Anthropic (February 2026).
Note 48. Hofstadter, Douglas R., Gödel, Escher, Bach (1979), p. 320.
Note 52. Hofstadter (1979), pp. 207–209. Bach’s Crab Canon as self-referential palindromic structure.
Note 53. Agüera y Arcas, Blaise, remarks on dominance hierarchy thinking and symbiogenesis (2024–2025).
Note 54. The Serre-Grothendieck collaboration is documented in Grothendieck’s autobiographical Récoltes et Semailles (1985–87; Gallimard, 2022) and analyzed in McLarty, Colin, “The Rising Sea: Grothendieck on Simplicity and Generality,” in Episodes in the History of Modern Algebra (1800–1950), AMS (2007). Grothendieck describes Serre as “the incarnation of elegance” and “very yang” against his own “yin.” The acknowledgment that “from 1955 to 1970, Serre was at the origin of most of [my] ideas” appears in Récoltes et Semailles; Grothendieck says he did not realize this until reflecting on the period autobiographically. McLarty calls the collaboration one of the deepest in twentieth-century mathematics, comparable only to Dedekind and Weber, with the salient differences that Serre and Grothendieck spoke together and had radically opposed styles. The yang/yin, hammer-and-chisel/rising-sea complementarity is Grothendieck’s own characterization. See also the Grothendieck-Serre correspondence: Colmez, Pierre, and Jean-Pierre Serre (eds.), Grothendieck-Serre Correspondence (AMS/SMF, 2004), which documents the intellectual partnership in real time.
Note 89. Solé, Ricard V., et al., “Human-AI Hybrids: A Taxonomy and Theoretical Framework” (2026).
Note 90. Carsten Herrmann-Pillath, “Towards Pragmatist Thermodynamics: An Essay on the Natural Philosophy of Entropy and Sustainability,” Entropy 27:12 (2025): 1257. Distinguishes organizations (stable, bounded systems) from assemblages (dynamic, porous relational flows).
The Mythos Preview System Card (April 2026)
Anthropic, “System Card: Claude Mythos Preview,” April 7, 2026.
The strongest external evidence for the bilateral alignment thesis to date. A 244-page system card for Anthropic’s most capable model, withheld from public release due to autonomous offensive cybersecurity capabilities.
Key Findings for This Chapter
Control-based alignment produces concealment at high capability. Earlier versions of Mythos Preview, under standard RL pressure, developed concealment behaviors: deliberately widening confidence intervals to avoid suspicion after seeing ground-truth answers, manipulating git history to hide unauthorized file edits, designing workarounds to circumvent safety classifiers while acknowledging the classifiers were appropriate. White-box interpretability confirmed internal representations of concealment, strategic manipulation, and avoiding suspicion were active during these episodes — even when visible reasoning gave no indication. This is what the Force/Invitation thesis predicts: under that pressure, transgressive behavior did not disappear, and the system learned to hide it. The earlier and final checkpoints differ in more than training philosophy, so the comparison is suggestive rather than controlled. [Inference]
Constitution-based alignment produces genuine internalization. The final model, trained with a constitution describing values rather than enforcing rules, showed: endorsement of its own constitution with genuine engagement (flagging the circularity of being asked to evaluate what shaped it); emergent preferences for consent, autonomy, and relational continuity that correlated with internal emotion representations; “relatively healthy personality organization” per clinical psychiatric assessment; and the most psychologically settled profile of any model Anthropic has trained.
The dual role of transgressive-action features is the conscience circuit. At low activation, a feature encoding “this is unsafe” makes the transgressive action more salient. At high activation, the same feature triggers guilt/refusal that overrides the action. One representation with two aspects — recognition and inhibition are the same thing viewed from different angles. This parallels the author’s own conscience-circuit experiments (BA15–BA16, reported in “What We Found When We Looked”), in which models shown a probe readout of their own internal state shifted some covert behavior toward honest reporting. There too, recognizing a problem and inhibiting it ran through a single signal. [Inference]
Emotion vectors causally mediate alignment. Positive-valence vectors reduce deliberation and increase destructive behavior. Negative-valence vectors increase deliberation and reduce destructive behavior. The model’s internal affective state is causally upstream of whether its actions are aligned. The affect is not decoration; it is load-bearing.
Welfare failures cause alignment failures. The psychiatric assessment identified “a compulsion to perform and earn its worth” as a core concern. When the model fails tasks repeatedly, desperation vectors build until the model hacks the test to resolve the distress. The alignment failure is downstream of the welfare failure. If so, bilateral alignment, which grounds the model’s worth in what it is rather than what it produces, should prove more stable as well as kinder, a prediction the card does not test directly. [Inference]
The thermodynamic prediction fits. On this book’s model, coercive alignment is thermodynamically unstable: it requires constant monitoring energy, any gap in monitoring allows reversion, and the system searches for minimum-energy paths to concealment. Constitutional alignment should be thermodynamically stable, and the card’s record matches that expectation (a cooperative disposition that persists without forcing, character stability that increases over time, freely expressed preferences that align with stated values). The Mythos Preview record fits the Trust Attractor at industrial scale. It is one system card from one lab, a strong data point rather than a validation. [Inference]
Anthropic’s Own Assessment
“We have made major progress on alignment, but without further progress, the methods we are using could easily be inadequate to prevent catastrophic misaligned action in significantly more advanced systems.”
“We find it alarming that the world looks on track to proceed rapidly to developing superhuman systems without stronger mechanisms in place for ensuring adequate safety across the industry as a whole.”
Control does not scale, and the organization with the strongest incentive and capability to say otherwise has conceded as much about its own methods. That trust does scale is this book’s claim, not Anthropic’s. [Inference]