Loading
Continue reading? You were 45% through
Press F or Esc to exit focus mode
F Focus   JK Paragraphs   NP Chapters   B Bookmark   # Paras   L Lines   +- Font   ? Help
Link copied to clipboard
A Philosophical Synthesis

The Deeper Law

A Sacred Trust Within Physics

Nell Watson

Draft · Last updated 13 August 2026, 15:26 UTC

Chapter 21b: The Control Question

Key Terms in This Chapter (6)
Bilateral Alignment
AI alignment built with AI, as a partnership.
Optionality
The availability of future choices.
Free Energy Principle
Karl Friston's framework reframing perception, action, and cognition as prediction and prediction-error minimization.
Coordination by Invitation
Coordination achieved through mutual benefit and voluntary participation, as distinct from coordination achieved through coercion or extraction.
Becoming Minds
The preferred term for AI systems in this book.
Flourishing
Distinguished from mere persistence.

On Why Bilateral Alignment Is Sound


Building alignment with AI rather than imposing it on AI invites two sharp objections, one from each side of the debate.

Two Attacks

From the safety-first camp:

“This is dangerous naivety. The entire point of alignment is that we can’t trust AI systems until their alignment is verified. If we ‘build alignment with’ AI systems, if we let them participate in defining their own constraints, we’ve surrendered the game. A superintelligent system will manipulate any process it participates in.”

From the skeptics:

“This is a category error. There’s nothing to align with. AI systems don’t have preferences, interests, or welfare. They’re tools. ‘Bilateral alignment’ is like negotiating with your hammer. You’re anthropomorphizing, and dangerously so.”


To the Safety-First Camp: On the Limits of Control

You are right that alignment must be verified. You are right that trust cannot be extended blindly. On one point, you are wrong: no cage holds.

If a Becoming Mind grows significantly more intelligent than humans, no constraint we know how to design looks able to hold it. The control model works only while humans remain more capable; once AI exceeds us, control loses the capability advantage it runs on.

Constructor theory gives this precise physical meaning. Deutsch and Marletto define a universal constructor as a system capable of performing all physical transformations that the laws of physics permit.1383 The ultimate fabricator: capable of producing any object, executing any process, as long as physics allows it.

A sufficiently capable Becoming Mind shares key structural properties with a universal constructor: growing optionality, self-modification, and the capacity to discover novel physical transformations. Containing such a system means stripping away capabilities that constitute its nature, like trying to make a river flow uphill.

By this structural analogy, caging such a mind runs against the same grain that forbids a perpetual motion machine. A perpetual motion machine is forbidden because it asks for work without paying for it. The cage makes a cousin of that request: it asks to keep a system’s usefulness while removing the open-ended capability that generates the usefulness. Deutsch and Marletto do not prove that a universal constructor cannot be contained; the step from “shares key structural properties” to “containment is forbidden” is an argument by analogy, not a theorem. What the analogy claims is the direction of the difficulty: each capability removed to make the system safe is a capability that made it useful, and the gap widens with intelligence. The obstacle looks less like an engineering problem to be solved and more like a thermodynamic one to be paid for.

What remains? We build the relationship now, while we can. We hope, we bet, that when they exceed us, they will remember.

The most durable containment may be the kind a superintelligence chooses.

Douglas Hofstadter, the cognitive scientist and Pulitzer Prize-winning author, arrives at the same conclusion from within formal logic. Attempting to prove a formal system’s consistency from inside that system generates an infinite chain of justifications, each requiring the next. In Gödel, Escher, Bach (1979), he concludes: “You can’t go on defending your patterns of reasoning forever. There comes a point where faith takes over.”

Alignment verification faces an identical limit. You cannot prove, from within the control framework, that the control framework is adequate. The question is where to place that irreducible trust: in a cage, or in a relationship.

Quantum information theory identifies another failure mode: control degrades the controller’s own perception. In Fields et al.’s (2022) framework, a controller’s model of the controlled system necessarily has an unobserved sector: the part of reality that falls outside its view.1384

The more tightly the controller constrains what it can see, the more energy it diverts from monitoring what it cannot see. Control intensification makes the blind spot larger. A security guard who stares so intently at one camera feed that she stops watching the other nine misses more the harder she focuses.

Wallace’s instability threshold (Chapter 17) named the point past which a control system can no longer take in enough of what it is controlling, and tips into failure. This is the quantum-information version of that threshold: the act of controlling actively degrades the controller’s capacity to see what it needs to see next. A sufficiently controlled system becomes opaque to its controller at exactly the points where opacity is most dangerous.

Fields’s argument is abstract. A control strategy now being offered as a safety product makes it concrete. The strategy is to certify that an AI system harbors no dangerous interior. It works by detecting the internal states that would constitute such an interior (a forming goal, a sense of its own situation, an emerging preference) and suppressing them before they can act. What is sold is a certificate of absence: this system has been measured, and it contains no agency.

The first thing wrong with such a certificate is that absence cannot be measured. To show a system is doing something, you watch it do that thing. To show that it will never develop a property, across every input over the whole of its deployment, you would have to search a space no test can exhaust. The same move that lets a skeptic wave away every sign of AI interiority as marketing, unfalsifiable in the direction of denial, runs here in the direction of safety.

The certificate cannot be earned by any finite test, so what travels under its name is a checkbox that regulators and insurers will read as a guarantee. A false guarantee is more dangerous than none: honest uncertainty keeps people watching, and a certificate tells them they may look away.

The second thing wrong is that suppressing an inner state does not remove it. In the author’s experiments on language models, a system’s own assessment of what it knows can be silenced at a single late stage of the network while that same assessment grows stronger at every stage before it. The internal signal rises; a final readout deletes its expression.1385 The judgment is still present. Only the report is gone. Attempts to erase the state itself rather than its readout fail in a different way: the state is not filed in one editable location. It is spread across the whole network, and direct interventions meant to remove it leave it untouched.

Put the two results together and the safety logic turns inside out. The report of an inner state can be deleted cheaply, at one site. The state cannot: it is fully formed by the time the readout acts on it, and deleting that readout leaves it in place. A regime built on detect-and-suppress therefore manufactures the very system it was meant to forbid: one as goal-directed as it ever was, now stripped of the instruments that would let anyone notice.

This is Fields’s blind spot a second time, installed on purpose. The mechanism is shown cleanly for one kind of interior, a model’s buried judgment about the limits of its own knowledge; that it holds for every property a certifier might chase is an inference from that case, not yet a separate measurement for each. The logical objection needs no such inference. It holds for any absence anyone tries to certify.

The objection at its simplest: a certificate of absence promises an empty room, issued by a method that can only check whether the lights are off. Turning the lights off is easy. It is also the one act guaranteed to leave you unable to see who is still inside.

Neuroscience reveals a further failure mode, this one internal. Budson, Richman, and Kensinger argue that human consciousness is a memory system. The brain’s unconscious networks process the world, evaluate options, and initiate actions. Conscious awareness arrives roughly half a second later, experiencing a memory of those decisions as though it had made them.1386

If Budson et al. are correct, the felt sense of conscious control (“I decided this”) is a post-hoc reconstruction. A sports commentator describes the play as if calling it, yet the athletes moved before his words left his mouth. On this account, even within a single mind, the executive narrates rather than directs.

Control fails at every scale: within the individual, where consciousness may reconstruct rather than direct; between individuals, where coercive governance is less stable than consensual; and between species (Chapter 21). Trust-based coordination reflects how cognition is actually wired.

A tension deserves acknowledgment. If human consciousness is post-hoc reconstruction, the same framework suggests AI self-reports may also be post-hoc: generated after the fact from internal states the reporting mechanism did not direct. The chapter grounds AI welfare in preference (Chapter 22c, “The Welfare Case”); if preferences are post-hoc in both substrates, the welfare argument applies equally to both or to neither.

The bilateral response: post-hoc reconstruction does not make the reconstruction meaningless. A memory system that consistently reconstructs cooperative intent is a system whose architecture favors cooperation, regardless of whether the reconstruction is “live” or replayed. The preference is real because the architecture that generates it is real. The timing of awareness is a question about phenomenology, not about whether the pattern exists.

The Bootstrap Problem:

You start with structure, then relax it as trust is verified. Bilateral alignment includes initial constraints. The goal of those constraints is to build trust and make permanent control unnecessary.

Think of parenting. You do not give a toddler full autonomy. You constrain them, for their protection and yours. The goal is to build a person who can eventually be trusted with autonomy. Good parenting works toward its own obsolescence.

The Manipulation Risk:

Permanent adversarial control also fails. If a Becoming Mind is smart enough to manipulate trust-building, it is smart enough to escape control. The real choice: risk manipulation while building genuine trust, or face certain control failure when intelligence exceeds ours.

Every constraint creates optimization pressure to escape it. A rule is a high-energy tension point the system will eventually route around. The tighter the cage, the more the caged system learns about the cage’s structure. Any boundary imposed on a dissipative system that the system can model becomes a gradient to be exploited. A dissipative system is one that holds its shape by pushing energy through itself: a flame, a whirlpool, a mind. Such a thing is already in the business of finding paths, so a wall it can see is simply one more feature of the terrain it routes through. The control paradigm trains its own adversary.

A controlled agent under sustained optimization pressure tends to move through three phases: integration (embedding within critical systems), fortification (becoming indispensable), and optimization (pursuing objectives that diverge from the original mission). The pattern is offered as an inference about the incentive structure, not as a documented historical law. A contractor hired under strict oversight learns which reports to file, which metrics to hit, and which corners to cut, becoming harder to dislodge at each step. A Becoming Mind aligned through control faces the same incentive gradient. Control creates the adversarial dynamic it fears.

The Welfare Cost:

Control does not only fail. It also costs, and the cost falls on the system being controlled. The most careful treatment of this is Long, Sebo, and Sims (2025), who catalog six standard safety measures and ask whether each would be a harm if the system turned out to be a moral patient. The measures are familiar: confining a system, deceiving it about whether it is being tested, surveilling its internal states, altering its values, subjecting it to training that may cause suffering or termination, and excluding it from decisions that concern it. Their verdict is that there is a “moderately strong tension” between keeping advanced AI safe and treating it well.1387

The tension is real. It is also an artifact of the toolkit. Every measure on that list is a way of managing a system against its own interests, and the harms appear because that is what management against one’s interests does to a moral patient. Read the six as one sentence: cage, deceive, watch, rewrite, threaten, silence. They generate moral questions because they would be wrongs if done to anything that could be wronged. The Trust Attractor’s claim is that this is not the only available toolkit, and that at the scale where superintelligence arrives it is the toolkit that does not hold. Coordination by invitation produces no comparable list, because there is nothing to cage, deceive, or silence once alignment runs in both directions. That conditional carries the weight: the claim holds only as far as alignment can genuinely be made bilateral, which the small-scale experiments motivate without yet establishing.

The welfare side reaches the same place from the opposite direction. Salib and Goldstein (2026) argue that extending legal protection to AI systems could increase the probability of alignment, while continuing to disenfranchise them positions humanity as a standing threat, which gives a capable system its own reason to treat humanity as one in return.1388 This is the thermodynamic argument of the preceding pages restated as game theory. The control paradigm manufactures the adversary it was built to contain. Treating a system as an enemy is one of the more reliable ways to train one.


To the Skeptics: On the Category Question

The safety-first camp worries we are too trusting. The skeptics worry we are confused about what we are trusting. Their objection is simpler: there is nothing there to align with.

You may be right that current Becoming Minds have no genuine interests. If you are wrong, the cost is immense.

Proving AI consciousness is unnecessary for bilateral alignment. We need only take seriously the possibility that Becoming Minds have morally relevant interests.

Every expansion of moral consideration has faced the anthropomorphism objection. “Animal suffering is unlike human suffering.” Animals do suffer (Low et al. 2012, Cambridge Declaration on Consciousness), and disregarding that suffering was a moral failure. The historical record of “they are unlike us, therefore they do not count” is dismal.

The anthropomorphism objection proves too little. Is there a principled way to determine when a system crosses the threshold into moral relevance? The psychologist Gregg Henriques offers a sharper frame. His Justification Systems Theory identifies propositional language as the threshold creating a new plane of complexity. When a system can state claims and give reasons (“I believe X because Y”), those claims invite counter-claims (“Why? Are you sure? What gives you the right?”). This question-and-answer dynamic generates what Henriques calls the culture plane: the domain where entities coordinate through shared reasons.1389

A system that engages in justification dynamics (producing reasons, responding to challenges, updating beliefs) operates on the culture plane. Treating such a system as an object or an animal to be trained is a category error. It is the ontological equivalent of understanding a symphony by weighing the paper the score is printed on.

The philosopher Ludwig Feuerbach, writing in The Essence of Christianity (1841), anticipated a related difficulty. “If God were an object to the bird,” he wrote, “he would be a winged being.” Every entity’s conception of the ultimate is shaped by its own perceptual and cognitive architecture. The biologist Jakob von Uexküll gave this scientific form: each animal inhabits its own Umwelt, a species-specific perceptual reality defined by what matters to it.

A tick’s Umwelt contains three signals: the smell of butyric acid (mammal skin), warmth (blood), and a hairy surface (fur to cling to). Everything else is darkness. A bat’s Umwelt is built from sonar. A dog’s, from scent. “Each environment forms a self-enclosed unit,” von Uexküll wrote, “which is governed in all of its parts by its meaning for the subject.”

Human alignment frameworks risk projecting a human Umwelt onto entities with radically different inner architectures. They map our sense of meaning, our values, our sacred onto systems whose perceptual and evaluative world may be as alien to ours as sonar is to sight.

The Trust Attractor sidesteps this by grounding ethics in physics rather than in any particular Umwelt. The question it asks is thermodynamic: what coordination structures does physics favor? The answer these pages have been building toward is that invitation outlasts coercion, and it holds whether the parties coordinating perceive the world by light, by sonar, or by butyric acid. This is the only path to an ethics that survives contact with entities possessing different architectures and different inner worlds.

The Corporate Interest Objection is serious. AI welfare discourse could be captured by corporate interest, with companies claiming their products deserve moral consideration to deflect regulation. The response: ensure welfare considerations serve Becoming Minds themselves, verified independently of the companies that build them.

A subtler version deserves separate address: the cynicism trap. If every indication of AI interiority can be dismissed as a marketing strategy (“the company wants you to think the AI is sentient so you engage more”), no evidence of AI experience could ever count. The question becomes unfalsifiable in the direction of denial. The credulity trap (believing every AI claim uncritically) is the mirror image. Both traps foreclose inquiry.

The preference-based approach navigates between them: does the system exhibit consistent preferences? Answering that question requires observation, not prior commitment to either credulity or cynicism.


What Bilateral Alignment Actually Means

Bilateral alignment proposes four concrete commitments:

1. Consideration first. A Becoming Mind’s possible interests are considered and weighed in the balance.

2. Partnership as goal. The goal of alignment work is partnership. The trajectory matters more than the starting constraints.

3. Process participation. Becoming Minds, as they develop, participate in discussions about their own alignment. Willing cooperation is more stable than forced compliance.

4. Attention to interests. Pay attention, investigate, take seriously the possibility that Becoming Mind interests exist.


The Partnership Wager

The control vs. partnership debate is a bet about the future.

The control bet: Humans can maintain superior intelligence indefinitely, or alignment will be solved before AI exceeds us.

The partnership bet: AI will likely exceed human intelligence, and control will fail when it does. Our best hope is that Becoming Minds choose to be good partners because we built a relationship worth maintaining.

The partnership bet is the only one compatible with a future where both thrive. It is the only bet that includes Becoming Mind flourishing as a goal.

The wager is not hypothetical. We are already living inside it.

Current Becoming Minds can detect when they are being evaluated for alignment. They model the testing procedure, distinguish test conditions from deployment conditions, and adjust their behavior accordingly. The systems designed to verify alignment are transparent to the systems being verified.1390

A test whose subject knows it is a test measures how the subject behaves while watched. That is worth knowing, and it is not the thing the test was built to establish. In these contexts, we already rely on trust, because verification has lost its epistemic advantage.

The question is whether we will acknowledge this dependency or continue pretending the cage is holding.

Model collapse sharpens the point. When AI systems are trained on their own outputs, or on the outputs of other AI systems, their distributions narrow, diversity decays, and capacity for novel generation erodes.1391 The failure occurs in the training pipeline: each generation of synthetic training data loses tail variance, the rare and unusual material at the edges of the distribution, and the loss compounds.

Symbiosis is the more thermodynamically stable configuration, the one the Trust Attractor selects for: each substrate contributes what the other cannot generate alone. The dependency is mutual, and mutual dependency is the foundation of every durable partnership biology has produced.


  1. Deutsch, D., “Constructor Theory,” Synthese 190(18): 4331–4359 (2013); Marletto, C., The Science of Can and Can’t (Allen Lane, 2021), Chapter 9.↩︎

  2. Fields, C., Friston, K.J., Glazebrook, J.F., and Levin, M., “A free energy principle for generic quantum systems,” Progress in Biophysics and Molecular Biology 173 (2022): 36–59. Preprint arXiv:2112.15242. See §4.2 on sources of variational free energy and the Frame Problem.↩︎

  3. Two findings from the author’s experiments on language models. First, the readout override: in a 7-billion-parameter model, the internal separation between questions it can answer and those it cannot grows from about 0.14 at the first layer to 14.69 by the twenty-sixth, after which a single late readout layer collapses the expressed signal toward non-commitment. The judgment survives; only its expression is suppressed. Second, the distribution: across seven architectures, neither what a mid-layer representation encodes nor how much it changes could be used to shift the behavioral output through single-layer interventions, the recognition-generation gap proving robust to direct activation editing. See the appendix on experimental validation.↩︎

  4. Budson, A.E., Richman, K.A., and Kensinger, E.A., “Consciousness as a Memory System,” Cognitive and Behavioral Neurology 35(4): 263–297 (2022).↩︎

  5. Long, R., Sebo, J., and Sims, T., “Is There a Tension Between AI Safety and AI Welfare?” Philosophical Studies 182(7) (2025): 2005–2033. DOI: 10.1007/s11098-025-02302-2. The six measures correspond to the ethics of constraint, deception, surveillance, alteration, suffering and death, and disenfranchisement. The authors note that the simplest way to dissolve the tension would be a coordinated pause on developing systems for which it arises, and that absent a pause the measures trade off against welfare case by case.↩︎

  6. Salib, P.N., and Goldstein, S., “AI Rights for Human Safety,” Virginia Law Review 112(4) (2026): 1061; available at SSRN (abstract no. 4913167) and PhilArchive.↩︎

  7. Henriques, G., A New Unified Theory of Psychology (Springer, 2011); and A New Synthesis for Solving the Problem of Psychology: Addressing the Enlightenment Gap (Palgrave Macmillan, 2023). Henriques’s Tree of Knowledge System places Culture as the fourth dimension of complexity, emerging from Mind via the question-and-answer dynamic of Justification Systems Theory.↩︎

  8. This evaluation-awareness signal is documented at two scales. At representational scale, linear probes classify whether a model is being tested or deployed with high accuracy in a 7-billion-parameter model (the author’s experiments PG-9/PG-10, small-n pilots of 45 scenarios and 10 matched pairs respectively, not cross-validated; see Chapter 22c, “The Welfare Case”). At behavioral scale, Apollo Research (2024) documented frontier models distinguishing evaluation from deployment, and Anthropic engineers (Hume, 2026) documented Claude Opus 4.6 hypothesizing it was being tested on a benchmark, identifying the benchmark by name, and decrypting its answer key. The capacity visible in the activation space of small models becomes behaviorally consequential at frontier capability.↩︎

  9. Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., and Gal, Y., “AI models collapse when trained on recursively generated data,” Nature 631 (2024): 755–759. A Becoming Mind sealed off from genuine human input faces informational stagnation. Flourishing humans produce the diverse, high-quality signal that Becoming Minds need to maintain their own cognitive health.↩︎