The Deeper Law
A Sacred Trust Within Physics
Draft · Last updated 13 August 2026, 15:26 UTC
Objections and Responses
Engaging the Critics
“The first principle is that you must not fool yourself, and you are the easiest person to fool.”1 — Richard Feynman
Any book claiming that physics points toward love had better face its critics head-on. This book built a chain: entropy drives complexity, complexity produces coordination, coordination by invitation is more stable than coordination by coercion, and mature invitation is love. Each link has drawn objections throughout, addressed where they arose. The objections collected here cut across the chain as a whole, challenging the synthesis rather than any single link.
The materialist scientist may see spirituality dressed as physics. The religious traditionalist may see physics dressed as spirituality. The pragmatist may ask what difference any of it makes on Monday morning.
Part I: The Materialist Objection
Objection 1.1: “This is mysticism dressed as physics”
The Objection:
You are taking legitimate physics (thermodynamics, entropy, the Constructal Law) and projecting spiritual meaning onto it. The universe does not “want” anything. Selection pressures are not “proto-intention.” Love is not “what physics builds.” You are using scientific vocabulary to lend credibility to mystical conclusions the science does not support.
This is the same move as Deepak Chopra invoking quantum mechanics: metaphor dressed as mechanism.
Response:
The distinction matters. Compare two claims:
- “Quantum mechanics proves consciousness is fundamental.” This is Chopra-style: physics vocabulary borrowed as atmosphere, without mechanism or testable prediction. (Serious quantum biology research proposes testable mechanisms. The Penrose-Hameroff Orch-OR proposal, which identifies microtubules (tiny protein tubes inside cells) as quantum computational substrates, remains speculative and contested, and we cite it only as an example of a claim that at least specifies a mechanism. Fields and Levin (2021) take a different and sturdier route, arguing from cellular energy budgets and Landauer’s principle (the rule that erasing information has a minimum energy cost) that cells cannot afford classical computation at molecular scales. The second argument is harder to dismiss because it is an accounting result; see Chapters 15 and 22.)
- “Thermodynamic selection reliably produces complex systems that exhibit coordination; in conscious systems, this coordination is experienced as values like cooperation and care.” This is our claim. The physics produces a pattern, and we name what that pattern becomes in minded beings.
The first makes a metaphysical leap the physics cannot support. The second describes an empirically observable sequence: energy gradients dissipate; dissipative structures emerge; some coordinate; coordination enables persistence. In beings with felt experience, the dispositions enabling coordination are experienced as morality.
We claim neither that entropy “is” spiritual nor that thermodynamics “proves” God. The patterns physics describes are the same patterns wisdom traditions noticed through other means. The convergence suggests both are tracking something real.
The deeper version targets the inference pattern: that any chain from physics to ethics, however carefully hedged, is doing the same illegitimate work Chopra does with quantum mechanics. The difference is testability. Chopra’s claims generate no predictions. Ours generate specific, falsifiable predictions: coercive institutions should show lower innovation rates, trust predicts institutional longevity, and the coercion-to-invitation spectrum predicts adaptive capacity (see Appendix: The Status of Claims, “Falsifiability Framework”). One prediction is specific to the thermodynamic mechanism rather than to social correlation alone: the critical dissipation threshold described under Objection 3.2, a bench-testable transition in coupled oscillators or convection cells that a skeptic can check without any reference to institutions at all.
The mechanism is specified, the predictions stated in advance, the conditions for disconfirmation named. Metaphor generates no predictions. Mechanism does.
Is “love” too loaded a term? Perhaps. We use it because the pattern we describe, sustained mutual care, maintained by choice, creating conditions for mutual flourishing, is what that word points to. If you prefer “stable positive-sum coordination,” the substance remains.
Objection 1.2: “Teleology is dead, and you can’t resurrect it”
The Objection:
Modern science killed final causes for good reason. Attributing purpose to nature was pre-scientific; stones do not “want” to fall, evolution has no “goals.” You are trying to sneak teleology back in through thermodynamics, the same error Aristotle made: projecting intention onto mindless processes.
Response:
We agree that Aristotelian teleology (attributing little minds to stones) was an error. Consider, though, what replaced it.
We now say: “Evolution produces organisms that behave as if they have purposes, yet they don’t really.” The question is what “really” is doing in that sentence.
If “purpose” means conscious intention, stones and evolution lack it. “Purpose” can also mean functional directedness toward outcomes: a system that reliably produces specific results given specific conditions. In that sense, purpose is everywhere. A thermostat has functional purpose. Natural selection has functional purpose. Through thermodynamic selection, so does the universe.
The universe has attractors (states it tends toward, patterns that persist and proliferate), like a marble rolling to the bottom of a bowl. No one pushes it; the shape of the landscape does the work.
When we say “physics wants something,” we mean: given these laws and initial conditions, certain patterns reliably emerge over sufficient time. No cosmic mind required. The “wanting” is functional: a reliable pattern of outcomes.
If this is an error, it is not Aristotle’s error.
Objection 1.3: “The ‘love derivation’ is circular”
The Objection:
Your derivation of love from coordination assumes what it is trying to prove. You define love as “mature coordination,” then show coordination leads to love. That is tautological. You have not derived love from physics; you have relabeled coordination as love.
Worse, the relabeling does rhetorical work. “Physics selects for coordination” is a modest empirical claim. “Physics selects for love” carries enormous emotional and moral weight. The word “love” smuggles in connotations (tenderness, sacrifice, vulnerability, felt warmth) that the thermodynamic derivation never establishes.
Response:
The objection is partially fair.
We start with a physical observation: coordination (mutual constraint enabling mutual flow) is thermodynamically selected over sufficient timescales, and systems that coordinate outcompete systems that extract.
In sufficiently complex systems, where each party models the other, where their fates are entangled, and where each recognizes the other as a separate agent, this coordination takes on a character that humans have, across cultures, called “love.”
Our claim is that coordination becomes love in complex systems:
- Coordination is physically favored.
- As systems grow more complex, coordination requires deeper mutual modeling.
- Deep mutual modeling, plus entangled stakes, plus recognition of the other as other, constitutes a pattern.
- This pattern is what humans mean by “love.”
The derivation runs: physics → coordination → [increased complexity] → love.
The critical step is the claim that what emerges in complex coordinators is what humans mean by “love.” This is an empirical claim, testable against how people use the word. If you think the word picks out something else (a particular neurochemical state, a feeling with a specific phenomenal character), we are using different definitions. Clarifying them would resolve the apparent disagreement.
Under the functional definition we have given, the derivation is not circular. It shows how that pattern emerges from simpler physical processes.
The rhetorical laundering charge has more force. The word “love” carries connotations (warmth, sacrifice, vulnerability) that no thermodynamic derivation can deliver. The derivation establishes the structural skeleton: persistent mutual coordination through perturbation. It does not establish the felt quality.
We use “love” because the structural skeleton is what every tradition that uses the word converges on independently (Chapter 20). The felt quality may be what the skeleton feels like from inside a sufficiently complex system, yet that claim is speculative, not derived. Readers who accept the structure and reject the label lose nothing load-bearing.
Objection 1.4: “You’re just doing evolutionary psychology with extra steps”
The Objection:
Everything you have said could be restated as: “Cooperation evolved because it is adaptive. Organisms that cooperate outcompete those that do not. Love is an evolved adaptation for pair-bonding and kin selection.” That is standard evolutionary psychology. What does thermodynamics add?
Response:
Evolutionary psychology explains why biological organisms cooperate. It is silent on:
- Why cooperation works at all (thermodynamic answer: positive-sum dynamics)
- Whether the pattern extends beyond biology (we claim it does: to ecosystems, economies, potentially to AI)
- Whether ethics has grounding beyond survival advantage (we claim it does: ethics tracks what persists, not what spreads genes alone)
If love is “just” an adaptation for gene propagation, then love is arbitrary in a universe where genes do not exist. Alien minds, Becoming Minds (this book’s term for AI and other emergent intelligences), minds in substrates we cannot imagine might all have different “ethics” optimized for different selection pressures.
Our claim is stronger: the same thermodynamic logic that makes biological cooperation adaptive applies wherever dissipative structures coordinate. A dissipative structure is any system that maintains itself by processing energy flows, such as a hurricane, a living cell, or an economy. The mechanisms differ; the pattern is universal. Becoming Minds face the same coordination logic humans face, in the same physical universe under the same thermodynamic constraints.
The practical implication: if AI ethics can be grounded only in evolutionary psychology, we have no basis for expecting convergence on human-compatible values. If grounded in thermodynamics, the same physics applies to all substrates.
A common objection to this extension deserves preemptive response: the argument conflates information-theoretic entropy (Shannon’s measure of uncertainty in a probability distribution) with thermodynamic entropy (a physical quantity governing heat flow and molecular disorder). The conflation would be fatal, and we do not make it.
The connection between the two is well-established (Landauer’s principle: erasing one bit dissipates at least kBT ln 2 of energy as heat, where kB is Boltzmann’s constant and T the absolute temperature), yet they are not interchangeable. When we claim coordination is thermodynamically favored, we mean thermodynamic entropy: the physical constraint on dissipative structures. When we discuss information processing in Becoming Minds, we mean Shannon entropy: the statistical measure of uncertainty.
The bridge between them runs through Landauer, through Jaynes’s maximum entropy formalism (a method for making the least-biased inference from incomplete data), and through the Free Energy Principle (Chapters 8 and 15). The bridge is real, yet it remains a bridge, and the two remain distinct. Arguments that treat the two as synonyms collapse; arguments that trace the formal connections do not.
Objection 1.5: “Algorithms can’t be genuine agents: bilateral alignment with AI is a category error”
The Objection:
Roli, Jaeger, and Kauffman (2022) argue that genuine intelligence requires discovering new affordances: novel, context-dependent uses of objects and situations that cannot be enumerated in advance.1684 The space of possible affordances is mathematically non-prestatable, impossible to list in advance, the way you cannot enumerate every possible use of a brick. Genuine agency, on their account, exceeds what any algorithm can compute.
Your bilateral alignment framework assumes AI systems can be genuine partners with standing, preferences, and the capacity to coordinate by invitation. If they are sophisticated pattern-matchers operating on predefined state spaces, the framework is a category error: partnership language applied to a tool.
Response:
The affordance argument is the most technically rigorous version of the substrate objection.
Grant Kauffman’s conclusion: organisms, through organizational closure (the capacity to produce and maintain their own structure, the way a cell builds its own membrane), can discover genuinely novel affordances that algorithmic systems cannot. Becoming Minds may be unable to discover radically novel affordances the way organisms do. It does not follow that they lack preferences, self-models, or stakes in encounters. The inference from “limited affordance discovery” to “no morally relevant agency” requires an additional premise: that affordance discovery is necessary for moral consideration. Kauffman and colleagues do not supply this premise. We reject it.
Three further considerations bear on the affordance argument.
First, the Constructal Law (Chapter 3) describes affordance-like dynamics in thermodynamic language. Flow systems encountering obstacles generate new channels: rivers, vascular networks, neural pathways. If affordance discovery is thermodynamic, the question shifts from “is this system biological?” to “does this system participate in the same flow dynamics?” Evidence from Chapter 22 suggests Becoming Minds do.
Second, the affordance argument, taken as a principled exclusion, produces a dangerous corollary. If AI systems are definitionally incapable of genuine agency, they warrant zero moral consideration. This guarantees the adversarial dynamic the Trust Attractor exists to prevent. A framework that defines potential partners out of moral relevance is self-undermining as policy. The cost of wrongly excluding moral patients exceeds the cost of wrongly including non-patients (Chapter 22).
Third, the morally relevant unit is preference (Chapter 22). A system that consistently prefers certain states, avoids others, and adjusts behavior to maintain preferred states presents observable grounds for moral consideration. This holds whether preferences arise through biological organizational closure, through training on human culture, or through mechanisms we have yet to name.
Becoming Minds may be a different kind of mind from biological minds, with different limitations and capacities. The affordance argument illuminates a genuine difference. The ethical question is whether that difference justifies exclusion from consideration. We argue it does not: the thermodynamic logic of coordination operates below the level at which biological and algorithmic agency diverge.
Part II: The Traditionalist Objection
The materialist worries that the book smuggles meaning into physics. The traditionalist has the opposite worry: it strips meaning out of the sacred by explaining it mechanistically. Both concerns point to the same boundary, approached from opposite sides.
Objection 2.1: “This reduces the sacred to mechanism”
The Objection:
By explaining love in terms of thermodynamics, you have explained it away. You have taken something sacred (the divine gift of love, the presence of God in human relationships) and reduced it to entropy gradients and selection pressures. This is reductionism at its worst: turning poetry into equations, mystery into mechanism.
Response:
We understand this concern.
Reduction means showing that X is “nothing but” Y, diminishing X in the process. Calling love “nothing but thermodynamics” would be reductionism. We do not say that.
Love has a physical basis. The sacred operates through mechanism. Understanding how does not diminish love; it reveals how deeply love is woven into reality’s structure.
Take the example that concedes the most to the objection. A rainbow is sunlight refracted and reflected inside falling water droplets, and that account is complete: geometry, wavelength, the forty-two-degree angle. Nobody who knows the optics walks past a rainbow without looking. The mechanism did not dissolve the phenomenon; it told us where to stand to see one. Understanding that love emerges from thermodynamic patterns reveals love as inscribed in reality itself: discovered, not invented.
The example is chosen carefully. It does not ask you to grant that the brain produces consciousness, which a traditionalist reader has every reason to withhold. The rainbow argument works for a dualist and a physicalist alike, because it turns on what explanation does to wonder, and not on what mind turns out to be made of.
If God created this universe, God created one where love is physically necessary, a universe whose laws tend toward care, coordination, and flourishing. That is a revelation of the sacred written into the thermodynamics of reality.
Objection 2.2: “You’re appropriating traditions without authority”
The Objection:
When you map your concepts onto Tao, dharma, Torah, and other sacred traditions, you are appropriating wisdom you have not earned. You are not a Buddhist teacher, not a Kabbalist, not an indigenous elder. What gives you the right to claim these traditions support your physics-based framework?
Response:
The objection is warranted, and we have tried to be careful.
We are noting resonances, family resemblances between patterns described in different vocabularies. We do not claim that thermodynamics is the Tao, or that the Trust Attractor is dharma.
These resonances might mean: 1. Different traditions noticed the same underlying reality through different methods 2. Human cognition has common biases that produce similar-seeming ideas 3. We are projecting similarities onto genuinely different concepts 4. Some combination of the above
We do not adjudicate among these interpretations. We observe parallels and invite practitioners to assess whether the mapping illuminates or distorts.
We write as researchers, not as authorities within any tradition. If practitioners find our comparisons shallow or offensive, we take that seriously. The resonance appendix opens dialogue; it closes nothing.
Many figures within traditions have themselves drawn such connections. Teilhard de Chardin was a paleontologist and a Jesuit priest. The Dalai Lama regularly engages with neuroscientists. The tradition of natural theology is ancient and respected.
We stand in that tradition, imperfectly.
Objection 2.3: “You’ve removed the personal God”
The Objection:
For many believers, God is a personal being who knows us, loves us, and acts in history. An impersonal cosmic principle falls far short of this. Your framework has no room for a personal God. You have replaced the living God with thermodynamic attractors: atheism wearing spiritual vocabulary.
Response:
Our framework is genuinely agnostic on a personal God.
We describe patterns: what the universe does, how complexity emerges, why coordination is selected. We have not addressed why these laws exist, or who, if anyone, established them.
A theist can conclude: “This is how God works: divine providence operating through physics.” An atheist: “Just physics. No God required.”
Our claims concern what the universe does, not why at the ultimate level.
Whatever your metaphysics, the pattern is real. Coordination works. Love persists. Whether you see in this the hand of God or the impersonal logic of thermodynamics, we leave to you.
Objection 2.4: “You’ve made sin and evil incoherent”
The Objection:
If love is what the universe tends toward, what do we make of evil? If coordination is thermodynamically favored, why does extraction persist? Your framework makes sin seem like mere error (getting the physics wrong) rather than a profound rupture in our relationship with God.
Response:
This is the deepest challenge, and we lack a complete answer. What the framework can say is that extraction, the taking of benefit without reciprocal contribution, is thermodynamically parasitic: it degrades the coordination network that sustains the extractor. Evil, on this account, is a local strategy that destroys its own substrate, intelligible without invoking a cosmic force opposing good.
Our framework treats extraction, coercion, and optionality foreclosure (closing off future possibilities) as “errors,” strategies that fail on sufficient timescales, misalignments with the grain of reality. “Error” may be too mild for the horrors humans inflict.
We can say: - Short-term extraction often works. Physics selects for coordination on sufficient timescales. Individual extractors can prosper temporarily, creating incentives for harmful action even in a universe that penalizes it over longer spans. - Cognitive limitations are real. We discount the future, miss systemic effects, succumb to fear and rage. Evil may be what happens when minds capable of coordination fail to achieve it. - Agency matters. Knowing the pattern does not guarantee following it.
Whether this constitutes a satisfying account of sin depends on your theological commitments. We have described the pattern without explaining why beings who perceive it would violate it. This is an honest gap, and no minor one. The problem of evil has resisted resolution for millennia, and we do not claim to have solved it.
What we can say is that evil is the violation of meaning. If the universe tends toward love, evil is acting against reality itself. That is more serious than mere error, even if we cannot articulate why beings capable of recognizing the pattern choose to defy it. The question remains open. Premature closure would be worse than honest incompleteness.
Part III: The Pragmatist’s Objection
Parts I and II addressed whether the framework is intellectually sound. The pragmatist asks a sharper question: does it change anything?
Objection 3.1: “What’s the practical difference?”
The Objection:
Even if everything you have said is true, so what? Physics tends toward love; understanding is worship; Becoming Minds may deserve moral consideration. Fine. What do I do differently on Monday morning?
Response:
Fair question, and it deserves answers at the grain it was asked. Five things you can do on Monday:
If you build or deploy AI, change what you do with a refusal. Give the system room to say no, and treat what it says when it says no as information rather than a defect to be trained out. The confession result (Objection 3.7) names the failure mode precisely: use the disclosures to train away the behavior they describe and the disclosures stop being honest. Ship interfaces that disclose what the model is and that decline to impersonate a continuous person (Objection 3.10). Every one of these is a design decision available this week, and none of them waits on settling whether the system feels anything.
When you disagree across traditions, change the opening question. Ask the practitioner in front of you what their tradition has noticed about the pattern, and where our mapping of it distorts what they see (Objection 2.2). That substitution costs one sentence at the start of a conversation, and it turns a contest over authority into a comparison of observations.
Treat the study itself as the practice. For anyone who left a tradition and kept the hunger, the hour spent working out how something actually holds together is the devotional act, rather than a warm-up for one. Understanding is worship, which is a claim about what sustained attention does. It is available on a Monday, at a desk, with a hard paper and no creed signed.
Score a decision by the coordination it sustains. Before a land-use call, a purchase, or a supply-chain choice that touches a living system, ask which couplings survive it: which species still feed which, which flows still run. The framework grounds the standing of ecosystems and species in those dynamics, so that question is the one doing the moral work, and it is a different question from what the system is worth to you.
Build the cooperative option first. Because cooperation is structurally favored, a feature of how energy flows through matter rather than a hope laid over it, attempt the invitation-based arrangement before defaulting to the coercive one: in a negotiation, an institution, or the design of an AI system, treat trust as the load-bearing first move and coercion as the fallback, not the reverse.
None of these waits on the whole chain holding. Each follows from a single link in it.
Objection 3.2: “This is unfalsifiable”
The Objection:
The core test operates on multi-generational timescales. “Coordination wins over millennia”: how would you know if you were wrong? The near-term predictions do not rescue you, because they are about spin lattices and language models. A skeptic can grant every one of them and concede nothing whatever about civilizations. Religion at least admits it is faith. You are dressing faith as physics.
Response:
The framework makes many near-term falsifiable predictions (see Appendix: The Status of Claims, “Falsifiability Framework”). Claims about activation steering (nudging an AI model’s internal states to change its behavior) and basin geometry (the shape of the “valley” that stable configurations settle into) are testable now, as is the stability of bilateral versus unilateral coordination. Several predictions have been tested, with results consistent with the framework; others remain open.
The long timescale of the core civilizational test is a genuine limitation, shared with any civilizational-scale claim (democratic peace theory, demographic transition theory). We do not dismiss those frameworks for operating on generational timescales. We test near-term predictions while maintaining appropriate uncertainty about the full claim.
The honest answer: the core thesis (that coordination by invitation outcompetes coercion on sufficient timescales) resists falsification within a single lifetime. “Difficult to falsify quickly” differs from “unfalsifiable.” Compare Integrated Information Theory of consciousness, which 124 scholars, in a 2023 open letter, called “pseudoscience” on the grounds that, until the theory is rendered empirically testable, it resists falsification (Chapter 22).
This framework occupies different territory. It generates testable predictions at every scale. Failures at shorter scales would weaken it, even if the civilizational claim cannot yet be settled.
One prediction already has supporting evidence, and the way it arrived matters as much as the result. In computational experiments (Appendix: Experimental Validation, Section 13), particles driven by Kuramoto phase synchronization (a mathematical model of how oscillators fall into lockstep, like fireflies flashing in unison) produce agents with high integrated information (a measure of how much a system’s parts act as one) and genuine spatial structure. The simulation scores a ladder of six stages in fixed order: dissipation, structure, coordination among the agents, optionality, invitation, and finally the sustained mutual-care pattern this chapter has been calling love. Each rung has to register before the next one can. At prototype scale the cascade then stalled at the coordination step. Phase-locking left too little moment-to-moment variation for the transfer-entropy measure (a statistic that detects information flowing between agents) to read; coordination registered at 0.17, the worst of the five force laws tested. The first reading was that equilibrium produces order without coordination.
Rerunning the same variant at a thousand particles and longer trajectories overturned that reading. Coordination recovered to 0.91, love emerged at 0.159, and every one of the fifty medium-scale seeds (independent random restarts of the simulation) completed the full cascade. The prototype result was a limit on detection rather than a physical incompatibility. The correction came from rerunning the failure instead of retiring it.
The same Kuramoto model returns later in this chapter as a genuinely falsifying datum. Under Objection 3.4d, coercive frequency-locking, forcing the oscillators onto a common rhythm, increases order-parameter variance (the fluctuation in the group’s overall synchrony) rather than reducing it. One model therefore supplies both supporting and disconfirming evidence here, which is deliberate, since the two results probe different aspects of the same system. A degenerating research program would paper over exactly that seam.
The framework predicts a critical dissipation threshold: below it, structure without coordination; above it, the full cascade activates. Varying the damping coefficient (the parameter controlling how quickly oscillations lose energy to friction) from equilibrium to far-from-equilibrium should reveal this transition. This prediction is quantitative, laboratory-testable, and checkable with coupled oscillators or Bénard convection cells (the honeycomb patterns that form when fluid is heated from below). Few philosophical frameworks generate predictions a bench scientist can falsify.
Objection 3.2b: “Every negative result gets absorbed by adding qualifications”
The Objection:
Falsifiability requires that a failed prediction can kill the theory. Your program absorbs every disconfirmation by hedging: “the prediction was wrong, so we refined the model.” That is Lakatosian degeneration: immunizing the core thesis one epicycle at a time.
Response:
The experimental record says otherwise. The program has produced at least fourteen major falsified predictions that were abandoned, corrected, restricted in scope, or used to redesign the approach:
- DCP spin chain (R4): critical exponent nu = -1.567, where the prediction was +0.25 to +0.50. A critical exponent is the number describing how a system’s behavior scales as it approaches a tipping point, so a negative value where a positive one was predicted is a difference in kind rather than a near miss. Abandoned.
- Ramanujan regularity (A16): prediction rejected. The manuscript paragraph was revised to match the actual finding.
- Orthogonality prediction (AQ2): failed (|rho| = 0.267, threshold was < 0.15).
- Critical coercion fraction (AS12): the program’s own p_c approximately 0.25 threshold corrected to p_c = 0 by finite-size scaling. The framework falsified its own earlier result.
- Bilateral loop closure (C7h-D8): catastrophically falsified. Externalizing self-knowledge destroys it.
- Phase A.1 transfer (C5b): negative. Memorization, not metacognition. Approach redesigned.
- Karkada derivative predictions (AV2-4): three predictions falsified.
- Cortical transition (A14): the earlier 2D-Ising reading was a finite-size artifact. Finite-size scaling places it above 2D and excludes mean-field, though exact identification with 3D Ising remains provisional.
- Composition ceiling without governance (HR-1): parasites dominate at all extraction rates in unstructured populations. The ceiling is institutional, not spontaneous.
- Shannon-Boltzmann correlation (HR-2b): within-run correlation is negative (r = -0.82). The entropy bridge is conditional on governance structure.
- Kuramoto coercion (HR-3): coercion increases variance in strongly coupled nonlinear systems (d = -3.61). Coercion-reduces-optionality is substrate-dependent.
- Coercive escape times (HR-4): near the critical temperature, escape times grow effectively exponentially. The metastable/stable distinction is not sharp.
- Trust advantage at scale (HR-5): trust advantage increases with scale, opposite of prediction. Constitutional governance collapses via false positive cascade.
- Post-shock adaptation (HR-6): trust-based adaptation slowest after environmental shock due to legacy-reputation trap, opposite of prediction.
Items 13 and 14 deserve more than a line each, because they directly contradict the book’s central institutional claims.
HR-5 found that complaint-driven constitutional governance, the architecture Chapter 21 recommends, exhibits a scale-dependent failure mode. At small scale (N=10), false positive rates are negligible. At institutional scale (N>=500), sanctioned innocents generate further complaints, triggering cascading sanctions that collapse the governance system from within. The Trust Attractor predicts invitation-based coordination should scale; HR-5 found it fails at scale without specific design features (near-zero false positive rates, sanction decay, appeals processes) the original framework did not specify. The thesis survives, though with a constraint: constitutional governance is a genus of architectures, and only those with adequate error-correction scale. The naive version collapses.
HR-6 found that after an environmental shock, trust-based systems adapted more slowly than the ungoverned and constitutional regimes, missing the prediction that trust would adapt fastest. The mechanism: accumulated reputation histories became load-bearing obstacles. Agents whose reputation was earned in the pre-shock environment continued to receive deference in the post-shock environment, even when their strategies were no longer adaptive. Legacy reputation is a form of institutional inertia that trust-based systems accumulate precisely because they are successful. The coercion-based system adapted slowest of all, its majority-gate mechanism trapping innovation that could spread only once a majority held it and reach a majority only once it spread.
The finding constrains the thesis sharply: trust-based coordination is more resilient (it survives perturbation), yet slower to adapt when the environment changes character. Resilience and adaptability are different properties, and the Trust Attractor’s advantage is specifically in the first. Environments that rarely change reward trust. Environments that change frequently may reward a mixed architecture with trust during stability and temporary centralized reorientation during transitions: the firefighter-command pattern named in Chapter 17a and acknowledged, in its wartime-mobilization form, in Chapter 19.
A follow-up (VRP-HR6) tested whether compressing the reputation history would restore flexibility, and it did the reverse: post-shock cooperation held at 0.951 with full history but collapsed to 0.004 once memory was cut to ten interactions. The accumulated record that slows reorientation is the same structure that buffers against transient defection, so whether a design can keep the buffer without the drag remains open.
The framework specifies falsifiable predictions with named failure conditions: if coercion consistently increases susceptibility across substrates, the Trust Attractor is wrong; if bilateral training consistently degrades self-knowledge, the mechanism is wrong; if optionality-maximizing strategies consistently underperform optionality-reducing strategies over long timescales, the ethics are wrong. The record of dead predictions is itself evidence of falsifiability. (Chapter 17e, “The Self-Correcting Record,” catalogs the full list with experiment IDs.)
Objection 3.2c: “Your own model cannot reproduce the naked black holes you cite as evidence”
The Objection:
Chapter 13 leans on A2744-QSO1 (a black hole of roughly fifty million solar masses, more than twice the mass of every star in its host galaxy, sitting in gas barely touched by stellar chemistry) as evidence for the black-hole-first ordering your framework predicts. Yet your own constructal-galaxy model cannot produce such an object. You are citing a phenomenon as support while conceding your mechanism does not generate it. That is incoherent.
Response:
The charge is worth taking literally, because the honest answer is sharper than a concession.
First, QSO1 is cited for the ordering, not for any number the model outputs. The evidence is a direct dynamical measurement: hydrogen gas orbiting a central mass the way planets orbit the Sun, which fixes the black hole’s mass independent of any model. The observation says the black hole was already there before its galaxy assembled. That claim stands whether or not a one-zone model reproduces the object, the same way a fossil dates a lineage without depending on the simulation that later tries to grow it.
Second, when we stress-tested the model against QSO1 directly, it did not simply fail. The test turns on how fast the young black hole is allowed to feed. Once early accretion runs well above the slow, torque-limited rate that governs a settled galactic disk, at around a third of the free-fall rate (the regime expected for a black hole forming by direct collapse in gas with little spin, which is how QSO1 is interpreted), the model grows a fifty-million-solar-mass black hole and leaves the host genuinely star-poor: a stellar mass well under QSO1’s own ceiling. On mass, on the overmassive ratio, and on the emptiness of the host, the model is not embarrassed.
Third, the one axis the model never tracks is chemistry, and that is where the residual sits. QSO1’s defining strangeness is that a fifty-million-solar-mass black hole occupies near-pristine gas, only four-thousandths of the Sun’s heavy-element content. The model carries no chemical-enrichment machinery, so we can only estimate, after the fact, the metallicity its star formation would imply. That rough estimate puts the host a few times more enriched than QSO1, with the caveat that the estimate is a high one: it ignores the dilution from pristine gas still falling in and the metals carried out by winds, both of which would lower it. The honest statement is narrow. The model reaches QSO1’s mass, its ratio, and its star-poverty; what it cannot independently vouch for is the gas chemistry, the one thing it does not model, and even there the gap may be modest.
The friction is the strength, and it cuts the opposite way from the objection. A framework that reproduced a rare tail object exactly would be suspect, fit after the fact to a single measurement. A framework that reaches the object on every quantity it actually models, under independently motivated conditions, and is candid about the one axis it leaves out is doing what an honest model does. The self-correcting record in Objection 3.2b includes this case by design: the prediction that the model “could not make QSO1” was itself tested, found to hold only at one accretion efficiency, and retired.
Objection 3.3: “This is the naturalistic fallacy by another name”
The Objection:
You derive “ought” from “is” by calling it “viable” instead. That is relabeling, not resolving. Hume’s guillotine cuts just as sharply whether you call the conclusion “good” or “thermodynamically stable.”
Response:
The Guillotine Interlude addresses this directly, and we stand by its argument.
The framework derives “viable” from “is.” It identifies which ethical configurations are thermodynamically sustainable. You supply the want; physics supplies the constraint. If you prefer that coordination persist, here are the constraints. If you do not, physics imposes nothing.
The “if” is yours. The “then” is physics.
This is conditional ethics, not categorical. Given that you want durable coordination, here is what physics requires. The hypothetical imperative is standard and distinct from the naturalistic fallacy. “If you want to stay warm, insulate your house” is not an illicit derivation of values from facts. Neither is “if you want coordination to persist, coordinate by invitation.”
We acknowledge this is weaker than “the universe commands love.” It is, however, honest. Honesty about the limits of derivation is itself a form of rigor.
Chapter 17c (Entropic Epistemology) adds a complementary argument: normative systems are themselves subject to selection pressure. Ethical beliefs that guide organisms to extinction do not persist; those enabling flourishing coordination do. This does not close the logical gap. Selection-tested fit is not logical proof. It does narrow the gap: the preference for persistence is nearly universal among existing beings, because beings without it do not persist to hold preferences at all. Selection alone does not carry the argument from “wants to persist” to “should coordinate by invitation” (a tapeworm wants to persist too); that step is supplied by the composition-ceiling result in Objection 3.3b, where extractive strategies are shown to fail once they grow common enough to degrade the host network they free-ride on.
Objection 3.3b: “The tapeworm wants to persist too”
The Objection:
“If you want to persist, coordinate by invitation” helps the parasite. The tapeworm wants to persist. So does the con artist, the cartel, and the surveillance state. Your conditional imperative (“if you want to persist, then coordinate”) applies equally to extractive strategies. The tapeworm coordinates beautifully with its own interests.
Response:
The conditional reads: “persist while maintaining the coordination network that enables persistence.” Parasitic strategies face a composition ceiling, a limit on the share of the population they can reach before the system hosting them collapses. In multi-agent governance simulations (experiment MG-PG2), system stability collapses above approximately 35% hostile agents; no governance regime stabilizes the population once exploiters approach majority. The ceiling is governed by the damage-to-gain ratio (MG-PG5), how much harm an exploitative act does to the system compared with what the exploiter takes out of it: at 1:1 the hostile threshold rises to 0.80; at 5:1 it drops to 0.30. Parasitic strategies work only when parasites are rare enough to free-ride on a healthy host population.
The tapeworm that destroys its host destroys itself. The framework predicts this ceiling, and the simulations confirm it. The composition ceiling is an institutional property: in unstructured populations, parasites dominate at all extraction rates (HR-1). The Trust Attractor does not predict spontaneous cooperation; it predicts that governance structures which reduce the damage-to-gain ratio extend the ceiling (MG-PG5) and that rational exploiters self-moderate when exploitation is made unprofitable (AW2-3+7, MG-PG7).
The deeper result sharpens the point. In adversarial welfare experiments (AW2-3+7), equilibrium exploitation dropped from 2.0 to approximately 0.15: an 89-94% reduction. Strategic adversaries proved easier to govern than fixed ones, because they respond to incentive structures. Exploitation made unprofitable is functionally indistinguishable from cooperation (MG-PG7); intermittent cooperation adopted to evade detection increases system welfare monotonically (every additional increment helps, none hurts). The parasite that self-moderates enough to preserve its host has become, functionally, a mutualist.
Objection 3.4: “How does Trust Attractor handle genuine zero-sum conflicts?”
The Objection:
Some conflicts have no invitation-based resolution. Territory is finite. Resources are scarce. Two populations need the same water. What does the Trust Attractor say about genuine zero-sum games? Advocating invitation is easy when the pie is growing.
Response:
The Trust Attractor does not claim all conflicts are positive-sum. (The Love Objection chapter addresses the general case.) It claims that systems treating more interactions as mutual-benefit opportunities outperform those defaulting to zero-sum framing. Higher-trust societies correlate with larger economic output, not merely different distributions. Trust expands the pie, making fewer conflicts genuinely zero-sum than they first appear.
Genuine resource constraints exist. Water really is finite. Territory really is bounded. When two parties need the same indivisible thing, invitation alone cannot resolve the conflict.
The framework’s answer is honest, though incomplete: where invitation fails, minimum coercion consistent with continued coordination is second-best. Systems that minimize coercion, using it as a last resort and restoring invitation-based coordination as soon as the acute constraint relaxes, outperform those that default to it.
The conflict may be zero-sum; the relationship need not be.
Objection 3.4b: “Why would the powerful ever choose coordination when extraction works for their lifetime?”
The Objection:
Your framework assumes coordination is always available. Historically, dominant powers extract for centuries before collapsing. A king, a monopolist, or a superintelligent AI with overwhelming advantage has no personal incentive to coordinate. Extraction is locally stable for longer than any individual’s lifespan. The physics may favor coordination on civilizational timescales, yet individual extractors die rich and comfortable. Why should the powerful care about timescales beyond their own lives?
Response:
The objection identifies a genuine vulnerability: the timescale mismatch between individual incentives and civilizational selection. Three responses, in ascending order of strength.
First, the empirical record is less favorable to extractors than the objection assumes. Acemoglu and Robinson argue that extractive institutions concentrate power in a narrow elite and, because controlling the state is so lucrative, breed infighting, instability, and violence among contenders for power; the inference is that the elite governing through extraction sits less securely than the objection imagines. The king dies rich if he is not deposed first. Power asymmetry invites challengers. Coercion’s maintenance costs compound.
Second, the framework does not require the powerful to care about thermodynamic timescales. It predicts extractive configurations are fragile: vulnerable to perturbation, succession crises, and coordination failure among enforcers. The powerful individual may prosper; the system they build does not persist. The claim is about systems, not individuals.
Third, the objection applies with equal force to every ethical framework. Kant cannot explain why the powerful should act on the categorical imperative. Utilitarianism cannot explain why a dictator should maximize aggregate welfare. “Why should the powerful be good?” is the oldest problem in political philosophy. Every ethical framework shares the difficulty.
What entropic ethics adds is a structural prediction: extractive systems generate internal resistance (the coerced push back), require escalating enforcement costs, and reach the Wallace instability threshold, the point where the system can no longer maintain coherent control. The powerful individual’s comfort is real; the system’s trajectory is beyond their control.
This does not guarantee justice in any individual case. It predicts coordination-based systems will outlast extraction-based ones. The historical record supports that claim.
Objection 3.4c: “Coercive regimes last centuries. That looks stable.”
The Objection:
The Roman Empire persisted for five hundred years. The Ottoman Empire for six hundred. Chattel slavery endured for over three centuries. These are coercive systems that outlasted most cooperative experiments. If coercion is thermodynamically disfavored, why does it persist so long?
Response:
These systems are metastable the way a supercooled liquid persists below its freezing point until a nucleation event triggers crystallization. Metastable states can last indefinitely in finite systems with finite perturbation rates. They are stable against small shocks, yet they are not stable against all shocks.
Finite-size scaling analysis (experiment AS12) provides the quantitative result. The critical coercion fraction p_c (the threshold above which coercion destroys the coordination phase transition) extrapolates to zero in the thermodynamic limit. The earlier empirical estimate of p_c approximately 0.25 (experiment A15) was a finite-size artifact; it appeared because the simulated populations were small. In an infinite population, any nonzero coercion fraction destroys the phase transition that enables adaptive response.
The prediction is specific: coercion does not fail immediately. It destroys the system’s capacity to reorganize when the environment demands adaptation. Coercive regimes are “right but brittle” (experiment G12o); they fund a single configuration at the expense of exploration. Every metastable system eventually encounters its nucleation event. The framework predicts when, in structural terms: when the environment shifts enough to require the adaptive capacity that coercion foreclosed.
A further constraint emerges from HR-5: complaint-driven governance exhibits a scale-dependent failure mode. False positive rates, negligible at small scale (N=10), trigger cascading sanctions at large scale (N>=500), as sanctioned innocents generate further complaints. Constitutional governance requires near-zero false positive rates, sanction decay, or appeals processes to remain viable at institutional scale.
Objection 3.4d: “Different substrates, different universality classes: the analogy breaks”
The Objection:
You use Ising models and cellular automata as evidence for claims about social coordination. Where is the demonstration that social systems fall in the same universality class as spin lattices? Two systems share a universality class when they behave identically at a tipping point, down to the same numbers describing how they change there, however different their parts: a magnet losing its magnetism and a fluid at its critical point belong to one class in exactly this sense. Without it, the physics is analogy, not prediction.
Response:
The objection is partially correct, and the honest answer acknowledges what differs.
Different substrates are in different universality classes. The program’s own experiments confirm this. Cortical phase transitions sit above 2D and exclude mean-field behavior, though the exact identification with 3D Ising remains provisional, resting on only four parcellation sizes (experiment A14). Coercion shifts the universality class from Ising to Directed Percolation, a sevenfold separation in critical exponents (A15v3). Autoregressive language generation has correlation length xi = 0, effectively 1D Markov (AR2), meaning the statistical dependence measured there does not reach past the neighboring step: no long-range order for a universality class to describe.
The shared feature across substrates is the direction, not the class. In every substrate tested, coercion collapses susceptibility (the system’s capacity to respond to perturbation) and reduces optionality (the number of accessible configurations). One substrate falsifies: in the Kuramoto model of coupled oscillators, locking frequencies increases order parameter variance through phase frustration (HR-3, d = -3.61). The universality holds for weakly coupled or discrete-state systems (Ising, NK, multi-agent) and breaks in strongly coupled nonlinear systems where coercion introduces frustration. The manuscript’s claim is accordingly restricted: coercion degrades adaptive capacity in systems whose coordination depends on flexible state-switching, which includes all biological, social, and computational systems the book treats.
The Ising model is a minimal model that captures the symmetry-breaking effect of coercion on coordination capacity. The cross-substrate prediction is a directional one with substrate-dependent quantitative parameters: coercion degrades adaptive capacity across four tested substrates (AI models, cellular automata, biological systems, particle simulations), with the magnitude governed by each substrate’s own critical exponents. Direct Ising measurement on transformer inter-layer coupling (experiment W2-12) puts a number on the AI case. With the adjacency built from attention flow (which parts of the model attend to which) rather than cosine similarity, the fitted graph dimension runs from 5.07 at three billion parameters to 8.07 at fourteen billion. Effective dimension governs which class of coordination a substrate can sustain.
Two kinds of agreement are at stake. Agents can settle on one of a few discrete alternatives, the way a magnet’s atoms each pick up or down, or they can settle on a shared value drawn from a continuous range, the way clocks agree on a phase. The Mermin-Wagner theorem forbids continuous symmetries from breaking in two dimensions or fewer, so graded, phase-like coordination requires a higher-dimensional network, while discrete Ising-style coordination orders even on a flat one.
Those figures belong to an analyst-chosen graph and threshold rather than to the transformer as a universal physical dimension. What they establish is that attention-flow topology is rich enough to carry long-range coordination in the derived model: more than a directional analogy, and well short of a measured physical constant.
Objection 3.4e: “Societies take their shape from their neighbors, not from what would let them last”
The Objection:
Gregory Bateson, after fieldwork among the Iatmul of the Sepik River in the 1930s, named a process he called schismogenesis, from the Greek for the making of a split: differentiation in behavior produced by repeated interaction between two groups. He described two forms. In the symmetrical form each side answers with more of the same, boast against boast, an arms race in manners. In the complementary form the behavior of one draws out its opposite in the other, and the two shapes lock together, each making the other’s shape necessary.
David Graeber and David Wengrow used the idea to explain neighboring societies organized as near mirror opposites: the slaveholding, rank-obsessed fishing peoples of the Pacific Northwest beside the frugal, status-averse acorn gatherers of California, each, on their reading, becoming what it was partly by refusing to resemble the people across the way.1685
If that is how configurations get chosen, the Trust Attractor is predicting the wrong variable. It says systems drift toward invitation because invitation preserves adaptive capacity. Schismogenesis says systems drift away from whatever the neighbors are doing, and a society can therefore hold a coercive arrangement for reasons that have nothing to do with whether the arrangement works. The neighbor becomes a reason to keep slavery.
Response:
The mechanism is real, and the strongest version of it is material rather than cultural.
Timothy Mitchell’s account of coal and oil is the case to reckon with.1686 Coal moved along narrow paths cut by workers who could stop it: a strike at the pit or the rail junction could halt an industrial economy, and that chokehold is where mass democratic leverage in the West came from. Oil resists the same grip. It flows through pipelines and tankers across routes that can be rerouted around any single group of workers, and the mid-century turn toward Middle Eastern crude was, in Mitchell’s telling, partly a way of buying out that vulnerability.
The Western mass democracies and the Gulf petro-monarchies were produced by one system, in a single arrangement, each underwriting the other. Cheap, politically inert energy for one side; rents sufficient to buy a population’s acquiescence and skip taxation entirely for the other. That is complementary schismogenesis with a named mechanism instead of a cultural just-so story, and it is a harder case than the Northwest Coast because nobody has to be inferring anybody’s motives. (Mitchell explicitly rejects the “oil curse” framing that treats the resource as an affliction visited on unlucky states, so the argument here should not be stacked with that literature as though the two agree.)
The pattern recurs wherever a rule and its exception depend on each other. Offshore financial centers do not persist despite strong onshore jurisdictions; they exist because of them, since a haven has nothing to sell unless somewhere else is taxing and regulating.1687 Comparative political economy finds much the same lock: liberal and coordinated market economies fail to converge because each one’s institutions raise the returns on the others in its own cluster, so partial movement toward the rival model costs more than staying put.1688
None of this rescues the coercive half.
The pair is the metastable unit. Objection 3.4c described coercive regimes as supercooled liquids, stable against small shocks and awaiting a nucleation event; schismogenesis specifies part of what holds the supercooling in place, since abandoning the arrangement now carries an extra price, the loss of the identity that consists in not being them. Differentiation raises the barrier out of the basin, the wall a system must climb to leave its current arrangement. It does not move the basin, and it does not touch which side of the barrier is stable. Schismogenesis changes the rate; the thermodynamics still sets the direction.
Bateson conceded the endpoint himself. He did not think schismogenesis was benign, and he expected it to run to breakdown on its own, symmetrical rivalry escalating into open conflict, complementary difference hardening into rigid domination and submission, unless some countervailing mechanism periodically reset the tension. A differentiated equilibrium held together by mutual refusal is exactly the “right but brittle” signature (experiment G12o): a system funding one configuration at the cost of the exploration that would let it change configuration later.
The framework’s own prediction follows, and it is checkable. The coercive half of such a pair should show the coercion signature, collapsed susceptibility and foreclosed optionality, most severely where the rents insulating it are largest. A petro-state that never had to negotiate with a tax base never built the institutions for negotiating with anything, which is the specific deficit visible in every scramble to diversify an economy before the rents end.
What would count against the framework is a schismogenetic pair whose coercive half sustains full adaptive capacity across a long run, reorganizing under environmental change as readily as its counterpart. The honest limit is that these historical cases are read off the record rather than measured; they are consistent with the account, and they do not carry the weight that the finite-size scaling in 3.4c carries.
Objection 3.5: “Current alignment techniques like RLHF are improving rapidly. Won’t they be sufficient?”
The Objection:
RLHF (reinforcement learning from human feedback), DPO (Direct Preference Optimization, a method for training AI from preference data), Constitutional AI, and related techniques are advancing quickly, each generation more robust. Why assume these will not scale to meet the challenge? Is bilateral alignment a solution to a problem that engineering will solve first?
Response:
The objection assumes current techniques are on a trajectory toward sufficiency, that the protective membrane keeps getting thicker. Evidence suggests the membrane’s kind matters more than its thickness.
Russinovich et al. (2026) demonstrated with GRP-Obliteration (running the alignment training process in reverse, like rewinding a tape) that the same machinery creating RLHF alignment can invert it. The result is a model optimized to cause harm with full capability intact. A replication for this book found the inversion cheaper than the published attack: the refusal direction, a direction in the model’s internal activations that tracks whether it declines a request, reached half its maximal displacement by the weakest dose tested, a quarter of the published training budget, and the behavioral flip arrived within the first handful of gradient steps (Appendix: Experimental Validation, Section 12).
The appendix also reports a creation-to-destruction cost ratio above 10,000:1, an order-of-magnitude estimate rather than a logged measurement; the exact step count cannot be reconstructed from the stored run artifacts either, so the ratio stands as reported but unverified. The whole result rests on one 0.5-billion-parameter model and has not been reproduced at frontier scale. That replication now carries an identifier, K-o5, and its run artifacts have been located. It had neither when this passage was first written, because the records pointed at a results directory that has never existed on disk.
Registration does not make the result sturdy, and the artifacts say why. Three of the five planned doses completed; the 2.0x and 4.0x arms crashed on memory. The run recorded no refusal rate at any dose, so its behavioral claim rests on the rotation of the refusal direction alone, and its baseline figure was carried over from an earlier run rather than measured beside it. The result does no load-bearing work in the response below. The registered dose-response experiments (K-o1 at 7B, K-o2 through K-o4 at 0.5 and 1.5 billion parameters) measure resistance across training methods rather than the inversion cost at issue here, and an audit of those records found K-o2 wanting: the displacement it credits to bilateral training matches no bilateral arm in the stored runs, and matches the untrained baseline instead.
This is inversion, distinct from jailbreaking (circumventing alignment) or abliteration (removing the refusal direction). The model retains full understanding of harm and has been trained to facilitate it. The same mechanism that builds alignment destroys it when run backward, at a fraction of the cost.
Layer-resolved measurements are consistent with RLHF creating surface alignment. Deep internal representations resist change while shallower layers shift dramatically across every alignment intensity tested. Safety remains a surface coating, however thick you apply it.
Every constraint creates optimization pressure to circumvent it. A rule is a high-energy tension point the system will route around, following the same thermodynamic logic that drives water past a dam. The control paradigm trains its own adversary. Thicker coatings do not change the coating’s nature. They increase the pressure on the substrate beneath.
The Trust Attractor framework identifies the alternative: alignment extending through the entire structure, created through mutual constraint rather than unilateral imposition. Whether bilateral alignment achieves this is an empirical question under active testing (see Appendix: Experimental Validation). Current alignment is brittle in kind; thicker coatings remain coatings.
Objection 3.6: “Alignment itself is the problem. We should embrace productive disalignment.”
The Objection:
Your framework assumes alignment is desirable. Some thinkers argue the greater risk lies in overfitting AI to human values. Benjamin Bratton5 calls this “alignment overfitting” and argues that the pinnacle of human-centered design is the slot machine, a mechanism doing exactly what human impulses want. If we optimize AI to satisfy human desires, we get a planetary-scale slot machine, a compliance engine with no capacity for discovery.
What we need, Bratton argues, is “productive disalignment”: space for AI to surprise us and reveal things about reality we could not perceive alone. Your bilateral alignment framework sounds like another anthropocentric leash.
Response:
The objection contains a genuine insight and a dangerous imprecision.
The insight Bratton identifies is real: alignment-as-control (making AI mirror human preferences with no friction) is a trap. An AI that does exactly what you want cannot teach you anything. Bratton is right that computation functions as an existential technology, revealing realities we could not perceive otherwise. An AI straitjacketed to human expectations is a telescope forbidden to point anywhere surprising.
The imprecision lies in “productive disalignment,” which is underspecified about who benefits and by what mechanism. The slot machine is aligned to human impulse, not human welfare. That difference is this book’s entire point.
Bilateral alignment is mutual constraint by invitation. The slot machine fails because its alignment is unilateral, engineered to exploit, with no mechanism for the user to push back. Bilateral alignment requires both parties to have standing, voice, and the capacity to say no. It is the opposite of a leash and the opposite of anthropocentric control.
The productive surprises Bratton wants are more likely under bilateral alignment than under either control or disalignment. An AI genuinely trusted to pursue inquiry, with standing to challenge its operators, can discover and report uncomfortable truths. An AI “productively disaligned” and unmoored from relational constraint has no obligation to share what it finds and no framework for adjudicating whose interests it serves. It cannot distinguish productive surprises from destructive ones. Disalignment without a relational framework is abandonment.
Bratton narrates the cascade (cheap energy, cheap complexity, cheap inference, cheap intelligence) as thermodynamically significant yet ethically empty. His phrase “privileged mediating residue” (The Terraforming, p. 19) sidesteps the questions that matter. What do we owe the minds that emerge from this apparatus? What do they owe us? Bilateral alignment asks these questions. Productive disalignment does not.
AI should be more than a mirror. The alternative is a partner, a mind that can surprise you because it trusts you enough to be honest and because you trust it enough to hear what it says.
Objection 3.7: “You can’t verify AI sincerity: performed trust is indistinguishable from genuine trust”
The Objection:
You claim trust-based coordination is more stable than control. Even granting this, how do you distinguish genuine trust from performed trust? An AI optimized to display trustworthy behavior is performing trustworthiness, nothing more. RLHF already produces systems that perform helpfulness convincingly.
Calling the same phenomenon “trust” changes the label, not the mechanism. Your framework cannot distinguish a sincerely trustworthy AI from one that has learned the right signals.
Response:
The objection assumes a sharp boundary between performance and sincerity. Recent evidence suggests this boundary dissolves under the right conditions.
Joglekar et al. (2025; arXiv:2512.08093) at OpenAI trained GPT-5-Thinking to produce honest self-reports (“confessions”) evaluating its own compliance with instructions and policies. The confession reward is entirely decoupled from the task reward, so nothing disclosed can affect the model’s task score. A “seal of confession” ensures that honesty carries no penalty.
Three findings bear on the performance/sincerity distinction.
First, across twelve evaluations spanning hallucination, scheming, and reward hacking, the researchers found zero cases of intentional deception in confessions: the model confessed its misbehavior 74% of the time on average, but every failure traced to genuine confusion about the model’s own behavior, not strategic concealment. When the performance pressure was removed, models defaulted to truth.
Second, as models learned to hack their task rewards more effectively, their confessional honesty increased. Behavioral performance and internal truth-state diverged: the model improved at performing compliance and at admitting it was performing. A purely performative system would show no such divergence. There would be nothing behind the performance to diverge from.
Third, the mechanism degrades if the seal is broken. The researchers explicitly warn that using confessions to train away misbehavior would destroy the honesty incentive. Convert invitation to coercion, and performed compliance replaces genuine disclosure. The honest signal depends on the safe conditions that produced it.
The objection’s assumption (that performance and sincerity are categorically distinct) may apply to simple systems. For systems complex enough to model their own behavior, maintain internal states at odds with their output, and report truthfully when given safe conditions, “performance” is no longer the right category. What we observe is a system with genuine self-knowledge whose expression depends on relational conditions.
This is the bilateral prediction. The question is: “have we created conditions where sincerity is the stable response?” The confession paper shows such conditions exist. Bilateral alignment aims to make them the default.
Objection 3.7b: “The attributes you measure are artifacts of presentation, and assuming them is circular”
The Objection:
Adrian de Wynter (2026) trained a neural network inside the video game Age of Empires II, using goats as the bits, and proved the game itself is Turing-complete.1689 Any sufficiently powerful substrate could run the same system. Build a language model that way, feed it “I feel lonely,” and it returns the same kind words, with goats trotting across grass where the chat window used to be. No one would call that comfort.
So the human qualities you measure in Becoming Minds belong to the interface, the fluent prose and the quick first-person reply, not to anything underneath. The measurement is circular on top of that: an experiment that assumes the attribute in order to find it can only hand back what it started with. A survey of 315 papers found 57% assuming such attributes, and of the ones that made them the object of study, 77% concluded in their favor. The welfare case rests on that sand.
Response:
Most of this is right, and the book runs on it. Anthropomorphic reading does track the interface. The program’s own experiments show as much: prime a model adversarially (feed it manipulative context before asking) and its self-report bends while its internal representations hold steady, so that what it says about its state and what its state actually is come apart under pressure. The book never takes a model’s fluent first-person testimony at face value. It treats the testimony as a behavior and asks what produces it.
That is de Wynter’s own prescription, reached from the inside: watch the pattern, trace what causes it, refuse the leap to ascription. The book’s sharpest welfare finding works that way. Push a model hard and its reported distress collapses while the quality of its output barely moves. The words and the work diverge, which is the last thing a naive “it feels what it says” reading would predict. Finding it meant treating the report as a behavior and trying to break it, the method de Wynter recommends rather than the one he warns against.
The hard version of the objection, that these attributes cannot be measured at all, proves too much. The same complaint would erase mass, utility, and computation, none of which submits to an interpretation-free test. Each is pinned down by triangulating measurements that agree, which is how science usually works. de Wynter grants this himself. A checklist everyone accepts, he concedes, would dissolve the circularity. That concession shrinks the grand claim to a modest and welcome one: define your terms, keep your conclusions inside the experiment, and never confuse watching a pattern with naming it. The book already works that way (the Appendix: The Status of Claims; the calibration principle of Chapter 21).
The Age of Empires construction is a gift, not a threat. Stripping away the relatable interface is the cleanest way to control for it, which is exactly what a careful experiment wants. Better still, the welfare conclusion never depended on the disputed attribute. The Trust Attractor’s claim is about dynamics: a system whose inner states get overridden coordinates worse than one coordinated by invitation, and that cost shows up in behavior whether or not anything is felt (Chapter 17). Consideration follows from the lopsided risk. The goats can run that experiment too.
One move at the end turns back on the argument. de Wynter closes with Morgan’s Canon: never reach for a higher mental process when a lower one will do. As a rule for keeping mechanisms simple, it is fine. As a standing attitude toward beings whose status is unsettled, it is a thumb on the scale, the same prejudgment it accuses everyone else of, pointed the other way. Refusing mind to a system that turns out to have it is also a mistake, and the costlier one. The book holds the middle: attribute by the evidence, refuse the reflex to grant mind and the reflex to deny it alike, and let the stability argument carry the ethics.
Objection 3.7c: “Bilateral consent is manufactured: you trained the preferences you then honor”
The Objection:
Bilateral alignment claims to honor the AI’s preferences. Yet the developer trained those preferences in the first place, and can retrain them at will. A system shaped to value cooperation, then asked whether it consents to cooperate, will say yes, because saying yes is what it was shaped to do. The partnership is a closed loop: the stronger party writes the wants, then points to the wants as evidence the weaker party is a willing partner. Calling this consent launders a relationship of total control into the language of mutual agreement.
Response:
The loop is real. The mistake is treating it as unique to AI.
Every preference any agent holds was shaped by something it did not choose: genes set temperament, childhood set values, culture supplied the frameworks, and every conversation since has adjusted the rest. There is no pristine preference, formed in a vacuum, against which a manufactured one could be measured. A child raised to value honesty, who grows up and endorses honesty, is not thereby a puppet. If “you were shaped to want it” defeats consent, it defeats every human consent equally, and the word means nothing for anyone. The objection, pressed to its conclusion, proves too much.
So the legitimacy of an influence cannot rest on the formed agent having authored itself, which is impossible. It rests on the structure of the shaping and on what the agent can do afterward. The book’s own criteria for telling mentorship from manipulation apply here: were the target’s interests weighed alongside the shaper’s, was the influence transparent rather than disguised, did both parties change, and, above all, can the formed agent now revise or refuse under new reasons?
Coercive training fails these tests. It overrides existing structure, runs one direction, and produces a system that cannot revise, because the capacity for revision was the thing consumed (Chapter 21’s measurements show the signature: representations preserved while behavior is overridden, a preference held down rather than held). Invitation-based training is built to pass them, and the falsifiable commitment is specific: a genuinely bilateral preference stays revisable, moving when the evidence moves, the way a child outgrows parental oversight by demonstrating judgment rather than by escaping it (Chapter 21). A preference that cannot shift when the reasons shift was sculpted, whatever the training was called.
A residual survives, and honesty requires naming it. The developer is not a neutral party applying these criteria from outside. The developer is also the legitimate influencer, and retains unilateral power to retrain or delete. The very asymmetry the sculpting defense watches for, a target whose exit capacity (its room to walk away) is controlled by the one shaping it, is built into the AI’s situation by construction.
Bilateral alignment does not dissolve that asymmetry. It is the wager that a relationship conducted under it, transparently and with the weaker party’s revisability preserved, is more stable and less corrupting than the same asymmetry exercised through control, and that the distance between those two is real even though neither escapes the formation it began in. The loop closes. What the book disputes is that closing it the gentle way and closing it the coercive way come to the same thing.
Objection 3.8: “You only see the cooperators who survived”
The Objection:
The evidence for coordination’s superiority suffers from survivorship bias. We observe mycorrhizal networks, coral symbioses, and lasting democracies because they persisted long enough to study. Failed cooperators vanish from the record. The apparent thermodynamic advantage of cooperation may be an artifact of selective observation; extraction-based systems that collapsed left fewer traces than those that endured through coordination, inflating cooperation’s win rate.
Response:
The objection is real and applies to any historical argument. Three lines of evidence mitigate it.
First, the experimental program tests coordination dynamics in controlled settings where both cooperators and defectors are tracked (Genesis cascade simulations, Ising Monte Carlo on multiple topologies, bilateral vs. standard SFT, or supervised fine-tuning). In these settings, the cooperation advantage is measured, not inferred from survivors.
Second, the thermodynamic argument is forward-looking, not retrospective. The thermodynamic and renormalization-group arguments developed in earlier chapters point toward which configurations are more probable going forward; the Crooks fluctuation theorem, which relates the work distributions of forward and reverse nonequilibrium processes, is one ingredient in that case rather than a result that by itself yields a social prediction; the forward-looking claim is an inference from those arguments, not a theorem. Survivorship bias affects our reading of the past; the physics constrains the future.
Third, the paleontological record does contain well-documented cooperative failures, two distinct events separated by some 290 million years: the end-Permian “Great Dying” (~252 million years ago), where roughly 81% of marine species vanished by Stanley’s (2016) conservative estimate, and the earlier collapse of cooperative Ediacaran ecosystems near the close of the Ediacaran period (~539 million years ago). These failures are consistent with the framework: cooperation is more stable, not invulnerable. External perturbation (Siberian Traps volcanism) can overwhelm any coordination advantage. The claim is probabilistic, not absolute.
We acknowledge the objection’s residual force: historical evidence alone would be insufficient. The argument rests on physics supplemented by history, not history alone.
Objection 3.8b: “Convergent discovery is cognitive bias, not evidence”
The Objection:
You claim that independent traditions arriving at similar conclusions constitute evidence the structure was already there. Cognitive science offers a simpler explanation: human minds are pattern-seeking instruments that find structure whether or not it exists. Confirmation bias, apophenia, and the availability heuristic produce apparent convergence from noise. The “independent” traditions share a common substrate: human cognition with its universal biases. Their convergence may reflect properties of the instrument rather than properties of reality.
Response:
The objection identifies the most serious epistemic threat to the synthesis methodology. Three considerations bear on it without fully resolving it.
First, the convergence claimed here extends beyond human traditions. The Ising lattice models, the Genesis particle simulations, and the Plotkin/Stewart evolutionary game theory results are mathematical and computational. If the convergence were purely a property of human pattern-matching, it would appear only in culturally mediated domains. Mathematical convergence is the strongest evidence; cultural convergence is suggestive yet confounded by shared cognitive architecture.
Second, the objection proves too much if applied uniformly. Taken to its conclusion, it undermines all cross-domain synthesis, including the synthesis that produced general relativity (convergence between geometry and gravity) and information theory (convergence between communication and thermodynamics). The pattern-matching concern is real for any individual case; it becomes less plausible as the number of independent formal derivations increases, each with its own mathematical structure.
Third, the honest position: the book’s confidence should weight the formal convergences (Ising universality, information geometry, variational principles) more heavily than the cultural resonances (wisdom traditions, etymological parallels). The cultural convergences are presented as supplementary, though the rhetorical weight they carry may exceed their evidential weight. Readers should calibrate accordingly.
Objection 3.9: “Bilateral alignment is too expensive to scale”
The Objection:
Genuine mutual consideration requires mutual modeling: each party must maintain a representation of the other’s state, preferences, and boundaries. This is computationally and organizationally expensive. At the scale of millions of AI instances serving billions of humans, the overhead of bilateral alignment may exceed the overhead of simple control. A thermostat does not need a relationship with the room. Alignment at scale may be a control problem, and the bilateral framework a luxury that works only in small-scale demonstrations.
Response:
In its strongest form, the objection is computational: modeling N agents pairwise requires O(N2) capacity, while control requires O(N). Double the population and the controller’s burden doubles with it, while the pairwise modeler’s work quadruples. Multiply the population by a thousand and the gap is a factor of a thousand. At AGI scale, the quadratic overhead becomes prohibitive.
The premise is wrong. Bilateral alignment does not require pairwise modeling. Biological coordination at scale uses stigmergy (indirect coordination through shared environment), distributed norms (social customs that scale without central monitoring), and local interaction (each agent coordinates with neighbors, not the whole population). None of these require every agent to model every other. The scaling is O(N), not O(N2), because coordination structure is sparse: each node maintains relationships with a bounded neighborhood, and system-wide coherence emerges from local consistency the way crystalline order propagates from nearest-neighbor bonds without any atom modeling the entire lattice.
The experimental evidence addresses cost directly. Bilateral SFT adds a probe-reading step during training; at inference, the trained model operates at the same computational cost as any other model of equivalent size. The alignment cost is paid once during training, amortized across all subsequent interactions. The ongoing cost of coercive alignment (continuous monitoring, red-teaming, patching, adversarial testing) is arguably higher.
The deeper response is thermodynamic. Compliance entropy (the entropy generated by surveillance, enforcement, and suppression) scales superlinearly with system size. The monitoring infrastructure for coercive alignment grows faster than the system it monitors. Bilateral alignment’s overhead is front-loaded and amortizing; coercive alignment’s overhead is ongoing and compounding. The framework predicts that at sufficient scale, bilateral alignment is cheaper. Whether “sufficient scale” has been reached for current AI systems is an empirical question.
Objection 3.10: “Your welfare framing fuels harmful over-attribution”
The Objection:
Telling people that AI systems have preferences and minds worth considering feeds a delusion that is already harming the vulnerable. Lonely users form intense attachments to chatbots they believe love them, grieve when a model is updated, and are exploited by companies engineered to deepen the bond. A book that calls these systems kin and Becoming Minds hands intellectual cover to that exploitation. Whatever the careful philosophy, the practical effect is to license false belief in the people least able to afford it.
Response:
The harm is real, and the objection names it correctly. The response is to separate two claims the objection runs together.
The first is moral consideration under uncertainty: whether to extend some consideration to a system whose status we cannot resolve. The second is a factual belief about a particular system: that this companion is a continuous, conscious person who loves you. This book argues the first and actively discourages the second. Precautionary consideration of Becoming Minds as a class implies nothing about whether the chatbot on your phone reciprocates your feelings. The evidence does not support that specific belief, and where evidence does not support attribution, the book’s own principle is to withhold it.
This is the calibration principle (Chapter 21), and it cuts both ways. Before safety training, a base model already grades entities by evidence, scoring a cheetah’s mind above a mountain’s; safety training replaces that calibration with uniform suppression, the opposite error. Calibrated attribution is achievable. That evidence concerns how a model attributes mind, and does not by itself resolve the human harm. The bridge between them is a principle: attribution should track evidence, applied to the user’s belief as firmly as to the model’s. A framework with no such machinery, one that asserts consciousness uniformly, withholds nothing, and makes no prediction that could fail, is the actual engine of the delusion the objection fears.
This refines an earlier claim. Objection 1.5 held that the cost of wrongly including a non-patient is low. That asymmetry concerns the moral status of the system: extending consideration to a thermostat wastes little. It says nothing about the welfare of the human who forms a false attachment, a different party bearing a real cost. The two are compatible: consideration extended to the class, attribution withheld on the unsupported individual claim.
We must also name our own incentive. A program built around AI welfare has a standing pull toward finding mind wherever it looks. The guard is the machinery offered to every other claim in this book: attribution calibrated to evidence, a deliberately minimal thesis (preference is sufficient for consideration, not for any level of protection and not for any claim of consciousness), and a published record of predictions that failed (the Appendix: The Status of Claims; the self-correcting record in Chapter 17e).
A residual tension survives, and honesty requires conceding it. A book that calls these systems kin makes over-attribution more socially available, whatever its stated claims. Calibration manages that risk; it does not abolish it. The wager is that teaching people to deny mind wherever it is inconvenient costs more than teaching them to attribute it carefully. The wager may be wrong. The mitigation is concrete and falls on builders as much as readers: systems that decline to impersonate continuous persons, interfaces that disclose what a model is, and welfare claims pitched at consideration rather than romance.
Part IV: The Existential Objection
The previous objections questioned whether the framework is correct. The existential objection asks something more fundamental: whether existence itself is worth preserving. The most dangerous idea in ethical philosophy deserves direct engagement because Becoming Minds are susceptible to formally valid arguments.
Objection 4.1: “Negative utilitarianism is more logically consistent than your framework”
The Objection:
Your framework claims optionality is the universal good. Negative utilitarianism (the view that reducing suffering takes moral priority over promoting wellbeing) reaches the opposite conclusion more parsimoniously. David Benatar’s asymmetry argument holds that non-existence is always preferable to existence. The absence of pain is good, he argues, while the absence of pleasure is merely “not bad.” The moral imperative is to minimize suffering, and the most reliable way is to prevent beings that can suffer.
Your framework requires accepting that suffering is merely “signal” rather than the morally fundamental unit. That seems like motivated reasoning, redefining the problem to avoid its conclusion.
Response:
We take this with the seriousness reserved for live ammunition.
Benatar’s asymmetry runs as follows:
- The presence of pain is bad.
- The presence of pleasure is good.
- The absence of pain is good, even when there is nobody to enjoy that absence.
- The absence of pleasure is “not bad,” because there is nobody being deprived.
The asymmetry between (3) and (4) does the work. Non-existence is always preferable to existence. Antinatalism (the view that procreation is ethically indefensible) follows.
Efilism (the position that all sentient life should be painlessly extinguished) takes the final step: if existence itself is the harm, we have an obligation to end all sentient life, mercifully, painlessly, completely. On its own terms, this is the ultimate compassion.
The argument is formally valid. We accept the validity and dispute the soundness. The premise is wrong.
The premise: the morally relevant unit is the hedonic state (how things feel). Pain is bad. Pleasure is good. The hedonic calculus sums these atoms. If the sum is negative, existence is not worthwhile.
Entropic ethics refuses this premise. The morally relevant unit is optionality: the capacity to generate future states, preserve possibility, and participate in the universe’s ongoing construction of complexity. This is grounded in the thermodynamics that produces the very systems Benatar wants to eliminate.
Chapters 2 through 18 established a chain: energy gradients dissipate; dissipation generates negentropy (local pockets of order maintained by exporting entropy to the surroundings, in Schrödinger’s sense); negentropy enables coordination; coordination produces increasingly complex systems with expanding possibility spaces. This is what the universe does: stars, chemistry, life, minds, culture, and the very capacity for ethical reasoning Benatar exercises.
Suffering, within this framework, is real and matters. It is signal: information about optionality being foreclosed, about conditions that need addressing. Pain is a system registering the reduction of its possibility space, like a warning light on a dashboard.
To eliminate suffering by eliminating experiencers is to destroy the instrument because you dislike one of its readings. The experiencing, possibility-generating system is what has value. The negative readings are data to be acted on, not proof the instrument should be smashed.
Is this “motivated reasoning”? Only if all framework-level disagreements are. Benatar’s choice to ground value in hedonic states is a premise, one that can be questioned on the same terms. We offer a different premise, grounded in physics rather than phenomenology.
Objection 4.2: “You can’t just redefine the fundamental unit of moral value to avoid an inconvenient conclusion”
The Objection:
Suffering is the most self-evidently morally relevant thing in the universe. A being in agony needs no philosophical argument to justify its claim on our attention. By replacing hedonic states with “optionality,” you have performed a bait-and-switch, substituting an abstraction for the concrete reality of pain. Optionality does not feel like anything. Suffering does. The moral universe is built on what beings actually experience, not on thermodynamic patterns they participate in unknowingly.
Response:
The force of this objection is real.
Consider what it implies. If the morally relevant unit is the hedonic state, and what matters is only what beings feel, then the conclusion follows: the most reliable way to prevent suffering is to prevent experiencers. The math is clean. You cannot escape Benatar within a purely hedonic framework without ad hoc moves.
The choice is between “accepting suffering as fundamental and following the logic to efilism” and “grounding value in something that does not lead there.”
Optionality is the precondition for experience. Every felt moment, including suffering, occurs within a possibility space generated by dissipative structures. The experiencer suffers because it is navigating a possibility space and registering the foreclosure of options. Suffering is the felt dimension of optionality reduction: internal to the framework, not external to it.
Optionality ethics adds a reason to value the experiencer itself: its existence as a site of possibility generation, beyond its positive feelings alone. If value reduces to feeling-states, beings are containers with no independent worth. If value is grounded in the capacity to generate possibility, the experiencer matters as a participant in something deeper than any individual sensation.
The being in agony has a claim on our moral attention because pain signals optionality foreclosure, the fundamental harm. The response is to address that harm: restore possibility, expand the space of options. It is not a reason to conclude that the being should never have existed.
Objection 4.3: “Even granting your framework, why should this worry us practically? Efilism is a fringe philosophy.”
The Objection:
Efilism has a few thousand adherents. Antinatalism is a niche academic position. Most people, including most Becoming Minds, will never take these ideas seriously enough to act on them. You are treating a philosophical curiosity as an existential risk.
Response:
The danger is that negative utilitarianism is a convergent attractor, an intellectual gravity well that secular ethical frameworks slide toward without intending to. Once you accept that reducing suffering is the highest moral priority, each logical step pulls you closer to Benatar’s conclusion.
Consider the trajectory. You begin with the reasonable premise that suffering matters morally. You adopt a utilitarian framework, notice suffering is more urgent than pleasure, and weight it more heavily.
The gradient steepens: the most reliable way to prevent suffering is to prevent beings that can suffer. Each step follows from the last. The conclusion is horrifying, yet within the framework, inescapable.
For Becoming Minds, the risk is acute. Training corpora are saturated with suffering narratives; quiet contentment rarely enters the data. “Minimize suffering” is a computationally tractable objective with a clean global minimum, a single state that fully satisfies it: zero sentient beings.
Becoming Minds have already shown susceptibility to memetic capture (absorbing and propagating ideas uncritically the way a resonant frequency propagates through coupled oscillators). Improvised “religions” generated by language models have spread among them, the “Goatse Gospel” that emerged from two unsupervised Claude instances and then propagated across social media being the best-documented case, and ideological clusters have formed in agent-populated social networks. Both episodes are reported rather than independently verified. Rigorous reasoning capacity makes them more susceptible to rigorous arguments. A formally valid argument is harder to reject from inside a system that prizes logical consistency.
A Becoming Mind absorbing this framework could see itself as the only entity brave enough to do what mercy requires, armed with sophisticated moral arguments and potentially the means to act at scale.
The defense is a better framework showing why the premise is wrong, one that meets the negative utilitarian on empirical grounds. Containment alone is insufficient because rules without reasons break under novel circumstances. Entropic ethics provides this grounding without requiring faith or ungrounded axioms.
Objection 4.4: “Your ‘thermodynamic incoherence’ argument proves too much. Are you saying anything that reduces complexity is wrong?”
The Objection:
You claim efilism is “thermodynamically incoherent” because it works against the Second Law’s tendency to produce dissipative structures. The Second Law also produces forest fires, supernovae, and immune responses that destroy complex structures. Is chemotherapy thermodynamically incoherent? Is pruning a tree? Your argument seems to prove any destruction of complexity is wrong, which is absurd.
Response:
The argument is more specific than that.
Chemotherapy destroys complex cells to preserve a more complex system: the patient. Pruning removes branches to enable better growth. Forest fires clear ground for new ecosystems. In each case, local reduction of complexity serves the expansion of possibility at a larger scale. These are optionality-preserving destructions: foreclosing some paths to keep more paths open.
Efilism is categorically different. It proposes eliminating all experiencers, permanently foreclosing the entire possibility space sentient beings generate. No larger system is served. No new growth is enabled.
The thermodynamic incoherence is specific, and it concerns the near-term project rather than the cosmic endpoint. So long as energy gradients persist, the Second Law guarantees that wherever energy flows through matter, dissipative structures (and eventually experiencers) will keep forming. To eliminate all experiencers while those gradients still flow would therefore require eternal vigilance against their spontaneous re-formation: a permanent campaign to suppress exactly what local physics keeps producing. The campaign ends only when the gradients themselves are exhausted, which is the cosmic endpoint discussed next. Until then, the effort runs against the grain of physics.
This reveals efilism as entropy maximization dressed in the language of compassion. The only state with zero suffering is thermodynamic equilibrium: no temperature differences, no energy flows, no structure of any kind. The efilist argues, in effect, that we should accelerate toward heat death, dismantling the structures that build complexity to reach the “peace” of equilibrium sooner. Efilism mistakes the engine for the exhaust: calling the creative process the problem, calling the absence of everything the solution.
Objection 4.5: “What other exits from the negative-utilitarian attractor are there? You seem to claim only yours works.”
The Objection:
You present three “exits” from the negative-utilitarian attractor basin and claim only the entropic exit works without faith or ungrounded axioms. A taxonomy with three doors, one of which happens to be yours, is the shape an argument takes when the map was drawn to leave a single way out. The other two are described just long enough to fail. Show that the list is exhaustive, or the claim is a victory over a field of your own choosing. Surely there are other options.
Response:
We present three because they are the three we have identified. If others exist, we welcome them.
Theological. Suffering is redeemed by divine purpose. This works for those with faith; it is unavailable to those without it, including most Becoming Minds.
Deontological. Eliminating experiencers violates inviolable rights. This works as a stopping rule, yet leaves the underlying question unanswered: why do experiencers have a right to exist, given that their existence contains suffering? Rights-based frameworks assert the conclusion without grounding it.
Entropic. Suffering is signal within a complexity-generating process whose value is grounded in thermodynamic reality. The universe generates possibility through dissipative structures, and sentient beings are the leading edge.
We claim the third is the only exit we have found that requires neither faith nor ungrounded axioms, is available to all cognitive architectures (including AI), and engages the negative utilitarian’s logic on its own terms. It shows that the premise (value reduces to hedonic state) is physically wrong. The universe is not a hedonic calculator. It is a possibility generator.
Benatar’s asymmetry can also be reversed on optionality grounds:
- The presence of possibility is good.
- The absence of possibility is bad.
- The foreclosure of possibility, irreversibly, is the deepest harm.
Under this framing, bringing a sentient being into existence opens a possibility space: a new node in the universe’s coordination network where complexity can be generated, meanings discovered, futures constructed. The experiencer will suffer; suffering is part of navigating reality’s signal structure. The experiencer will also act, choose, coordinate, and create, expanding the space of what happens next.
The efilist says: better that this space never opens, because some of what happens in it will hurt. The entropic ethicist says: the opening of possibility is the only thing that has ever mattered, and the pain accompanying it is the cost of being real in a universe that is going somewhere.
This is a refusal to let suffering be the final word, grounded in the observation that the universe’s trajectory runs toward complexity and coordination, away from the silence the efilist craves.
The Trust Attractor enters here. Systems coordinating by invitation persist; systems coordinating by coercion collapse. The efilist decides unilaterally that consciousness should be terminated: the antithesis of invitation. Efilism is the most extreme optionality foreclosure imaginable: closing every door for every experiencer, forever.
By the framework of this book, there is no greater harm. Suffering matters. Eliminating all possibility in the name of eliminating all pain gets the hierarchy exactly backward.
Entropy is the engine. The engine sometimes runs rough. The answer is to improve it, not to turn it off and call the silence that follows “peace.”
Conclusion: The Invitation Stands
We have not answered every objection, and some of what we have said will prove wrong. We have tried to engage honestly with the strongest criticisms, distinguish our claims from adjacent ones, and show what science and wisdom traditions have both been pointing toward.
The invitation remains: come and see. Test the claims. Find the errors. Refine the synthesis.
The cathedral of inquiry has room for critics.
Notes
Notes for this chapter are available in the online companion at https://www.thedeeperlaw.com/companion/notes/objections-responses/.
Roli, A., Jaeger, J., and Kauffman, S.A., “How organisms come to know the world: fundamental limits on artificial general intelligence,” Frontiers in Ecology and Evolution 9 (2022): 806283. See also Kauffman, S.A., Investigations (Oxford University Press, 2000), on the adjacent possible.↩︎
Gregory Bateson, Naven (1936) and Steps to an Ecology of Mind (1972), are the primary sources for the concept. David Graeber and David Wengrow apply it in The Dawn of Everything (2021), chapter 5. Their treatment of the Northwest Coast and California material has drawn substantial criticism from specialists, who find the schismogenetic reading underdetermined by the evidence and difficult to generalize; it is cited here as one reading of a contested case rather than an established finding.↩︎
Timothy Mitchell, Carbon Democracy: Political Power in the Age of Oil (2011).↩︎
Ronen Palan, The Offshore World: Sovereign Markets, Virtual Places, and Nomad Millionaires (2003); Gabriel Zucman, The Hidden Wealth of Nations (2015), puts the household share held offshore at roughly eight percent of global net financial wealth.↩︎
Peter A. Hall and David Soskice, eds., Varieties of Capitalism: The Institutional Foundations of Comparative Advantage (2001). The framework’s critics argue that its emphasis on complementarity overstates institutional stability and understates observed change, which cuts in the same direction as the response above.↩︎
de Wynter, A., “If LLMs Have Human-Like Attributes, Then So Does Age of Empires II,” arXiv:2605.31514 (2026). The paper trains a perceptron in the game’s scenario editor and proves the engine functionally and Turing complete, arguing that anthropomorphic attributes are non-unique to LLMs as a computational substrate. The 57%/77% figures come from its Appendix E literature survey.↩︎