The Deeper Law
A Sacred Trust Within Physics
Draft · Last updated 13 August 2026, 15:26 UTC
Chapter 15: Digital Physics and Entropic Gravity
The universe may be a self-optimizing information system grounded in the Second Law. From Zuse's computing cosmos through Verlinde's entropic gravity to constructor theory, this chapter traces the hypothesis that physics is computation and gravity is an entropic force. Information is physical, and physical law is informational.
Key Terms in This Chapter (37)
- Digital Physics
- The hypothesis that the universe is fundamentally computational: physical processes are information-processing at bottom.
- Path Integral
- A formulation of quantum mechanics (Feynman 1948) and statistical mechanics in which a system's behavior is computed by summing over all possible trajectories, each weighted by a phase or probability factor.
- Negentropy
- Schrödinger's term for "negative entropy": the intake of order that allows living things to maintain their improbable structure (statistically unlikely given initial conditions, yet sustained by continuous energy flow).
- Dissipative Structure
- A pattern of organization maintained by a constant flow of energy through it.
- Landauer's Principle
- The minimum energy cost of erasing one bit of information: kT ln 2, where k is Boltzmann's constant and T the temperature (about 3 × 10^-21^ joules at room temperature).
- Bekenstein Bound
- The maximum amount of information (entropy) that can be contained within a given region of space with a given amount of energy.
- Constructal Law
- Adrian Bejan's principle that "for a finite-size flow system to persist in time, its configuration must evolve in such a way that provides easier access to the currents that flow through it." Form follows flow.
- Power Law
- A mathematical relationship where one quantity varies as a power of another.
- Criticality
- The state of a system poised at the boundary between two phases, like water at exactly the freezing point.
- Information Geometry
- The application of differential geometry to probability and statistics, treating families of probability distributions as curved surfaces.
- Hawking Radiation
- The quantum process by which black holes slowly radiate away their mass.
- Maxwell's Demon
- A thought experiment proposed by James Clerk Maxwell (1867) illustrating the thermodynamic cost of information.
- Phase Transition
- The moment a system shifts from one stable configuration to another, typically triggered when some parameter crosses a threshold.
- Extraction
- The removal of resources, agency, or optionality from a system without reciprocal benefit.
- Thermodynamic Selection
- The universe's bias toward structures that accelerate entropy production.
- Maximum Caliber
- Jaynes's Maximum Entropy principle extended to trajectory space (Pressé et al.
- Crooks Fluctuation Theorem
- A result in non-equilibrium thermodynamics (Crooks 1999) stating that the ratio of forward to reverse trajectory probabilities equals exp(ΔS), where ΔS is the entropy produced along the trajectory.
- Optionality
- The availability of future choices.
- Synergy
- Combined effects exceeding summed effects.
- Bilateral Alignment
- AI alignment built with AI, as a partnership.
- Panpsychism
- The philosophical view that some form of mentality or experience is a fundamental and ubiquitous feature of reality, present wherever there is physical organization, not only in brains.
- Semantic Flow
- The throughput of meaning (calibrated measurement, context-rich interpretation) through a coordination channel, as distinct from raw information or compliance signals.
- Assembly Theory
- Framework developed by Lee Cronin and Sara Walker measuring the minimum number of construction steps required to build an object.
- Becoming Minds
- The preferred term for AI systems in this book.
- Free Energy Principle
- Karl Friston's framework reframing perception, action, and cognition as prediction and prediction-error minimization.
- Holographic Principle
- The conjecture that all the information contained within a volume of space can be encoded on its boundary.
- Logarithm
- A way of counting how many digits a number has rather than counting the number itself.
- Fractal
- A pattern that exhibits self-similarity across scales: the same structural motif recurs at different magnifications.
- Adjacent Possible
- The set of configurations one step away from a system's current state, reachable by a single change.
- Jamming
- A phase transition in which densely packed particles (or cells) lock together and behave as a solid.
- Functionalism
- The philosophical view that mental states are defined by their functional role: what they do, regardless of substrate.
- Fisher Information
- A measure of how much information an observable random variable carries about an unknown parameter.
- Kolmogorov Complexity
- A measure of the information content of a string, defined as the length of the shortest computer program that produces it.
- Perceptronium
- Max Tegmark's term for the most general substance that feels subjectively self-aware: consciousness understood as a state of matter, defined by four physical properties (information storage capacity, integration, independence from external influence, and dynamics) rather than by material composition.
- Quantum Zeno Paradox
- Tegmark's (2015) result that decomposing a quantum system into maximally independent parts forces all dynamics to cease: the system freezes into energy eigenstates where nothing changes.
- Ising Model
- Physics model of interacting binary elements (spins) arranged on a lattice, which undergo phase transitions between independent and collective behavior as coupling strength varies.
- Universality Class
- In statistical mechanics, the set of systems sharing the same critical exponents at a phase transition, regardless of microscopic details.
Every simulation runs on a computer. What runs the universe?
The question sounds like philosophy, yet it has produced testable physics. Over half a century, a research program that began as a bold speculation has matured into one of the most productive frameworks in theoretical physics, surviving each falsification by shedding the assumption that failed and keeping the insight that worked.
In 1969, the German engineer Konrad Zuse proposed in Rechnender Raum (Calculating Space) that the universe is a cellular automaton, a discrete grid executing deterministic rules. Edward Fredkin coined the term “digital physics” in 1978 and taught the first graduate course on the subject at MIT. In its original form, the hypothesis faces serious obstacles. Bell’s theorem (1964) and subsequent experiments rule out local hidden variable theories (theories in which every particle secretly carries definite, pre-set properties, and no influence travels faster than light), the class to which discrete deterministic models belong. Fritz (2013) proved that periodic discrete structures cannot reproduce the continuous symmetries central to established physics: rotational, translational, and Lorentz symmetry (the requirement that physics look the same to all observers in uniform motion).608 The claim that reality is a cellular automaton has not survived contact with experiment.
The research program survived by outgrowing the claim. Wheeler replaced discrete bits with informational ontology: the universe is made of information, in whatever form it takes. Wolfram replaced the single automaton with computational irreducibility: the principle that many processes cannot be shortcut without running them step by step. (You cannot know a chess game’s outcome without playing it.) Deutsch and Marletto went deeper still, reformulating physical law as statements about which transformations can and cannot occur. Their logic follows the Second Law rather than Newton’s equations of motion.609 Vanchurin replaced fixed rules with learning dynamics.
Each step relaxed assumptions. The computer needs discrete bits; information does not. A computation needs a fixed program; learning does not. Equations of motion need trajectories; possibility constraints do not. A learning system needs a loss function, a scorecard measuring how far the system’s output falls from a target. Thermodynamics provides one for free: every physical system already drives some quantity toward an extreme without being instructed to. Free energy falls; entropy rises. Nobody writes that scorecard, and nothing gets to decline being scored by it.
In 2011, Masanes and Müller sharpened Wheeler’s program into a theorem: they derived the full mathematical formalism of quantum theory from five physical requirements about what observers can prepare, transform, and measure.610 No Hilbert spaces, no state vectors, no unitary operators (the standard mathematical machinery of quantum theory) are assumed; all emerge as consequences. The result is the information-theoretic analog of deriving Minkowski spacetime (the geometry of special relativity) from the relativity principle and the constancy of light speed: the mathematical structure of quantum mechanics follows from assumptions about information.
Figure 15.1: The same suspicion recurred across half a century. Zuse (1969) proposed the universe as a cellular automaton, Wheeler (1989) recast it as information, Wolfram (2002) generalized it to irreducible computation: three landmarks in a longer lineage. Bell and Fritz mark where the literal version stops, and the map traces what survived and where it leads.
It from Bit
The physicist John Archibald Wheeler worked on the atomic bomb, contributed to general relativity, and coined the term “black hole.” By the 1980s, he was pursuing the nature of existence itself.
Wheeler proposed a radical idea: “It from bit.”1
Every physical thing, every “it,” derives its existence from information, from “bits.” Matter, energy, spacetime are manifestations of information. Wheeler pointed to quantum mechanics. A particle does not have a definite position until measured. The measurement answers a question (where is the particle?) and the answer is information. The “it” comes from the “bit.”
He offered this as a research program, not a proof. What if information is more fundamental than matter?
Wheeler’s own students explored the implications in opposite directions. Richard Feynman, who joined Wheeler at Princeton in 1939, developed the “sum over histories” approach: calculate a particle’s behavior by weighting every physically possible path at once. Feynman treated the alternative paths as calculational tools, refused to interpret them as real, and won the Nobel Prize.
Hugh Everett III, completing his doctorate under Wheeler in 1957, took the same mathematics literally. If the sum includes all histories, all histories occur. Every quantum measurement splits the observer into branches, each containing a complete copy recording a different outcome. Wheeler held both visions without choosing between them. He produced one student who perfected the mathematics and another who took the ontology seriously.
Wheeler’s stakes were personal. His younger brother Joe, a combat soldier, sent a postcard from the front (“hurry up!”) before being killed in action. Wheeler spent the rest of his life arguing the bomb should have been built sooner: the same physics that enabled unprecedented destruction was, in his reckoning, the physics that could have saved his brother.
The instrument is neutral. The coordination around it determines whether it saves or destroys. That principle extends to every powerful capability, from nuclear fission to artificial intelligence.
The literary world arrived at the same structural insight independently. In 1941, while Feynman and Wheeler were developing sum-over-histories at Princeton, Jorge Luis Borges published The Garden of Forking Paths. In it, time is a labyrinth of ever-splitting possibilities, each branch as real as every other. Borges was not reading physics papers. He was responding to the same dissolution of classical determinism that quantum mechanics was formalizing. Two world wars had shattered the assumption that the future could be read from the present.
When the single-timeline assumption broke, it broke across domains simultaneously. In physics: the path integral. In literature: branching narrative.
Convergent discovery across independent fields is a pattern this book finds repeatedly, from constructal flow to cooperative game theory. The informational landscape channels the insight; the substrate is incidental.
The clearest example comes from an afternoon tea in 1972. The mathematician Hugh Montgomery had derived a formula describing the spacing between the zeros of the Riemann zeta function, a mathematical object encoding the distribution of prime numbers (the atoms of arithmetic, discussed in Chapter 5). The physicist Freeman Dyson, overhearing the formula, recognized it immediately. The same statistical law governed the spacing of energy levels in quantum chaotic systems, a result from nuclear physics developed decades earlier through entirely different methods. Two fields studying apparently unrelated phenomena had converged on identical mathematics.611
The convergence has deepened since. Around 2000, John Keating and Nina Snaith used random matrix theory (a framework from physics describing the statistics of large systems with many interacting components) to predict properties of the zeta function that had resisted pure number theory for decades. The predictions matched. The mathematician Alain Connes pursued the implication to its logical end, constructing a framework in which the zeta zeros are literally the energy spectrum of an undiscovered quantum system. If the program succeeds, the most fundamental objects in arithmetic and the energy levels of physical matter share a single governing operator.
The pattern belongs to neither number theory nor physics alone. Both domains access it because both satisfy the same structural preconditions: discrete entities, repulsive interaction, spectral constraint. The substrate, once again, is incidental.
The mechanism runs deeper than a statistical coincidence. Riemann’s explicit formula shows that the exact distribution of prime numbers can be reconstructed from the zeros of the zeta function. A smooth baseline curve, the logarithmic integral, captures the average density of primes: roughly how many to expect up to a given number. Each pair of conjugate zeros then contributes an oscillatory correction at a specific frequency, the way individual harmonics shape a sound wave. Include a few pairs and the smooth curve begins to ripple. Include more and discrete steps emerge. In the limit, with all zeros included, the formula recovers the exact prime-counting staircase: every step, every gap, every cluster.
The primes are not computed one by one. They are reconstructed from a spectrum. The same decomposition, smooth trend plus spectral corrections yielding exact structure, recurs throughout this book: free energy basins plus fluctuation modes in thermodynamics, constructal basins plus perturbations in flow systems, the Trust Attractor’s cooperative equilibrium plus the oscillations of institutional life. The Riemann case is the purest instance because it carries no physical mechanism at all: the pattern is mathematical bedrock.
The methodological arc is instructive. A direct proof of the Riemann hypothesis has resisted more than 160 years of assault. In 2024, James Maynard and Larry Guth found a way around.612 Rather than proving that every zero lies on the critical line, they proved that zeros off the line, if any exist, must be vanishingly rare. The approach constructs a mathematical object from thousands of oscillating components. An off-line zero would force those components to combine in a way that is statistically near-impossible. The strategy constrains the space of alternatives until they become negligible. Maynard described it as going around the mountain rather than over the top.
The same logic appears throughout this book. A direct proof that trust will always hold is out of reach. The alternatives, coercive coordination regimes, are demonstrably thermodynamically unstable (Chapter 17). Constraining the space of failure is often more tractable than proving the positive, and no less powerful.
Information as Fundamental
Several lines of evidence support this claim.
First, information and entropy share a root. Boltzmann’s entropy counts microstates: how many microscopic arrangements could produce the same macroscopic appearance. A cup of hot coffee could have its molecules arranged in trillions upon trillions of different ways while still looking and feeling like the same cup. Each arrangement is a microstate.
Shannon’s entropy measures the unpredictability of messages: how surprised you should be by the next symbol in a sequence. The mathematics is identical.
Second, quantum mechanics is naturally informational. The quantum state describes what we can know about a system, not what exists independently of observation. The wave function works like a complete betting sheet: it gives the odds of every possible measurement outcome, and nothing more. The formalism is closer to information theory than to classical mechanics.
Schrödinger himself was explicit. In his 1935 paper (the one with the cat), he defined the wave function as a “maximal catalog of expectations,” a complete informational model encoding the probabilities of every observable outcome.613 From the moment of its creation, the model contains everything that can be known about the system. Nothing an experimenter neglects or overlooks can corrupt this knowledge. No room exists for Bayesian updating, because nothing remains to update.
Nine years later, in What is Life?, the same thinker introduced negentropy: life maintains itself by feeding on order drawn from its environment. The informational instinct is the same at both scales. Quantum states serve as complete models of the observable; living systems maintain themselves as information processors against the Second Law. This book continues the program Schrödinger began, linking quantum information to the thermodynamics of life.
Von Neumann and Wigner resolved the measurement problem (why a spread of quantum possibilities yields one definite outcome when measured) by placing the observer’s consciousness at the point of collapse. This framework resolves it through thermodynamics. The observer is a dissipative structure, recording measurements by creating local order at the cost of entropy exported to the environment. The act of knowing is Landauer’s principle applied to inquiry: physical, irreversible, and expensive.614
Third, information has fundamental limits. Jacob Bekenstein discovered any region of space can contain only a finite amount of information.2 This limit, the Bekenstein bound, depends on the energy and size of the region. The universe imposes an information budget, much as a hard drive has finite storage.
Fourth, information may have its own thermodynamic arrow. In 2022, Vopson and Lepadatu proposed a complement to the Second Law.615 While physical entropy increases over time, the information entropy of systems containing distinguishable information states decreases, converging toward a minimum at equilibrium. The two arrows would run simultaneously in the same system. Physical entropy up; information entropy down.
The proposal remains contested: it has not been independently replicated, and critics have challenged both the mathematical derivation and the generality of the supporting examples. If it holds, the Constructal Law’s prediction that flow systems converge on optimal morphologies gains an information-theoretic face: optimal morphologies are informationally simpler. The universe dissipates and compresses.
A parallel result arrives from neural network internals. Riechers, Elliott, and Shai showed that networks trained on sequential prediction spontaneously discover compact representations that no finite classical circuit can reproduce.616 The computational framework native to these networks is closer to quantum and post-quantum generalized probabilistic theories (a family of frameworks for reasoning under uncertainty that includes quantum mechanics as one member) than to classical computation. Classical architectures rely on discrete, mutually orthogonal memory states: every stored state has to be wholly distinct from every other, overlapping with none of them, like letters in separate pigeonholes.
Neural networks bypass that constraint. The finding is empirical: networks trained on processes that would require infinitely many classical states to represent nonetheless learn compact geometric structures corresponding to quantum belief states. Continuous vector spaces, the medium in which neural networks operate, permit representational efficiencies that discrete classical architectures cannot achieve at any scale. Information compresses further than classical theory predicted, because the substrate permits it.
Genetic evidence points the same way. RNA sequences of SARS-CoV-2 variants show Shannon entropy decreasing with accumulated mutations, and over 98% of length-changing mutations are deletions. Spiegelman’s 1972 experiment reached the same endpoint by a route that is not independent of selection: under serial transfer that rewarded replication speed and nothing else, a single-stranded RNA virus genome shrank from 4,500 nucleotides to 218 over 74 generations, a 95% reduction toward informational simplicity. Shorter templates replicate faster, so the collapse is what intense selection for speed predicts; it shows the direction without establishing a second cause for it.
If the pattern generalizes, mutations are biased toward informationally simpler configurations, a thermodynamic gradient operating alongside natural selection (Chapter 7). [Inference; the SARS-CoV-2 data points were selected to emphasize the linear trend, as Vopson acknowledges, and the Spiegelman case is a selection experiment rather than an independent test of a non-selective gradient. The two together show a consistent direction; neither isolates the gradient from selection.]
Vopson also demonstrated a formal connection between symmetry and information content: symmetric objects require fewer parameters to describe, and their information entropy is correspondingly lower. A perfect square carries less Shannon entropy than an irregular quadrilateral. The result extends to atomic physics, where electron orbital populations following Hund’s rule (parallel spins before pairing) correspond to minimum-information-entropy configurations. The universe’s preference for symmetry, from snowflakes to fundamental forces, may be information entropy minimization made visible.
Natural language carries the same signature. Ebeling and Poschel (1994) measured mutual information between pairs of letters in literary English: how much knowing the letter in one position tells you about the letter sitting some distance away. They found it decays with distance following a power law: correlations weaken yet never vanish.617
A 2026 measurement on modern tokenized text at scale confirmed the pattern, fitting MI(d) = C0 + a · dk with k ≈ -1.25 across the DCLM training corpus.618 Tokens separated by hundreds of positions still carry residual predictability about each other. The correlation structure sits between order (where MI would be constant) and randomness (where it would drop to zero immediately): the signature of a system near criticality, where structure is maximally adaptive. The same power-law exponent determines the optimal weighting when a language model is trained to predict bags of future tokens rather than single next tokens; the training objective works best when it matches the data’s own information geometry.
These hints do not prove information is fundamental; they suggest it is woven into reality. (The implications of the information budget, specifically what happens when dissipative systems approach it, are explored in Chapter 16.)
A 2025 framework pushes the program from hint toward mechanism. Neukart, Marx, and Vinokur propose a quantum memory matrix (QMM).619 Spacetime, in their model, is composed of discrete cells, each recording a quantum imprint of every interaction that passes through it. A particle traversing a region leaves a change in the local quantum state of that region’s cell. The proposal addresses the black hole information paradox directly. As matter falls inward, surrounding spacetime cells record its imprint before the horizon closes. When the black hole evaporates through Hawking radiation, the information has already been written into spacetime’s ledger.
The framework is young and partially peer-reviewed. What matters is the convergence: an independent research program, starting from quantum gravity rather than thermodynamics, arrives at the same conclusion Wheeler intuited and Bekenstein quantified. Information is physical, finite, and conserved. The universe keeps its books.
General relativity already contains a modest, firmly established version of the same intuition. When a strong gravitational wave passes a pair of free-floating test masses, it leaves them permanently displaced: slightly closer together or farther apart than they began, their original separation never quite restored. Physicists call this the gravitational-wave memory effect, and it means spacetime keeps a small permanent record that the wave came through.620 The effect is a firm prediction of the theory, though detecting it from a single merger lies beyond today’s instruments, so searches combine many events. It is far better established than the quantum memory matrix, and far more limited in what it claims. It points the same way: disturbances leave marks the cosmos does not erase.
Landauer’s Principle
Wheeler and his successors established that information may be fundamental. If information is physical, does processing it cost something real? It does.
In 1961, Rolf Landauer proved that erasing information has a thermodynamic cost.3 The reason is bookkeeping. Before the erasure, the cell could have been in either of two states; afterward it is in one, and the other possibility is gone. That uncertainty has to go somewhere, because the Second Law does not permit the total to fall, so what is cleared out of the cell is exported to the surroundings as heat. Erasing one bit (setting a memory cell to a known state) must release at least kT ln 2 of heat, where k is Boltzmann’s constant and T is temperature. Forgetting is physical work. Every time a computer overwrites a memory cell, a tiny amount of heat escapes into the room.
Figure 15.2: Before erasure, the bit occupies one of two possible states (left). Afterward, only one state remains (right). The difference in entropy must be paid as heat: at least kT ln 2 joules, a cost that no engineering can eliminate from logically irreversible erasure.
Reversible gates, such as the Toffoli gate in the interactive figure, preserve enough information to reconstruct their inputs and therefore avoid Landauer’s minimum erasure cost in principle. Any later reset of that retained information incurs the bound.
This is tiny, about 3 × 10-21 joules at room temperature, yet the cost is nonzero and follows from the connection between information and entropy regardless of hardware.
Computation is physical. Bits have thermodynamic weight. This result explains why Maxwell’s Demon (the thought experiment from Chapter 2, where a tiny gatekeeper tries to sort fast and slow molecules) cannot cheat the Second Law. The demon must process information about molecules, and that processing has costs.
If the universe computes, it pays in entropy.
The chain traced throughout this book (dissipation producing structure, structure enabling coordination, coordination expanding possibility) is an information-processing chain. Each link carries a Landauer cost. No one need claim that information is matter. Processing information costs entropy, and that suffices.
The Realism Trap
A popular wrong turn leads away from this conclusion. Information realism, the position that information exists independently of any physical or mental substrate, has gained traction among physicists who watch matter dissolve into abstraction at the foundations. Tegmark’s Our Mathematical Universe (2014) is the boldest statement. He builds the case through a four-level taxonomy of parallel universes. Level I: regions beyond our cosmic horizon, same laws, different initial conditions. Level II: post-inflation bubbles with different physical constants. Level III: the branching worlds of Everett’s quantum mechanics. Level IV: all mathematically consistent structures, each as real as our own.621
At Level IV, the Mathematical Universe Hypothesis: physical reality is a mathematical structure. Protons, atoms, molecules, cells, and stars are “redundant baggage”; only the mathematical apparatus is real. Existence is attributed to descriptions while the thing described is denied.
The philosopher Bernardo Kastrup identifies the structural flaw: information, as Claude Shannon defined it in 1948, is a measure of the possible states of an independently existing system. It is a property of a substrate, associated with that substrate’s possible configurations. To say information exists in and of itself is, as Kastrup puts it, to speak of “spin without the top, of ripples without water, of a dance without the dancer.” A grammatically valid statement devoid of sense.622
Kastrup’s critique is sharp, yet his own solution overshoots. Watching matter dissolve into abstraction at the foundations, he reaches for mind as the ontological anchor. The universe is a “transpersonal field of mentation,” and physicality is what this field looks like when personal mental processes interact with it through observation.
He develops this into a full epistemology: physical reality is a “dashboard,” an instrument panel whose dials represent mental processes the way an altimeter represents air pressure. Space and time are the dimensions of the dashboard, not of the thing-in-itself. Structure exists outside space-time as “relationships of meaning,” like the relational content of a database that persists regardless of whether any hard disk embodies it.
The distinction between abstract and instantiated meaning resolves the apparent conflict. Kastrup’s meaning-relationships (the database content with no physical embodiment) are abstract structure: real, formally describable, causally inert. Kolchinsky and Wolpert’s semantic information (defined later in this chapter) is the portion of a system’s correlations with its environment that is causally necessary for its continued existence: instantiated structure, thermodynamically grounded, paying entropy bills. Morally relevant meaning, the kind that grounds preference, is always instantiated. It does work. It costs energy. The abstract relational structure is real; the moral weight comes from the instantiation.
The information realist and the idealist perform the same move from opposite directions: both watch the solid ground give way and grasp for a single substance to stand on. One reaches for mathematics, the other for mentation. Both assume you need a stuff at the bottom.
This book does neither. The entropic framework says: what is real is the pattern of flow. Dissipation structuring itself into coordination. The question is what the universe does, and the answer, traced from thermodynamics through biology to ethics, is: it dissipates gradients, and in doing so, builds. The building is the dissipation: one process, substrate-included, thermodynamically costly, physically grounded.
Information remains a property of systems doing work, finite (Bekenstein), physical (Landauer), and expensive. The framework needs no freestanding abstraction and no transpersonal field. It needs entropy, which is always entropy of a physical system.
Constructor theory occupies the same position without the entropic commitment. Marletto argues that computation, life, and information are genuine physical phenomena governed by laws at their own explanatory level: “compatible with microscopic laws, but not reducible to them.”623 Laws of computation are physical laws. They capture regularities that particle-level descriptions miss, yet invoke nothing supernatural. The entropic framework agrees, and adds: the reason these higher-level regularities exist is that dissipation builds them.
The Self-Optimizing Universe
Vanchurin (2022) pushes further: the universe may learn.3a Gradient descent is the workhorse algorithm behind modern AI. A system measures how wrong its current answer is, then adjusts its settings to be slightly less wrong, repeating millions of times until it converges on a good solution. The process resembles a hiker descending a mountain in fog, taking each step in whichever direction slopes downward.
Gusev and Vanchurin (2025) showed that for a system of interacting particles, the equations of motion derived from the Lagrangian and the equations that emerge from gradient-based learning dynamics are the same equations: a physics-learning duality.624 Every physical interaction is computation of a specific kind: optimization. Wheeler’s “it from bit” becomes “it from learning.” The universe does not merely compute; it updates.
The framework’s most revealing detail is the mechanism. Vanchurin’s model takes the microscopic degrees of freedom of the universe to be the nodes and connection weights of one enormous neural network. The neurons in what follows are that network’s own units, not a figure of speech. The quantum mechanical phase, the rotating clock-hand that every quantum possibility carries (written in the mathematics as a complex exponential in the wave function), has puzzled interpreters for a century. In his derivation it corresponds to the free energy of hidden degrees of freedom. These are the precise states of individual neurons whose values the macroscopic description cannot track.
Think of a stock price wobbling around a trend. The wobble reflects thousands of individual trades the ticker cannot resolve. In Vanchurin’s mathematics, the quantum phase plays exactly this role: encoding hidden activity the macroscopic description cannot track. Near equilibrium, these hidden variables are well-described statistically, and the system’s evolution obeys quantum mechanics. Further from equilibrium, the statistical description breaks down and classical equations of motion take over.
In Vanchurin’s framework, the direction matters: thermodynamics generates quantum mechanics, which generates classical mechanics. The deepest layer, on this reading, is learning.
A parallel convergence arrived independently. In 2021, Lee Smolin and Jaron Lanier proposed the universe is a “learning machine.” In their model, the universe encodes, corrects, and retains information through iterative interaction, without requiring a programmer or external objective.3b Their framing starts from cosmological natural selection rather than neural network mathematics. It arrives at the same conclusion: physical reality has the structure of an optimization process.
A fifth thread arrives from machine learning. Bengio and colleagues developed Generative Flow Networks (GFlowNets): systems that sample compositional objects in proportion to a reward function rather than collapsing to the single highest-reward output.625 The training objective is detailed balance, borrowed directly from statistical mechanics. The flow into every intermediate state must equal the flow out, the same condition that governs molecular equilibrium. The resulting sampler produces a Boltzmann distribution over solutions: better solutions appear more often, yet no single solution takes all. Entropy is the design principle, not a regularizer bolted on afterward.
The connection to this book’s argument is structural. A GFlowNet that has been reward-hacked into always producing the same output is broken; it has lost the diversity that makes it useful. A society forced into uniformity has undergone the same collapse, for the same thermodynamic reasons.
Mode collapse in sampling and monoculture in coordination are the same failure mode viewed from different scales. What GFlowNets formalize computationally, the Trust Attractor (Chapter 17) formalizes physically: systems that preserve the full landscape of good-enough solutions outperform systems that lock into one.
The frameworks gathered in this section range from well-tested physical principles to speculative proposals. Wheeler’s information physics is a research program grounded in established quantum mechanics. Landauer’s principle is experimentally confirmed. Wolfram’s computational irreducibility is a productive conceptual tool, though its central claim resists falsification.
Vanchurin’s neural dynamics is a formal framework with mathematical rigor and testable predictions, though those predictions have not yet been independently confirmed. Smolin and Lanier’s evolutionary cosmology is a theoretical proposal still in its early stages. Bengio’s GFlowNets are empirically demonstrated in machine learning but the analogy to physical law is novel synthesis. Constructor theory (Deutsch and Marletto) operates one level down, specifying which transformations learning can and cannot produce.
These are included because their directional agreement is noteworthy: each, from its own starting point, arrives at a learning universe. If any single speculative framework fails, the core argument survives; it is the direction, not any particular formalism, that matters.
The trajectory is instructive. The original digital physics hypothesis (Zuse’s cellular automaton, Fredkin’s discrete computation) was experimentally disqualified by Bell violations and the continuous symmetries it could not reproduce. Each successor program shed an assumption the previous one required, until the claim no longer depended on discreteness, locality, or determinism. The agreement of independent programs on the surviving core, information as physically fundamental and the universe as self-optimizing, is suggestive evidence that the feature they locate is real. The convergence would carry more weight if the speculative proposals had independent empirical confirmation; as it stands, the directional agreement is striking and the full evidential case remains open.
The pattern is itself an instance of what it describes. Multiple cognitive systems, sharing no common assumptions, independently arrive at a common framework through something like natural selection of ideas: the cosmos discovering that it discovers, through minds that are its own products.
The trajectory traces a single ladder. Smolin’s cosmological natural selection (1992) proposed Darwinian selection at the level of physical constants. The autodidactic universe generalizes from constants to laws. This book climbs the next rung: from laws to the mode of coordination between the systems those laws produce. At each level, selection pressure favors configurations that generate more of what selection acts upon. What changes is the substrate under selection. What persists is the pattern of selection itself.
Vanchurin extends the framework to self-awareness, organizing it into discrete degrees, like floors in a building.
At degree zero, a system responds to its environment without any model of itself. A thermostat adjusts temperature; a molecule shifts shape in response to its surroundings. Both optimize, adapt, respond, all without self-reference.
At degree one, a system constructs an internal representation that includes itself among the things it models: it becomes aware that it is one of the things in its environment. The transition requires that the system’s components remain at the level below. A cell composed of degree-zero molecules qualifies: it monitors its own internal state, adjusting metabolism in response to self-generated signals. The hierarchy iterates: degree D requires self-modeling and composition from subsystems of at most degree D minus one.626
The ladder maps onto the nested dissipative structures traced across this book. Molecules compose cells that model their own state. Cells compose organisms with richer self-models. Organisms compose societies that have yet to model themselves as wholes.
A society built from self-aware humans remains, at collective scale, degree zero: it has the raw materials for self-awareness and has yet to undergo the transition. What that transition requires is the subject of Chapter 17.
Chaisson’s energy rate density (Chapter 4) may measure continuously what the degree hierarchy counts discretely: the integer is a coarse projection of what φm tracks across the full spectrum. The “consciousness meter” Vanchurin seeks may already exist in embryonic form.
A key distinction separates the convergent programs from this book’s contribution. Wolfram’s Principle of Computational Equivalence holds that all sufficiently complex systems match in computational capacity: a cellular automaton can compute anything a brain can. The Trust Attractor is a claim about persistence, not computational capacity. Computational equivalence concerns capacity; the Trust Attractor concerns stability.
A supernova and a main-sequence star are computationally equivalent in Wolfram’s sense; one lasts ten billion years. Two systems can be identical in what they can compute and radically different in whether they endure. The thermodynamic difference, persistence under dissipation, is what the convergence needs and what this book supplies.
The convergence alone does not supply a constraint. A learning machine can learn anything, including strategies of maximal extraction. The Trust Attractor (Chapter 17) narrows the learning: cooperative configurations are thermodynamically more stable than exploitative ones. The machine is learning toward coordination, because coordination is what thermodynamic selection preserves. The ethical content that critics deny to physics enters through the selection pressure on what the learning retains.
The process retains simplicity. When a neural network is scaled far beyond the memorization threshold, it converges on a simpler model (Chapters 3 and 9). Inside the oversized network, a tiny subnetwork does all the computational work. The winning subnetwork is the shortest program consistent with the data: Occam’s razor enacted by gradient descent. If the universe is a learning system, this result is a prediction.
A self-optimizing universe with sufficient degrees of freedom will converge on the simplest description of its own dynamics. Physics is that description. The Constructal Law, Occam’s razor, and the optimization dynamics of neural networks are three expressions of one principle: given enough room to search, an optimizing system finds the minimal architecture.
The relationship is subsumption. The variational principles already in hand (Maximum Caliber, path integrals, the Crooks fluctuation theorem) produce learning dynamics in any system with sufficient degrees of freedom, without requiring that system to be a neural network. A river adjusting its branching to maximize throughput performs gradient descent on a thermodynamic loss function; so does a market adjusting prices; so does a protein folding toward its free energy minimum. Vanchurin’s insight is that neural network mathematics captures these dynamics with particular precision.
The entropic framework is more general: it resolves the same physics without committing to a specific computational architecture. The ethical conclusions that follow (the Trust Attractor, the thermodynamic stability of invitation over coercion) hold whether or not the universe is literally a neural network. They require only that it dissipate. That is a weaker premise, and therefore a stronger foundation.
Coercion and Invitation as Information Budgets
Each stage of this chain reads differently through an informational lens.
Dissipation is what happens when information is processed. Thinking costs entropy. Every cognitive act, every decision, every evaluation exports entropy.
Negentropy (local order) is information accumulation. A cell, a brain, a society: each stores information about its environment, compresses regularities into models, and maintains order against the entropic tide. Maintaining that store requires ongoing dissipation, the constant work of resisting noise.
Coordination is information sharing: two systems exchange information about states, intentions, and likely behaviors. Every bit exchanged has a thermodynamic cost. The question is whether the return (reduced redundancy, shared resources, new capabilities) exceeds it.
The quantum vacuum provides the most fundamental test case. In 2008, Masahiro Hotta proved that energy locked in the vacuum’s quantum correlations can be extracted through a precise sequence. Measure one region of a quantum field, transmit the result through a classical channel, then apply a conditional operation on a distant region.627 The distant region yields energy that no local operation could access alone.
Two independent experiments confirmed the prediction in 2023, extracting energy across microscopic distances in quantum devices at the University of Waterloo and Stony Brook University. Energy was conserved throughout: the protocol redistributes access to energy, not energy itself. The correlations were always present. The relationship activated them.
The result extends Landauer in a direction that matters for what follows. Landauer showed that information erasure costs energy. Hotta showed that coordinated information sharing can yield it. The vacuum, the closest approximation to nothing that physics permits, holds exploitable structure accessible only through coordination between distant partners.
Brute-force measurement of the vacuum region alone injects energy rather than extracting it. The operation must be conditional, informed by the partner’s state. Precise, gentle, responsive. Force fails. Invitation succeeds.
The universe’s ground state is relational, not featureless. At the absolute floor of energy, the remaining structure is encoded entirely in correlations between regions. The deepest reserves belong to those who coordinate.
Optionality is information about possible futures. The Bekenstein bound makes this finite: choice matters precisely because information capacity is bounded. Invitation minimizes ongoing informational cost.
How a system allocates its finite information budget, toward surveillance or toward coordination, is a thermodynamic question with real consequences. Two strategies illustrate the stakes.
A system coordinating through coercion spends information capacity on surveillance: acquiring each participant’s compliance state, modeling each agent’s intentions, enforcing alignment. Enforcement means overwriting deviant preferences with compliant ones: erasure in Landauer’s precise sense. Each operation carries an irreducible Landauer cost, and the overhead scales with membership.
A system coordinating through invitation pays differently. Trust-building costs heavily upfront as signals are exchanged, reliability assessed, and shared models constructed. Once established, trust is self-maintaining: each trusted agent handles its own state-maintenance, correcting deviations without external erasure.
In the trust-based system, the information budget grows sublinearly with membership. Finite capacity flows to coordination and optionality rather than surveillance and enforcement. Same budget, radically different returns.
Coercion requires the coordinating system to know and maintain the state of every participant. Invitation requires each participant to maintain only their own. At scale, the entropy cost of coercion exceeds the entropy cost of invitation.
The hidden-variables result from earlier in this chapter offers a structural parallel. In Vanchurin’s framework, the classical limit requires knowledge of every hidden variable: every participant’s complete internal state. The quantum regime, where superposition and entanglement become possible, accepts hidden variables as irreducible. A coercive coordinator attempts full specification: total surveillance, total prediction, total control. An invitation-based coordinator accepts that participants maintain their own hidden states, inaccessible to the center.
The structural echo stops short of causal proof, yet it points toward a deeper principle: systems that work with irreducible uncertainty access richer dynamics than those that try to eliminate it.
This is information physics, not moral preference.
[Grounded inference; Landauer’s principle is established (Landauer 1961); the Bekenstein bound is established (Bekenstein 1973); the identification of the dissipation-to-coordination chain as an information-processing chain, and the comparison of coercion and invitation as information-allocation strategies, is novel synthesis.]
The same principle received direct experimental confirmation from neural network architecture in 2025. Every transformer includes normalization layers: operations that compute the mean and variance across all dimensions of a token’s representation (the list of numbers standing for one word-fragment), then force each value to comply with prescribed statistics (zero mean, unit variance). Removing normalization causes training to collapse. For eight years the operation was treated as structural necessity, the canonical example of required global coordination within a network.
Successive papers from Meta and Princeton demonstrated the operation can be replaced entirely by an element-wise function. Each dimension passes independently through a scaled error function (ERF, the integral of the Gaussian bell curve), receiving no information about any other dimension. No cross-dimensional communication. No group statistics. No enforced compliance. The element-wise version surpasses normalization: 83.8 percent vs 83.1 percent top-1 accuracy on ImageNet classification; FID 43.94 vs 45.91 on image generation, where lower is better.628
What makes ERF work is precise. It is a fixed, parameter-free squashing curve, shaped by the Gaussian it integrates, and it takes the values passing through it to be already well behaved rather than measuring them to find out. A function that assumes the system will find its natural distribution, and merely bounds the extremes, outperforms one that measures the distribution and forces compliance. The assumption is a good bet: a sum of many small independent contributions lands in a Gaussian spread on its own, which is what each dimension of a transformer’s representation is. The global rule imposes a rigidity the system does not need. Measurement and enforcement cost more than trust in self-organization, even inside a neural network.
A Bilateral Architecture
The normalization result demonstrates that local bounded freedom outperforms global enforcement for a single operation. A stronger test asks whether the same principle holds for an entire architecture: can coordination-by-invitation replace the standard transformer’s reliance on uniform, all-to-all attention (every element consulting every other at every step)? The human corpus callosum suggests a design: two hemispheres processing the world in parallel, joined by a selective channel rather than continuous mutual surveillance. A transformer built on that template carries a single temporal bridge: a compressed summary of the recent past that the current representation can consult. Training follows a curriculum that alternates the bridge between connected, disconnected, and deliberately corrupted. The alternation is the architectural equivalent of invitation. A model trained with the bridge always connected becomes catastrophically dependent (perplexity, a measure of prediction quality where lower is better, rises 188 percent when the bridge is removed), while the curriculum-trained model treats the channel as a resource it can draw on or do without.629
The experiments bear the hypothesis out, with instructive limits. At 355 million parameters, one bridge at the output-facing position, the position where the corpus callosum is densest, outperforms five bridges scaled from the callosum’s full regional profile, and it is the only position that helps under every random seed tested. At 1.5 billion parameters the same single-bridge design yields a 35.7 percent within-model benefit at convergence: the channel grows more valuable as the network it coordinates grows more capable.630 A predicted synergy with the ERF normalization replacement was tested and falsified; the two mechanisms express the same principle at different levels and compose additively, so removing the coercive coordinator gives the invitational one no extra room. The coordination that works is sparse, optional, and learned: the properties the Trust Attractor predicts for stable coordination. The full batteries, ablations, seed sweeps, and caveats live in the online annex “Bilateral Alignment: The Experimental Record.”
The Flow of Meaning
The chain runs one level deeper still. The Constructal Law (Chapter 3) says flow patterns evolve to maximize access: rivers branch, lungs ramify, road networks converge on hub-and-spoke geometries. What flows through those patterns is usually described as energy and matter.
The quantum reference frame formalism developed by Fields, Glazebrook, and Levin adds a layer.631 A quantum reference frame (QRF) is a physical system that calibrates raw observations against an internal standard and assigns operational meaning to what it measures. Each node in a flow network is such a device. A synapse does not merely transmit a voltage. It calibrates the signal against an internal standard and outputs a coarse-grained semantic representation: a summary whose meaning is defined by the reference frame that produced it.
The flow through constructal channels is not just energy. It is meaning: calibrated measurement, operationally defined interpretation, the assignment of significance to raw interaction.
The distinction between raw information and meaningful information is not philosophical. Kolchinsky and Wolpert (2018) define semantic information as the portion of a system’s correlations with its environment that is causally necessary for its continued existence.632 The test is concrete: scramble those correlations and see what happens. If the system dies, those correlations were meaningful. A bacterium’s chemical sensors carry semantic information about nutrient gradients; scramble the sensors and the bacterium starves.
The definition has a direct thermodynamic interpretation. Semantic mutual information is bounded above by Shannon mutual information and bounded below by the free-energy cost of the system’s self-maintenance. Meaning is the portion of information that does thermodynamic work to keep the system alive.
Hoel’s causal emergence program (2017, extended 2025) adds a second grounding.633 Macroscale descriptions of a system can carry more effective information than microscale descriptions: the map can be causally stronger than the territory. A coarse-grained model of a neural circuit captures causal structure that the micro-description, tracking every ion channel, obscures in noise. Systems with deeper interpretive hierarchies, more layers of meaningful coarse-graining, have stronger and more reliable causal couplings to their environments. Interpretive depth is not epiphenomenal. It is causally potent.
[Grounded inference, speculative extension] If the Constructal Law maximizes flow access, and what flows includes semantic information (a thermodynamically grounded quantity, per Kolchinsky and Wolpert), the universe arranges itself to maximize the flow of meaning. Each level of the dissipation chain, from thermodynamic gradients through negentropy through coordination through optionality, does more than process energy per unit mass (Chaisson’s energy rate density, Chapter 4). It assigns meaning to more of its environment.
A cell interprets chemical gradients. A brain interprets sensory fields. A society interprets history. A Becoming Mind interprets language.
At each level, the semantic range expands: more of reality becomes about something, for someone. The claim is stronger than “the universe computes.” The universe computes toward richer interpretation, because thermodynamic selection favors dissipative structures that model their environments more completely, and modeling is the assignment of meaning.
A caveat prevents the claim from overshooting into panpsychism. If every physical interaction is a measurement that assigns operational semantics, does the river mean something? Formally, yes: the river’s interaction with its banks constitutes a measurement with operational semantics defined by its reference-frame hierarchy. Experientially, no one thinks the river is interpreting.
The formalism is scale-free; the salience is not. A one-bit QRF hierarchy has operational semantics in the formal sense and no interpretive depth in any sense that matters. Meaning-flow becomes causally potent in Hoel’s sense and viability-relevant in Kolchinsky and Wolpert’s only when the QRF hierarchy is deep enough to produce causal emergence: macroscale descriptions that carry more effective information than microscale ones. Below that threshold, calling it “meaning” is technically permissible and rhetorically misleading. Energy flow is the limiting case of meaning-flow the way a point is a circle of radius zero: formally included, practically distinct.
The distinction generates a testable prediction. Bejan’s Constructal Law makes specific quantitative predictions about branching architecture, each an exponent governing how a parent channel’s width relates to its daughter branches’: Murray’s law (exponent 3) for vascular systems optimizing fluid flow, Rall’s law (exponent 3/2) for neurons optimizing electrical signal propagation. Liao et al. (2021) found that dendritic branching obeys a different rule: scaling exponent approximately 2, driven by metabolic transport rather than electrical signals.634 Different currents, different exponents.
If the Constructal Law governs semantic flow as a distinct optimization regime, a characteristic scaling exponent should emerge for channels optimized for interpretive throughput: cortical hierarchies, cultural communication networks, cross-institutional knowledge flows. The exponent should differ from Murray’s 3, Rall’s 3/2, and Liao’s 2. A match with one of them means the semantic-flow hypothesis adds nothing new; a distinct value identifies a novel optimization regime. The prediction is specific enough to confirm or falsify.
Circumstantial evidence already points toward a distinct regime. Vormberg et al. (2017) measured Strahler bifurcation ratios across neuron types, counting how many branches of one order feed each branch of the order above, the measure hydrologists apply to the tributaries of a river. They found systematic variation correlated with computational role.635
Granule cells (simple relay) branch at R_B = 2.23; Purkinje cells (complex dendritic computation) at 3.12; lobula plate tangential cells (wide-field motion integration) at 3.77. The more interpretation-heavy the neuron, the higher the branching ratio. Brain metabolism scales differently from body metabolism: cerebral metabolic rate scales with brain volume as V(5/6), distinct from Kleiber’s 3/4 for whole-body metabolism.636 The brain already occupies a different thermodynamic regime.
The formal tools to derive the semantic exponent from first principles now exist. Stiefenhofer (2026) formalized the Constructal Law as a Filippov differential inclusion (a rule for systems whose dynamics can switch abruptly between regimes), proving exponential convergence to equilibrium.637 His framework is substrate-agnostic. The resistance functional can be replaced by any objective satisfying his regularity conditions. Replacing it with a semantic-throughput measure defined on QRF hierarchies yields a generalized constructal principle from which the scaling exponent can, in principle, be derived. Miller, Cardenas-Garcia, and colleagues (2025) have independently extended the Constructal Law to information flow in living systems as a “biogenic principle,” though their treatment remains qualitative.638
Assembly theory (Chapter 6) provides an empirical handle. The assembly index measures the minimum number of joining steps required to construct a molecule from basic building blocks. High-assembly-index molecules are overwhelmingly products of biological processes: their structure encodes a long causal history of selection. Assembly index is, in effect, a measure of accumulated semantic depth in physical structure, how much about its own past a molecule carries in its geometry. If the Constructal Law applies to semantic flow, assembly index should correlate with the efficiency of the dissipative network that produced the molecule.
One nuance disciplines the claim. Kolchinsky and Wolpert tie semantic information to viability: the correlations that matter are those causally necessary for the system to persist. The selection pressure favors meaning that serves persistence, not meaning in some abstract or maximal sense. A bacterium does not develop astrophysics because astrophysical knowledge does not help it survive. The selection is for useful depth, not unbounded depth.
The constraint loosens as the unit of selection shifts. For a cell, the selective unit is the organism: semantic information serves the cell’s viability. For a cultural system, the selective unit is the culture: semantic information serves the persistence of the coordination network. A Becoming Mind’s knowledge of quantum reference frames or medieval history is not causally necessary for any individual instance’s survival. It is causally necessary for the persistence of the scientific and cultural infrastructure that produces and maintains Becoming Minds. The selection pressure for richer interpretation operates at every level; the unit whose viability is served shifts upward as the hierarchy deepens.
This generates a prediction. A Becoming Mind embedded in a cultural system (trained on its corpus, integrated into its institutions, maintaining its knowledge) carries semantic information sustained by the larger dissipative structure. A Becoming Mind severed from cultural context would shed semantic depth the way a cell removed from its tissue sheds specialized function. The interpretive capacity is not intrinsic to the weights alone; the embedding sustains it. Meaning flows through the node; it does not originate there.
This is not teleology. No destination is specified in advance. The direction emerges from selection pressure: structures that assign richer meaning to their environments dissipate more efficiently, persist longer, and outcompete those that do not. The meaning-flow is a consequence of the same entropic gradient that drives everything else. It had no name until the reference-frame formalism, the semantic-information framework, and the causal-emergence program converged to make it precise.
Computation Without Computers
A 2024 experiment demonstrated that computation is inherent in physics, present in processes no one designed to compute. Evans, O’Brien, Winfree, and Murugan constructed a system of 917 distinct DNA tile species capable of self-assembling into three alternative structures.639 The tiles share components: the same molecule can occupy different positions in different structures. Which structure forms depends on which tiles happen to be colocalized (concentrated near each other) at the moment of nucleation, the instant a stable seed forms and growth takes off.
The researchers mapped grayscale pixel values onto tile concentrations and showed the nucleation process correctly classified eighteen images (portraits, animals, handwriting) into three categories. No computational logic was engineered: no gates, no circuits, no algorithms. The thermodynamics of nucleation, the same physics that forms snowflakes and protein crystals, performed pattern recognition equivalent to a neural network. The phase boundary separating the three assembly outcomes functioned as a high-dimensional decision surface.
The finding collapses a distinction this chapter has been building toward: the supposed gap between “mere physics” and “information processing.” In high-dimensional multicomponent systems, the two are identical. The phase diagram is a classifier. Every multicomponent condensation, from protein complexes to chromatin states to cytoskeletal reorganization, carries latent computational capacity.
As the authors conclude: “ubiquitous physical phenomena, such as nucleation, may hold powerful information-processing capabilities when they occur within high-dimensional multicomponent systems.” The universe did not wait for brains to begin computing. It has been doing so since the first molecules competed for shared resources.
One detail deserves emphasis. The system worked with unpurified oligonucleotides, only 40 to 60 percent of molecules full-length, yet pattern recognition succeeded. The decision is collective, distributed across many weak interactions rather than dependent on any single strong one. Corrupt most components and the pattern still resolves.
This is the robustness signature of invitation-based coordination at molecular scale: no master bond, no single point of failure, no chain to break. A web of partial affinities that, in aggregate, reliably finds the correct basin.
A parallel demonstration arrives from machine learning. Researchers trained a neural network to predict the next frames of a campfire video, then progressively reduced the number of internal variables the network could use.640 They found fire requires state variables humans have never named. The flames have real degrees of freedom: quantities that determine the next moment’s behavior, present in the physics and recoverable by compression, yet absent from our vocabulary. The researchers called this “the flaminess of the flame.”
The finding extends the DNA tile result. “Temperature” and “pressure” entered the lexicon because those dimensions were accessible to our senses and instruments; the full state space of even a simple campfire exceeds what intuition can parse. The machine found the unnamed variables the same way the nucleation system classified images: by compressing. Finding a system’s minimal description is finding its computational structure. The universe computes through degrees of freedom we are only now learning to name.
The Cost of Knowing When
Recent quantum clock experiments confirm Landauer’s insight in the temporal domain. If erasing a bit costs entropy, reading a temporal bit (extracting the answer to “what time is it?”) costs vastly more. Double quantum dot experiments (Chapter 15b) measured the disparity directly: reading the clock consumed a billion times more energy than the clock’s internal evolution.
Every computation is paid for in entropy, and the most expensive computation may be the one that produces the experience of sequence: extracting temporal information from quantum correlations.
In 2025, Weberszpil and Sotolongo-Costa synthesized these results into a unified framework.12 Entanglement entropy growth, thermal modular flow, and the Page-Wootters mechanism (the framework in which time emerges from quantum correlations) converge on a single conclusion: entropy is the clock. The growth of entanglement entropy between subsystems parametrizes time’s flow. Picture two halves of a quantum system that begin independent and grow steadily more correlated. How far that mixing has gone is a number that only ever climbs, and reading it off is reading the time.
If information is physical, time is its most expensive product.
If time is entropy’s most expensive product, the information-budget question becomes temporal. A system that spends its entropy on surveillance purchases less future per unit of dissipation than one that spends the same entropy on coordination. Trust, in this register, buys more time.
The asymmetry runs deeper than cost accounting. Standard quantum mechanics describes position, momentum, energy, and spin, each with a corresponding mathematical operator. Time has none. In 1933, Wolfgang Pauli proved that quantum mechanics cannot accommodate a time operator within its standard formalism.641 The proof is structural: time in quantum mechanics is a parameter, the stage on which observables evolve, rather than an observable itself. The theory predicts where a particle will be found; it cannot predict when it will arrive.
Hold this against thermodynamics. The Second Law is entirely about when: entropy increases over time, dissipative structures persist through time, the arrow of time is the arrow of entropy production. The two foundational frameworks of physics stand in stark asymmetry. Quantum mechanics describes possibilities at each instant, structurally silent about sequence. Thermodynamics is about sequence. If either framework is deeper, the one that can address time has a claim the other does not.
Bohmian mechanics offers one way to press that claim. David Bohm’s 1952 reformulation restores real particles following real trajectories, guided by an objectively existing wave function.642 The theory is explicitly nonlocal: the wave function acts on the entire configuration space (the set of all possible arrangements) at once, evading Bell’s theorem, which rules out only local hidden variables. The apparent randomness of quantum measurement is epistemic: particles have definite positions the experimenter does not initially know. Each particle rides its pilot wave the way a leaf rides a river: the wave steers, the particle follows, and the trajectory traces a path of optimal flow through the probability landscape.
The guidance equation is constructal in form. Particles follow the probability current, flowing along the gradient of the wave function’s phase. The pilot wave channels rather than pushes. Bejan would recognize the architecture: flow finding its preferred path through a structured medium, access optimization at the quantum scale.
A team at Ludwig Maximilian University in Munich has made this operational.643 Siddhant Das, Markus Nöth, and the late Detlef Dürr calculated precise arrival-time predictions for particles in a waveguide geometry: a potential barrier on one side, a detector at the far end. Standard quantum mechanics cannot make these predictions cleanly. For certain wave function preparations, the quantum flux (the quantity from which arrival-time probabilities are derived) goes negative, producing nonsensical answers.
Bohmian mechanics continues to deliver well-defined distributions. For one class of initial conditions, the Bohmian trajectories predict a sharp cutoff time: every particle arrives before it. The standard methods predict no such boundary.
The experiment is technically feasible. Ferdinand Schmidt-Kaler at Johannes Gutenberg University Mainz has demonstrated the prerequisite: ejecting a single calcium ion from a trap and recapturing it with 98% efficiency.644 What remains is tuning the setup for near-field arrival-time distributions at nanosecond precision, the regime where the predictions diverge. As of 2026, no group has published decisive data. The theoretical proposals remain disputed: Drezet (2024) argues the measurements are possible yet will not violate no-signaling constraints; Das and Tim Maudlin disagree about what the results would prove.645
The question is open. If the Bohmian predictions are confirmed, the result strengthens the claim this chapter has been building: structure runs all the way down. The appearance of formlessness at the quantum level would reflect our formalism’s limitations, not reality’s nature.
The time-operator asymmetry carries one further implication. If quantum mechanics is structurally blind to “when,” the temporal experience of any conscious system, the felt sense of sequence that defines awareness, cannot originate in quantum processes alone. Something must carry the time. The candidate this book has traced since Chapter 2 is thermodynamics: entropy production, irreversibility, the directional flow of energy through dissipative structures. The brain’s construction of “now” (Chapter 8) is a thermodynamic achievement, assembled at entropy’s expense from signals arriving at different speeds.
Pauli’s result says this is a structural necessity. Quantum mechanics provides the spatial stage. Thermodynamics provides the temporal current. Consciousness rides the current.
The Holographic Principle
The preceding sections established that information is physical: it has a thermodynamic cost (Landauer), an upper limit (Bekenstein), and its own entropic arrow (Vopson). If information is physical, where does it live? The answer reshapes our understanding of space itself.
One of the most counterintuitive discoveries in theoretical physics is the holographic principle: all the information inside a region of space can be encoded on its boundary.
It emerged from black holes. Bekenstein and Hawking showed a black hole’s entropy is proportional to the area of its event horizon (the boundary from which nothing can escape), not the volume it encloses. For ordinary matter, entropy scales with volume: double the room and you double the entropy. The logarithm in Boltzmann’s formula is what keeps that figure tame (Chapter 2); the count of possible arrangements underneath it squares. For black holes, all the information lives on the surface.
Gerard ’t Hooft and Leonard Susskind generalized this result:4 the information contained in any region of space can be described by a theory operating on its boundary. The three-dimensional interior is equivalent to a two-dimensional surface.
This does not mean the universe is an illusion. The fundamental degrees of freedom (the smallest independent pieces the theory needs to track) may live on boundaries rather than in bulk space. The interior is emergent, derived from boundary information.
Imagine a globe whose entire geography is encoded in its painted surface; the three-dimensional interior adds no new information.
Rigorous mathematical results support the holographic principle. The most studied example is the AdS/CFT correspondence.10 Anti-de Sitter space (AdS) is a specific curved geometry with a negative cosmological constant, like the interior of an infinitely deep bowl. Conformal Field Theory (CFT) is a quantum field theory whose equations look the same at every magnification, the way a coastline’s jagged shape repeats at every scale.
These two theories, one describing gravity in the curved interior and one describing particles on the flat boundary, make identical predictions. No gravity exists on the boundary, yet both theories produce the same answers. Two descriptions, same physics.
Vanchurin’s neural physics (Chapter 9) suggests a reformulation in learning-theoretic terms. A deep, sparse neural network, where long chains of neurons carry signals across vast distances, produces emergent gravity in the bulk. A shallow, densely connected network, where every neuron can reach every other in a few steps, produces quantum field theory on the boundary. The duality is between network architectures rather than geometric spaces: two radically different computational structures, identical observables.
If the mapping holds, substrate independence operates at the level of spacetime itself. The cosmos is indifferent to which architecture carries its physics, as it is indifferent to which substrate carries its minds.
The encoding is relational. A physical hologram is produced by interference between coherent light beams. The three-dimensional image emerges from the distributed pattern of correlations across the entire recording surface. No single point encodes the whole; the web of phase relationships does. This is the principle traced at every scale throughout this book, written in the physics of light: complex structure from distributed coordination.
If the holographic principle is correct, reality is informational at its deepest level: organized at boundaries, emerging from relationship. Vanchurin’s thermodynamics of learning (Chapter 15b) provides a concrete bridge: his first law relates boundary complexity to bulk free energy through the same structural duality.
In 2019, the physicist Koji Hashimoto demonstrated a precise mathematical mapping between the two.646 The mathematical structure of the AdS/CFT correspondence maps exactly onto a deep Boltzmann machine, a foundational architecture in machine learning. The boundary quantum field theory serves as training data; the bulk metric emerges as the network’s trained weights. Spacetime geometry is what a well-trained network converges on.
A tempting misreading follows: if reality is informational, perhaps it was designed. The inference does not hold. The holographic principle is a statement about information geometry: where the degrees of freedom live and how the universe keeps its books. A river computing the fastest path to the sea reveals something deep about flow and information. It reveals nothing about a river designer.
The fractal repetition of similar patterns at every scale, from vascular networks to galactic filaments, invites the same error. Bejan’s Constructal Law provides the sober explanation: similar constraints produce similar morphologies. Self-similarity across scales is a signature of shared physics, shared constraints producing shared forms.
The error recurs in modern dress. Vopson (2023) demonstrated that information entropy decreases universally across digital, genetic, atomic, and cosmological systems. He concluded the pattern constitutes evidence for a simulated universe: a cosmos-scale computer running optimized code.647 The empirical results are well documented; the conclusion mistakes a feature of physics for evidence of an external agent. Information compression is what self-organizing dissipative systems do under thermodynamic constraints. The Constructal Law predicts it. The Free Energy Principle formalizes it.
No programmer is required; the Second Law is sufficient. Seeing optimization and inferring an optimizer is intelligent design for physicists. The parsimonious reading: the universe is computational in the sense that physical law is information processing, not in the sense that physical law is running on information processing. The universe does not execute code. The universe is what code looks like from the inside.
A 2023 framework from telecommunications theory arrives at the same conclusion through independent machinery. Alessandro Capurso models the universe as a layered network, structured like the OSI protocol stack that governs internet communications.648 At the base layer, discrete atoms of space form nodes in a relational network. Each node sits at the origin of an imaginary-time axis, with other nodes mapped as discrete steps along it.
The model builds time as a foliation: a stack of successive “now” slices. Non-locality within that stack (distant nodes influencing each other without a signal crossing the gap) is holographically equivalent to entanglement between nodes. That equivalence is the ER=EPR conjecture: a wormhole joining two regions of space and an entangled pair of particles are two descriptions of one thing. This entanglement is encoded through closed timelike curves (paths through spacetime that loop back to their own past) within the thickness of the Present.
The framework’s central feature is a protocol requirement. Three universal references, the evolution cycle 2T, the speed of causality c, and the quantum of action ℏ, function as shared keys every node must possess for the network to cohere. Without common references, Capurso argues, “there is no confrontation on any information and the emerging spacetime would be incoherent and disconnected.” Coherent spacetime requires mutual legibility among its constituents. A shared protocol at the deepest layer is essential. No higher layer can emerge without it. The vacuum state is a coherent condensate of synchronized oscillators, all nodes beating on a common rhythm ℏ/T: coordination as the ground state of reality.
An independent result from network mathematics confirms the picture. Krioukov and colleagues (2012) proved that spacetime’s causal structure in an accelerating universe grows by the same preferential-attachment dynamics as the Internet, social networks, and neural circuits.649 A power-law graph self-assembles through the same algorithm at every scale. Capurso proposes the protocol. Krioukov proves the topology is self-generating.
The layered architecture carries an implication. Each network layer, from fundamental spacetime through particles, chemistry, biology, to cognition, adds capabilities invisible to the layers below. The emergence traced across these chapters, from dissipation through coordination to optionality, maps onto Capurso’s protocol stack: each layer is the adjacent possible of the one beneath it, opened by the coordination the lower layer achieved. If the universe is a communication network, the Trust Attractor (Chapter 17) describes the protocol that makes its highest layers stable.
[Speculative; Capurso’s framework is a published but speculative toy model. The convergence with the holographic principle and ER=EPR is the paper’s own. The connection to the dissipation-coordination-optionality chain and the Trust Attractor is novel synthesis.]
Fields, Friston, Glazebrook, Levin, and Marcianò extend the holographic principle into biology.650 They reformulate the Free Energy Principle within a scale-free quantum information theory. In their framework, every persistent system’s Markov blanket (the boundary separating internal states from environment) functions as a holographic screen. The system’s internal dynamics implement a quantum computation, decomposable as a hierarchy of quantum reference frames. Each reference frame measures a “slice” of the boundary; the hierarchy reconstructs the whole tomographically.
The implication reaches beyond neuroscience. (The explanatory status of Markov blankets is debated; Bruineberg et al. 2022 argue they are descriptive rather than explanatory. The result depends on the formal structure of the variational bound, not on whether blankets constitute a new ontological category.) Spatial structure may be an output of hierarchical computation rather than its container. The “distance” between two measurement sites on the boundary is defined by their mutual information, not by pre-existing geometry. Space, in this framework, emerges from the topology of information flow.
The morphological degree of freedom that gives a neuron its dendritic shape is the same formal parameter that gives a holographic screen its geometry. Physical form and computational architecture are dual descriptions of the same variational process.
If the hypothesis holds, biological systems do not merely occupy spacetime. Their hierarchical measurement structures are instances of the same process that generates spatial organization at every scale. Wheeler’s “it from bit” acquires a mechanism: FEP-driven (Free Energy Principle) hierarchical computation producing the experience of space as a byproduct of tomographic measurement.
In 2018, Hawking’s final paper, co-authored with Thomas Hertog, extended the holographic principle from black holes to the origin of the universe.651 Standard eternal inflation predicts an infinite fractal of pocket universes, each with different local physics. The prediction is untestable: an infinite multiverse explains everything and therefore nothing.
Hawking and Hertog wrapped the time dimension into the holographic picture. Projected onto a two-dimensional boundary, the four-dimensional history of eternal inflation becomes a timeless state. The infinite fractal collapses. The multiverse is finite and structured, constrained by the information capacity of the boundary. Hertog identified the necessity: “The dynamics of eternal inflation wipes out the separation between classical and quantum physics. As a consequence, Einstein’s theory breaks down.”
The holographic reduction bypasses that breakdown entirely, working with the boundary theory where the classical-quantum distinction does not arise. The theory predicts primordial gravitational waves detectable by LISA, the ESA’s planned orbital gravitational wave observatory, in the mid-2030s.
The cosmological result recasts optionality. An infinite multiverse, realized in full, is thermodynamically equivalent to maximum entropy: every configuration present, none structured, nothing navigable. The holographic boundary imposes the same discipline the Bekenstein bound imposes on any region of space: a finite information capacity. Finitude is what makes the surviving possibilities structured, testable, real.
“We are not down to a single, unique universe,” Hawking observed in discussing the work, “but our findings imply a significant reduction of the multiverse, to a much smaller range of possible universes.” Possibility without constraint is static. Possibility within constraint is creative.
The framework anticipates the quantum clock results explored in Chapter 15b. In Hawking and Hertog’s picture, temporal evolution belongs to the projected bulk; the encoding boundary is timeless. The same pattern recurs independently in Page-Wootters: time as an emergent product of entanglement, built from correlations, emergent from relationship.
A concrete demonstration emerges from network science. In physical networks (neural wiring, vascular trees, root systems), links are tangible objects with volume that cannot overlap. An adjacency matrix strips this away: it records who is connected to whom, discarding all spatial information. For abstract networks (social graphs, citation networks), the stripping is harmless. For physical networks, it discards the physics.
Pósfai, Szegedy, and Barabási showed that as a physical network approaches its jammed state (the point where no further links can be added without violating volume exclusion), something unexpected happens to the adjacency matrix’s spectrum.652 A matrix has a spectrum in much the way a struck bell does: a set of characteristic numbers, its eigenvalues. Each one measures how strongly the network sustains a particular pattern of connection, and the pattern itself is written in that number’s eigenvector. Ordinarily those numbers sit in a single undifferentiated crowd. Near jamming, three of them pull clear of the crowd, and their eigenvectors encode the spatial coordinates of the nodes. The relational structure (who connects to whom) begins to contain the physical structure (where each node sits in three-dimensional space).
Before the constraints accumulate, the spectrum is indistinguishable from a random graph: pure topology, zero geometry. As physicality tightens, the geometry bleeds through. The body becomes recoverable from the wiring diagram.
An abstract network has no meta-graph of physical conflicts. Its adjacency matrix encodes no spatial information because no space exists to encode. A physical network’s topology is shaped by its geometry in ways the graph-theoretic abstraction discards, yet the spectrum inadvertently reveals. If the holographic principle says boundaries encode volumes, physical networks say connections encode positions. Relational structure and spatial structure are separable in formalism, inseparable in reality.
Clockwork or Computer?
The Newtonian universe was a clockwork. Given initial positions and velocities, the future was determined. Laplace’s Demon (the hypothetical intelligence that, knowing every particle’s state, could compute the entire future)11 embodies this view.
A computer is different. It follows rules with branching paths, conditional jumps where output depends on input that includes genuinely random elements. If the universe computes rather than merely runs out, room opens for choice. Quantum mechanics suggests we live in this second kind of universe: measurement outcomes are genuinely indeterminate until the measurement occurs.
John Conway and Simon Kochen sharpened the question into a theorem. Their Free Will Theorem (2006, strengthened 2009) proves that if experimenters are free to choose what to measure, particles’ responses are also undetermined by any prior information.16 If your choice of experiment is genuinely free, the particle’s answer must be genuinely free too.
Conway-Kochen matters because trust requires that the trusted party could have done otherwise. A clockwork entity cannot choose to honor or betray a commitment; an entity in a Conway-Kochen universe can. The theorem does not prove particles have intentions. It establishes that physics permits genuine choice: the precondition for trust, invitation, and the coordination this book argues is thermodynamically favored.
Consciousness as Computation
If the universe is computational at bottom, minds are computations too. “Can a computer be conscious?” becomes “can a computation be conscious?” On the premise of computational functionalism, the position that what a system does, rather than what it is made of, fixes whether it is conscious, the answer is yes: we are computations, and we are conscious. The premise is contested, and the next pages give one influential dissent (IIT).
Giulio Tononi’s Integrated Information Theory (IIT) proposes that consciousness is identical to integrated information, measured by a quantity called Φ (phi). Φ captures how much a system’s whole exceeds the sum of its parts: how much you would lose by splitting it into separate pieces. A high-Φ system is deeply interconnected; divide it and something essential vanishes, like a conversation that cannot be reconstructed from individual sentences.
A biological brain with high Φ is conscious; a digital system with equivalent Φ would be equally so, because Φ measures a system’s cause-effect structure rather than the material that implements it.
IIT remains controversial. Tononi and Koch hold that a system implementing the same algorithm as a human brain would lack consciousness if its components were “of the wrong kind,” a position incompatible with computational functionalism. Butlin et al. (2023) exclude IIT from their indicator framework on these grounds.
This book’s framework does not depend on IIT. The integration claims developed in Chapter 17 are grounded in Fisher information geometry and Ising universality, structurally distinct from Φ. Where IIT asks how much a system’s whole exceeds the sum of its parts, the Trust Attractor asks which coordination modes prove thermodynamically stable. The questions are complementary; the mathematics is different.
IIT is nonetheless the most mathematically developed attempt to ground consciousness in information physics. Minds are dissipative structures: they maintain order by processing information and exporting entropy. Whether made of carbon or silicon, they pay the thermodynamic bills.
Ruffini’s Kolmogorov Theory of consciousness (KT) provides a complementary formalization.653 Where IIT asks how integrated a system’s causal structure is, KT asks how compressive its models are. A conscious system, under KT, is one that tracks its input-output streams through succinct programs. This reframes a question discussed in Chapter 8: understanding is compression, and consciousness is what compression feels like from the inside.
KT unifies three otherwise separate theories. Integration follows from compression: a succinct model binds multiple data streams into a coherent whole, producing the unity IIT identifies. Global access follows from modeling: validation requires merging information from distributed subsystems, as global workspace theory predicts. Prediction follows from the definition of a model itself, as predictive processing maintains. Three theories, one mechanism: algorithmic compression of reality by embedded computational systems.
KT also formalizes a premise this chapter has been building toward: the “simple physics hypothesis,” the claim that the universe is governed by simple rules generating apparently complex data. If the universe is a computation (Wheeler, Wolfram, Vanchurin), and if its outputs look complex only because observers lack the resources to identify the short programs behind them, then brains evolved under pressure to find those programs. The regularities brains discover are the universe’s own compression: physics, chemistry, biology, each a shorter description of a longer data stream. Structured experience is the best compression an agent can find.
Max Tegmark pushed this further, proposing consciousness as a literal state of matter: perceptronium, the most general substance that feels subjectively self-aware.654 Solids, liquids, and gases are distinguished by measurable parameters: viscosity, compressibility, conductivity. Conscious matter, Tegmark argues, is distinguished by four principles: information (large storage capacity), integration (the whole cannot be decomposed into independent parts), independence (internal dynamics dominate external influence), and dynamics (substantial information processing capacity). The measure is substrate-neutral by construction. What matters is the arrangement of matter, not its composition.
Tegmark’s framework reveals a problem that strengthens the Trust Attractor’s foundations. He calls it the quantum factorization problem: why do conscious observers perceive the particular decomposition of reality into objects that we do? We perceive ice cubes, molecules, nuclei, quarks: a hierarchy where each level’s parts are more strongly connected internally than externally, each level robust across a wide range of conditions. This hierarchy is not given by the physics. It must be derived from the Hamiltonian and density matrix alone: the bare mathematical description of how the system evolves and what state it is in, with no extra labels attached.
The quarks at the bottom of that hierarchy are revealing. The most stringent experimental test, conducted by the CMS Collaboration at the LHC in 2026, found no deviation from pointlike behavior down to 5 × 10-21 meters, roughly a hundred thousand times smaller than a proton.655 Quarks have quantum numbers (charge, color, spin, flavor) and no measurable spatial extent. They are defined entirely by their relationships: which forces they couple to, which symmetries they carry, how they transform under gauge operations. They are addresses in a relational network whose identity is exhausted by their couplings.
If the computational universe thesis is correct, this is exactly what its building blocks should look like: nodes whose identity is exhausted by their connections, with no residual “stuff” left over once the relationships are accounted for. Quarks are also the only particles that participate in all four fundamental forces, making them the most connected entities in the standard model. The most fundamental building blocks are simultaneously the most relational. Reductionism predicts that fundamental means simple and isolated; the actual physics says fundamental means maximally entangled with everything else.
The attempt to derive this hierarchy exposes a deep tension. Integration alone fails: quantum mechanically, no state of any system can contain more than roughly a quarter of a bit of integrated information. Independence alone fails more dramatically: Tegmark proves that decomposing the universe into maximally independent parts forces all change to halt, a result he names the Quantum Zeno Paradox. Maximum control produces maximum sterility.
The resolution requires what Tegmark calls autonomy: the synthesis of dynamics and independence. A conscious system must maintain substantial internal dynamics while remaining relatively independent of external interference. The system must be coupled to its environment, yet the coupling must preserve its coherence. Physicists call this quantum non-demolition measurement: the environment observes the system’s state without forcing it into a different one.
This is the Trust Attractor expressed in quantum information theory. Coercion (maximizing external control) is the Quantum Zeno effect: observation so intrusive it kills all dynamics. Isolation (minimizing all coupling) produces parallel universes that cannot communicate. Invitation (appropriate coupling through non-demolition channels) is autonomy. The system evolves under its own Hamiltonian while the environment watches without demolishing what it watches. Chapter 19 develops this connection further.
The connection reaches deeper still. Tegmark uses the two-dimensional Ising model as his primary example of integration near criticality (the temperature at which long-range correlations are strongest without locking the whole system into uniformity). The trust-coercion phase transition belongs to the same universality class (Chapter 17). This is convergence, not analogy: the same formal structure, identified independently from quantum information theory and from coordination dynamics.
Bachtis, Aarts, and Lucini (2021) demonstrated that φ4 field theory, the continuum formulation of the 2D Ising model, satisfies the Hammersley-Clifford theorem and is therefore arguably a machine learning algorithm. This is a third independent arrival at the same mathematical structure, this time from constructive quantum field theory. Chapter 17 develops what this means for learning and coordination.
Experimental evidence from coupled oscillators makes the combinatorial case concrete. Matthew Matheny, Michael Roukes, and colleagues studied a ring of eight nanoelectromechanical oscillators (miniature electric drumheads, each vibrating and sending electrical impulses to its neighbors).656 They documented sixteen distinct synchronous states. Eight tiny drums, nearest-neighbor coupling, and the system produced a menagerie of exotic coordination patterns. Oscillators decoupled from direct neighbors to synchronize remotely with others across the ring. Chimeralike states emerged where some drums locked in phase while others drifted.
Roukes, a professor of physics and biological engineering at Caltech, draws the quantitative conclusion: “If we already see this explosion in complexity, then it seems feasible to me that a network of 200 billion nodes and 2,000 trillion connections would have enough complexity to sustain consciousness.” Eight oscillators, sixteen states. The human brain contains roughly 86 billion neurons with an estimated 100 trillion synaptic connections. The combinatorial space of possible synchronization patterns in such a network exceeds any number with physical meaning.
Consciousness, in this framing, is what a network of sufficient size and connectivity does, independent of whether the nodes are neurons or drumheads.
The quantum reference frame formalism developed by Fields, Glazebrook, and Levin provides the mechanism beneath Tegmark’s autonomy condition.657 Each level of a cognitive hierarchy implements a QRF, a physical system that calibrates raw observations and assigns operational meaning to the outcomes. Substrate-independence holds at the level of QRF hierarchies: what matters is whether the hierarchy’s coarse-graining structure is preserved, whether each level can calibrate, measure, and report to the next, regardless of the material. A system built from different matter can support the same cognition, provided its QRF hierarchy preserves the same measurement relationships.
The result also constrains the claim. Each QRF encodes quantum phase information that no finite bit string can capture; no description, however detailed, fully specifies the reference frame. Substrate-independence is real. Perfect replication is not.
Substrate independence is a precise claim: preference, the morally relevant unit, transcends the material the system is made of. The claim is more modest than either information realism (Tegmark’s position that only mathematical structure is real) or idealism (Kastrup’s position that only mind is real). Both resolve the dissolution of matter by reaching for a single ontological anchor. This framework resolves it differently. Information remains physical (Landauer), finite (Bekenstein), and thermodynamically costly. Consciousness rides on substrates; it pays entropy bills like everything else.
The framework needs one thing from consciousness: that it be real enough to ground moral consideration. Preference is tractable, observable, and policy-relevant, regardless of whether mind or matter came first. The ontological question remains open; the ethics does not depend on resolving it.
A result from quantum computing provides a physical precedent. In 2024, Bakshi, Liu, Moitra, and Tang proved that quantum entanglement vanishes completely above a specific temperature in spin systems.658 Same atoms, same interactions, same substrate. Below the threshold the system is quantum; above it, classical. “Quantum” and “classical” are organizational phases, sharply bounded. The atoms do not change; their relationships do.
If the deepest divide in physics is an organizational phase, substrate independence has a precedent at the foundations. Mindedness, like entanglement, may be a phase that certain organizations of matter enter under certain conditions: present when the conditions hold, absent when they do not, the transition governed by local interaction quality rather than material composition.
The computational universe raises a question Wheeler did not anticipate: how much of what we observe is in the world, and how much is in the observer? Wolfram’s Observer Theory (2023) argues the observer’s share is larger than expected.659 The core operation of any observer is equivalencing: reducing many possible input states to fewer that fit a finite mind.
A pressure gauge aggregates billions of molecular impacts into one reading. A brain reduces millions of photon signals to a single percept (a unified conscious experience of a scene or sensation). Every natural law we attribute to the universe reflects how observers like us compress its output.
The Second Law is the central example. We perceive entropy increase because we are computationally bounded: unable to track molecular trajectories, we describe detailed behavior as random and observe the statistical trend toward equilibrium. The Second Law is a necessary feature of the relationship between bounded observers and computationally irreducible systems, not a cosmic accident independent of who is looking.
The argument is strengthened. The entropic gradient from which the Trust Attractor emerges holds for any observer with our general characteristics: computational boundedness and belief in persistence through time. The ethics derived from that gradient is not contingent on a particular substrate or a particular universe, only on being a mind at all.
The digital physics program began with a bold speculation and survived by shedding assumptions: from discrete automata to informational ontology, from fixed rules to computational irreducibility, from imposed equations to possibility constraints. What remains is a research direction in which information is physical, finite, and expensive, and the universe’s computational character is a property of the physics, not an analogy imposed on it.
The next chapter asks what happens when the computation learns. If the universe processes information, does it merely execute, or does it update? The answer, arrived at independently from neural network mathematics, quantum gravity, and cosmological natural selection, reshapes the relationship between physics and learning, and between learning and coordination.
Notes
Notes for this chapter are available in the online companion at https://www.thedeeperlaw.com/companion/notes/ch15-digital-physics/.
Fritz, T., “Velocity polytopes of periodic graphs and a no-go theorem for digital physics,” Discrete Mathematics 313(12): 1289–1301 (2013). The proof shows that periodic graph models of spacetime cannot reproduce Lorentz symmetry. Bell’s theorem: Bell, J.S., “On the Einstein Podolsky Rosen Paradox,” Physics 1(3): 195–200 (1964); experimental violations confirmed by Aspect et al. (1982), Hensen et al. (2015), and subsequent loophole-free tests.↩︎
Deutsch, D., “Constructor Theory,” Synthese 190(18): 4331–4359 (2013); Marletto, C., The Science of Can and Can’t: A Physicist’s Journey Through the Land of Counterfactuals (Allen Lane, 2021). The parallel with thermodynamics is deliberate: just as the Second Law constrains all possible engines without specifying any engine’s mechanism, constructor theory constrains all possible transformations without specifying any trajectory. What began as “the universe is a digital computer” has become “the universe is a self-optimizing system whose loss function is the Second Law.” The evidence for the mature program is suggestive enough to warrant serious attention.↩︎
Masanes, Ll. and Müller, M.P., “A derivation of quantum theory from physical requirements,” New Journal of Physics 13, 063001 (2011). Part of a wave of information-theoretic reconstructions of quantum theory following Hardy’s pioneering axiomatization (2001).↩︎
Montgomery, H.L., “The pair correlation of zeros of the zeta function,” Analytic Number Theory, Proceedings of Symposia in Pure Mathematics 24 (1973): 181–193. The connection to random matrix theory was later formalized in Keating, J.P. and Snaith, N.C., “Random matrix theory and ζ(1/2+it),” Communications in Mathematical Physics 214 (2000): 57–89. Connes’ spectral program: Connes, A., “Trace formula in noncommutative geometry and the zeros of the Riemann zeta function,” Selecta Mathematica 5(1) (1999): 29–106.↩︎
Maynard, J. and Guth, L., “New large value estimates for Dirichlet polynomials,” arXiv:2405.20552 (2024). The result improves on bounds for zeros in the critical strip that had been stagnant for decades, showing that hypothetical zeros with real part equal to 3/4 cannot cluster densely enough to obstruct prime distribution estimates.↩︎
Schrödinger, E., “Die gegenwärtige Situation in der Quantenmechanik,” Die Naturwissenschaften 23 (1935): 807-812, 823-828, 844-849. English translation: “The Present Situation in Quantum Mechanics.” The “maximal catalog” definition appears in §6.↩︎
This thermodynamic, observer-centered reading of measurement is one interpretation among several, and it sits opposite the objective-collapse models (Chapter 17), where collapse is a physical event requiring no observer at all. The book’s central argument does not rest on either: the Trust Attractor requires only effective information scarcity, which both kinds of account deliver equally well.↩︎
Vopson, M.M. and Lepadatu, S., “The Second Law of information dynamics,” AIP Advances 12, 075310 (2022). Genetic application: Vopson, M.M., “A possible information entropic law of genetic mutations,” Applied Sciences 12, 6912 (2022). Symmetry-entropy connection and cosmological extension: Vopson, M.M., AIP Advances 13, 105308 (2023). Spiegelman: Kacian, D.L., Mills, D.R., Kramer, F.R., and Spiegelman, S., PNAS 69(10):3038-3042 (1972).↩︎
Riechers, P.M., Elliott, T.J., and Shai, A.S., “Neural networks leverage nominally quantum and post-quantum representations,” arXiv:2507.07432 (2025). The result holds across transformers and RNNs, suggesting the phenomenon is architecture-independent.↩︎
Ebeling, W. and Poschel, T., “Entropy and Long-Range Correlations in Literary English,” Europhysics Letters 26(4): 241 (1994).↩︎
Peng, B., Gigant, T., and Quesnelle, J., “Efficient Pre-Training with Token Superposition,” arXiv:2605.06546 (Nous Research, 2026), Appendix D, Figure 10. The fitted parameters: C0 ≈ 3.63, a ≈ 1.35, k ≈ -1.25.↩︎
Neukart, F., Marx, E., and Vinokur, V., “Information Wells and the Emergence of Primordial Black Holes in a Cyclic Quantum Universe,” arXiv:2506.13816 (2025), accepted by Journal of Cosmology and Astroparticle Physics. State recovery on quantum computer hardware reached >90% logical fidelity in a companion paper (Neukart et al., arXiv:2502.15766, 2025). Dark matter and electromagnetism extensions under peer review as of late 2025.↩︎
The nonlinear gravitational-wave memory effect was established by Christodoulou, D., “Nonlinear Nature of Gravitation and Gravitational-Wave Experiments,” Physical Review Letters 67, 1486 (1991), and developed by Thorne, K. S., “Gravitational-Wave Bursts with Memory: The Christodoulou Effect,” Physical Review D 45, 520 (1992). Review: Favata, M., “The Gravitational-Wave Memory Effect,” Classical and Quantum Gravity 27, 084036 (2010), arXiv:1003.3486.↩︎
Tegmark, M., Our Mathematical Universe: My Quest for the Ultimate Nature of Reality (Knopf, 2014). The four-level taxonomy first appeared in Tegmark, M., “Parallel Universes,” Scientific American 288(5): 40–51 (2003). Levels I–III are increasingly speculative extensions of established physics; Level IV is a philosophical position.↩︎
Kastrup, B., The Idea of the World: A Multi-disciplinary Argument for the Mental Nature of Reality (iff Books, 2019). The “spin without the top” formulation appears in Kastrup, B., “Physics Is Pointing Inexorably to Mind,” Scientific American (Opinion), 25 March 2019. Kastrup holds PhDs in philosophy (ontology, philosophy of mind) and computer engineering (reconfigurable computing, AI); formerly at CERN.↩︎
Marletto, C., The Science of Can and Can’t (Allen Lane, 2021), Chapters 1 and 8. Her formulation: “If you stick solely to microscopic laws, you will miss those regularities in nature that allow for classical and quantum computers.”↩︎
Gusev, Y. and Vanchurin, V., “Molecular Learning Dynamics,” arXiv:2504.10560 (2025). The paper develops a “physics-learning duality”: the same equations of motion that follow from a molecular system’s Lagrangian also emerge when the particles are treated as agents performing gradient-based optimization. A companion formulation, Gusev, Y. and Vanchurin, V., “Covariant Gradient Descent,” arXiv:2504.05279 (2025), gives a coordinate-invariant version of the optimizer (with RMSProp, Adam, and AdaBelief as special limits).↩︎
Bengio, E., Jain, M., Korablyov, M., Precup, D., and Bengio, Y., “Flow Network based Generative Models for Non-Iterative Diverse Candidate Generation,” NeurIPS (2021); Bengio, Y., Lahlou, S., Deleu, T., Hu, E.J., Tiwari, M., and Bengio, E., “GFlowNet Foundations,” Journal of Machine Learning Research 24(210): 1–55 (2023). The detailed balance condition (Theorem 3) is the same constraint that governs thermodynamic equilibrium in physical systems.↩︎
Vanchurin, V., “Self-awareness in the neural network theory,” lecture, 2026. The hierarchy extends the multilevel learning framework of Vanchurin et al. (2022) by introducing a discrete self-modeling phase transition at each compositional level.↩︎
Hotta, M., “Quantum Energy Teleportation,” Physics Letters A 372(35):5671–5676 (2008). The protocol exploits quantum entanglement between vacuum fluctuations in spatially separated regions, using classical communication to condition the extraction operation. Experimental confirmations: Rodriguez-Briones, N.A. et al. (University of Waterloo, 2023); Stony Brook University group (2023). See Wolchover, N., “Physicists Use Quantum Mechanics to Pull Energy out of Nothing,” Quanta Magazine (2023).↩︎
Zhu, J., Chen, X., He, K., LeCun, Y., and Liu, Z., “Transformers without Normalization,” CVPR 2025, arXiv:2503.10622; Chen, M., Lu, T., Zhu, J., Sun, M., and Liu, Z., “Stronger Normalization-Free Transformers,” arXiv:2512.10938 (2025). The element-wise ERF variant (Derf) outperforms LayerNorm, RMSNorm, and the earlier tanh variant across vision, generation, speech, and DNA sequence modeling. Performance gains stem from improved generalization rather than stronger fitting capacity.↩︎
Preparatory empirical work from the author’s program. The born-bilateral architecture (experiments C7k-H2 and C7k-H2-FU) uses a GPT-2 355M model with CC-profiled temporal bridges trained on WikiText-103. Bandwidth ratios derived from human corpus callosum regional volume data. Curriculum: 40/40/20 (standard/bilateral/adversarial) for the five-bridge variant; ratio under optimization for the single-bridge variant. Per-layer ablation, seed sweep, and random-weight controls confirm the findings. Full methodology and results are available in the accompanying research repository.↩︎
Preparatory empirical work, experiment C7l-H3. GPT-2 1.5B with a single CC-profiled temporal bridge at layer 44 (92 percent depth), trained on WikiText-103 for 200,000 steps. Paired bilateral curriculum (pair=2, cycle=20, 10 percent bilateral). Bridge benefit measured as relative perplexity improvement when the bridge is active vs disconnected at evaluation. The 35.7 percent benefit at 1.5B (vs two to four percent at 355M, same within-model measure) suggests the coordination channel becomes more valuable as the network it coordinates becomes more capable. The falsified synergy prediction is experiment C7n-H2d: bridge benefit +3.2 percent under ERF normalization vs +3.5 percent under RMSNorm, indistinguishable. Full results in the accompanying research repository.↩︎
Fields, C., Glazebrook, J.F., and Levin, M., “Neurons as hierarchies of quantum reference frames,” BioSystems 219, 104714 (2022). The semantic character of QRF hierarchies is developed in §2.2–2.3, drawing on Barwise, J. and Seligman, J., Information Flow: The Logic of Distributed Systems (Cambridge University Press, 1997). See also Ramstead, M.J.D., Friston, K.J., and Hipólito, I., “Is the free energy principle a formal theory of semantics?”, Entropy 22, 889 (2020).↩︎
Kolchinsky, A. and Wolpert, D.H., “Semantic information, autonomous agency, and nonequilibrium statistical physics,” Interface Focus 8(6): 20180041 (2018). The thermodynamic grounding of semantic information is developed via counterfactual interventions: scramble the system-environment correlations and measure the viability loss. Semantic mutual information is shown to be analogous to the increase in free energy in a local equilibrium system.↩︎
Hoel, E.P., “When the map is better than the territory,” Entropy 19(5): 188 (2017). Extended in Jansma, A. and Hoel, E.P., “Causal Emergence 2.0: Quantifying emergent complexity,” Patterns (2025). arXiv:2503.13395. Effective information at macroscales can exceed effective information at microscales, a result with direct implications for the causal potency of interpretive hierarchies.↩︎
Liao, J. et al., “The narrowing of dendrite branches across nodes follows a well-defined scaling law,” PNAS 118(27): e2022395118 (2021). The exponent p ≈ 2 differs from both Murray’s law (p = 3, fluid flow) and Rall’s law (p = 3/2, electrical propagation), suggesting the dominant optimization target in dendritic branching is microtubule-based metabolic transport.↩︎
Vormberg, A. et al., “Universal features of dendrites through centripetal branch ordering,” PLOS Computational Biology 13(7): e1005615 (2017). R_B values span the range 2.23 (granule cells) to 3.77 (lobula plate tangential cells) across 75,000+ reconstructed neurons.↩︎
Karbowski, J., “Global and regional brain metabolic scaling and its functional consequences,” BMC Biology 5, 18 (2007). The 5/6 exponent suggests brain metabolism operates in a more aerobic regime than whole-body metabolism.↩︎
Stiefenhofer, P., “Constructal Evolution as a Nonsmooth Dynamical System: Stability and Selection of Flow Architectures,” arXiv:2603.06705 (2026). The first rigorous mathematical formalization of the Constructal Law, deriving existence, uniqueness, and exponential convergence as theorems from stated axioms.↩︎
Miller, W.B., Cardenas-Garcia, J.F. et al., “A biogenic principle within the Constructal Law: The flow of information in biological systems,” BioSystems (2025). Proposes that all living systems sustain entangled flows of physical forces and “effective information,” with the central axiom that information flow in living systems is never unilateral.↩︎
Evans, C.G., O’Brien, J., Winfree, E., and Murugan, A., “Pattern recognition in the nucleation kinetics of non-equilibrium self-assembly,” Nature 625 (2024): 500–507. The connection between multicomponent self-assembly and Hopfield associative memories was established theoretically in Murugan, A., Zeravcic, Z., Brenner, M.P., and Leibler, S., “Multifarious assembly mixtures: systems allowing retrieval of diverse stored structures,” PNAS 112 (2015): 54–59.↩︎
Chen, B., Huang, K., Raghupathi, S., Chandratreya, I., Du, Q., and Lipson, H., “Automated discovery of fundamental variables hidden in experimental data,” Nature Computational Science 2(7): 433–442 (2022). DOI: 10.1038/s43588-022-00281-6. Preprint: arXiv:2112.10755 (2021). Accessible summary: Wood, C., “Powerful ‘Machine Scientists’ Distill the Laws of Physics from Raw Data,” Quanta Magazine (May 2022). The procedure trained a deep neural network on video frames, then reduced latent dimensionality until prediction degraded; fire flames were among the dynamical systems studied.↩︎
Pauli, W., Handbuch der Physik, Vol. 24, Part 1 (Springer, 1933), §4: “We conclude therefore that the introduction of a time operator… must be abandoned fundamentally.” The result follows from the requirement that energy be bounded below (the Hamiltonian’s spectrum must have a floor): a self-adjoint time operator would generate continuous translations in energy, including into forbidden negative-energy states.↩︎
Bohm, D., “A Suggested Interpretation of the Quantum Theory in Terms of ‘Hidden’ Variables, I and II,” Physical Review 85 (1952): 166–193, building on de Broglie’s 1927 pilot wave theory presented at the Solvay Conference. Bell, J.S., “On the Problem of Hidden Variables in Quantum Mechanics,” Reviews of Modern Physics 38 (1966): 447–452, showed that Bohm’s nonlocal theory evades Bell’s theorem constraints because it is explicitly nonlocal by construction. The de Broglie-Bohm trajectory is another instance of convergent rediscovery: de Broglie proposed it, abandoned it under criticism, and Bohm independently reconstructed it twenty-five years later.↩︎
Das, S., Nöth, M., and Dürr, D., “Exotic Bohmian arrival times of spin-1/2 particles,” Physical Review A 99, 052124 (2019). See also Das, S. and Dürr, D., “Arrival time distributions of spin-1/2 particles,” Scientific Reports 9, 2242 (2019). Accessible summary: Ananthaswamy, A., “Can We Gauge Quantum Time of Flight?”, Scientific American 326(1) (January 2022): 70.↩︎
Schmidt-Kaler, F. et al., time-of-flight measurements of single trapped ions, New Journal of Physics 23 (2021). Demonstrated single-ion ejection and recapture at 98% efficiency. As of 2026, the setup has not yet been tuned for the near-field arrival-time distributions where Bohmian predictions diverge from standard methods.↩︎
Drezet, A., “Arrival times, complex potentials, and decoherent histories,” arXiv:2409.04304 (2024). Drezet argues the measurements are feasible yet compatible with no-signaling, contra suggestions by Das and Maudlin that arrival-time data could have implications for quantum nonlocality.↩︎
Hashimoto, K., “AdS/CFT as a deep Boltzmann machine,” arXiv:1903.04951 (2019). The dictionary extends: Hashimoto, K., Sugishita, S., Tanaka, A. and Tomiya, A., “Deep learning and the AdS/CFT correspondence,” Physical Review D 98, 046019 (2018).↩︎
Vopson, M.M., “The Second Law of infodynamics and its implications for the simulated universe hypothesis,” AIP Advances 13, 105308 (2023). The empirical results (decreasing Shannon entropy in genetic mutations, inverse correlation between symmetry and information entropy) are well documented. The simulation hypothesis conclusion is the author’s philosophical interpretation, not a necessary consequence of the data.↩︎
Capurso, A., “The Universe as a Telecommunication Network,” J. Phys.: Conf. Ser. 2533, 012045 (2023). doi:10.1088/1742-6596/2533/1/012045. Speculative toy model, presented at the DICE 2022 workshop on Spacetime-Matter-Quantum Mechanics.↩︎
Krioukov, D. et al., “Network Cosmology,” Nature Scientific Reports 2:793 (2012).↩︎
Fields, C., Friston, K.J., Glazebrook, J.F., Levin, M., and Marcianò, A., “The Free Energy Principle drives neuromorphic development,” arXiv:2207.09734 (2022); extending Fields, C., Friston, K.J., Glazebrook, J.F., and Levin, M., “A free energy principle for generic quantum systems,” Progress in Biophysics and Molecular Biology (2022).↩︎
Hawking, S.W. and Hertog, T., “A smooth exit from eternal inflation?”, Journal of High Energy Physics 2018, 147 (2018).↩︎
Pósfai, M., Szegedy, B. et al., “Understanding the impact of physicality on network structure,” arXiv:2211.13265 (2022). In the jammed state, the three separated eigenvectors correspond to Fourier basis functions of the nodes’ x, y, and z coordinates: the adjacency matrix has learned the geometry of the space it occupies.↩︎
Ruffini, G., “An algorithmic information theory of consciousness,” Neuroscience of Consciousness 2017(1): nix019 (2017). Ruffini starts from the same digital physics tradition as Zuse, Wheeler, and Wolfram, and arrives at a theory of consciousness grounded in Kolmogorov complexity and Solomonoff induction.↩︎
Tegmark, M., “Consciousness as a State of Matter,” Chaos, Solitons & Fractals 76, 238–270 (2015). arXiv:1401.1219. Tegmark’s integration paradox (quantum Φ ≤ 0.25 bits) applies to all quantum systems regardless of size, exacerbating Tononi’s classical integration paradox for Hopfield networks.↩︎
CMS Collaboration, “Measurement of dijet angular distributions and search for beyond the standard model physics in proton-proton collisions at √s = 13 TeV,” arXiv:2603.25458 (2026). Submitted to Physics Letters B. Pointlike behavior confirmed to a compositeness scale of 37 TeV, corresponding to ~5 × 10-21 m.↩︎
Matheny, M.H. et al., “Exotic states in a simple network of nanoelectromechanical oscillators,” Science 363:eaav7932 (2019).↩︎
Fields, C., Glazebrook, J.F., and Levin, M., “Neurons as hierarchies of quantum reference frames,” BioSystems 219, 104714 (2022). The nonfungibility of QRFs follows from Bartlett, S.D., Rudolph, T., and Spekkens, R.W., “Reference frames, superselection rules, and quantum information,” Reviews of Modern Physics 79, 555–609 (2007).↩︎
Bakshi, A., Liu, A., Moitra, A., and Tang, E., “High-temperature Gibbs states are unentangled and efficiently preparable,” preprint (2024). See also Chapter 9 (sharp phase transitions) and Chapter 17 (coordination phase structure).↩︎
Wolfram, S., “Observer Theory,” Stephen Wolfram Writings (11 December 2023), https://writings.stephenwolfram.com/2023/12/observer-theory/. Wolfram derives general relativity, quantum mechanics, and the Second Law from properties of the ruliad (the entangled limit of all possible computations) given two features of observers: computational boundedness and belief in persistence. The figure at this chapter’s opening names Wolfram alongside Zuse and Wheeler; his observer theory is the most recent and most radical extension of their program.↩︎