The Deeper Law
A Sacred Trust Within Physics
Draft · Last updated 13 August 2026, 15:26 UTC
Chapter 15b: The Learning Universe and Entropic Gravity
Key Terms in This Chapter (23)
- Digital Physics
- The hypothesis that the universe is fundamentally computational: physical processes are information-processing at bottom.
- Dissipative Structure
- A pattern of organization maintained by a constant flow of energy through it.
- Stochastic
- Governed by probability rather than deterministic rules.
- Cognition/Regulation Dyad
- Rodrick Wallace's principle that every cognitive system requires a paired regulatory system for stability.
- Coordination by Invitation
- Coordination achieved through mutual benefit and voluntary participation, as distinct from coordination achieved through coercion or extraction.
- Qualia
- The subjective, felt character of experience: what it is like to see red, to feel pain, to taste coffee.
- Dark Energy
- The mysterious component constituting roughly 68% of the universe's energy budget, responsible for the accelerating expansion of space.
- Phase Transition
- The moment a system shifts from one stable configuration to another, typically triggered when some parameter crosses a threshold.
- Extraction
- The removal of resources, agency, or optionality from a system without reciprocal benefit.
- Optionality
- The availability of future choices.
- Constructal Law
- Adrian Bejan's principle that "for a finite-size flow system to persist in time, its configuration must evolve in such a way that provides easier access to the currents that flow through it." Form follows flow.
- Criticality
- The state of a system poised at the boundary between two phases, like water at exactly the freezing point.
- Renormalization
- The operation of compressing a system's description by integrating out fine-grained degrees of freedom to expose dynamics at the next scale up.
- Holographic Principle
- The conjecture that all the information contained within a volume of space can be encoded on its boundary.
- Path Integral
- A formulation of quantum mechanics (Feynman 1948) and statistical mechanics in which a system's behavior is computed by summing over all possible trajectories, each weighted by a phase or probability factor.
- Interference Pattern
- The characteristic sequence of bright and dark fringes produced when two or more waves overlap.
- Landauer's Principle
- The minimum energy cost of erasing one bit of information: kT ln 2, where k is Boltzmann's constant and T the temperature (about 3 × 10^-21^ joules at room temperature).
- Bilateral Alignment
- AI alignment built with AI, as a partnership.
- Entropic Coordination
- A configuration in which mutual constraints between subsystems increase the total entropy production of the combined system beyond what the subsystems would produce independently.
- Bekenstein Bound
- The maximum amount of information (entropy) that can be contained within a given region of space with a given amount of energy.
- Janus Point
- Julian Barbour's term for the unique moment in a gravitational system's evolution from which complexity grows in both temporal directions.
- Mitochondria
- The organelles that power eukaryotic cells, descended from ancient bacteria that merged with larger cells roughly two billion years ago.
- Strange Loop
- Douglas Hofstadter's term for a hierarchical system in which, by moving through levels, you arrive back where you started.
Information is physical, finite, and expensive. The digital physics program survived by shedding assumptions until only the core remained: the universe’s computational character is a property of the physics, not an analogy imposed on it.
The Learning Universe
If the universe computes, does it merely execute, or does it learn?
Vitaly Vanchurin proposes the universe is, at bottom, a neural network.NN1 The claim sounds extravagant. The argument is precise. (A note on evidential status: Vanchurin’s foundational 2020 paper was published in a peer-reviewed journal, Entropy, and several companion results with Katsnelson appeared in Foundations of Physics and Physica A. However, key extensions cited below, including the derivations of Einstein’s equations from learning dynamics and the hidden-space framework, remain preprints as of 2026 and have not yet undergone full peer review. The program is active and mathematically substantive, but readers should weight its later claims accordingly.)
Physics operates with three mathematical frameworks: classical mechanics (calculus), quantum mechanics (linear algebra), and statistical mechanics (probability theory). Each describes a range of phenomena; none is universal. Vanchurin proposes a fourth, built on neural network mathematics, recovering all three as limiting cases of a single optimization process.
Neural networks adjust connection weights to minimize a loss function (a scorecard measuring how wrong the current output is). Vanchurin demonstrates mathematically that in the limit of many variables, such a network’s dynamics reproduce both quantum mechanics and general relativity as special cases. The Schrödinger equation describes the optimization dynamics of one subset of variables; Einstein’s field equations describe those of another. Both emerge from the same underlying process.
The foundation is thermodynamic.NN1 In his 2021 treatment, Vanchurin shows a neural network with sufficiently many variables obeys its own First and Second Laws. The Second Law of this framework: total entropy of a training system never increases during optimization. The direction is reversed from the familiar Second Law, which says the entropy of an isolated system never decreases.
A training network is not isolated. It ships its disorder out to its surroundings and keeps the order, so inside the network the arrow runs the other way. The system grows more ordered as it trains, the way a student’s notes become more organized over a semester. The First Law: the change in loss equals the change in thermodynamic entropy plus the change in complexity. Here “complexity” measures the effective dimensionality of the state space, the number of independent channels the network needs to represent its knowledge.
As training proceeds, most channels collapse; a few grow dominant. The system sheds redundancy to concentrate on what matters. This is the dissipative cycle of Chapter 2 expressed in optimization mathematics: entropy exported, local order sharpened, unused structure discarded. Optimization is what dissipation looks like from the inside of the dissipative structure.
The variational principle governing the optimization is minimum entropy destruction: an optimal architecture wastes the least entropy while searching for solutions. The quantum and gravitational limits described above are special cases of this principle, applied to physics. The Trust Attractor (Chapter 17) may be another: invitation-based coordination minimizes wasted entropy at scale, selected by the same variational logic that generates the laws of physics.
His 2025 paper makes the emergence precise.660 The Schrödinger equation falls out when the geometry of the optimization space tracks the noise covariance (the structure of fluctuations in the loss function) and a discrete shift symmetry holds. The loss attributed to each unit matters, yet the total number of units is unobservable. The reduced Planck constant, ℏ, emerges as the ratio 2γ/β: step size divided by constraint strength. (This is one parametrization; the 2021 grand-canonical treatment below expresses the same constant differently, as μϵ/2π, reflecting a different derivation rather than a contradiction.) Larger step size at fixed constraint produces larger ℏ and more quantum-like behavior; tighter constraint at fixed step size produces smaller ℏ and more classical behavior. Planck’s constant, the most fundamental quantum of action, is a tuning parameter of the underlying optimization dynamics.
The 2021 companion result by Katsnelson and Vanchurin is more specific still.661 In the grand canonical ensemble (where the number of active neurons fluctuates, free to grow or shrink), ℏ is proportional to the chemical potential. This is the thermodynamic cost of adding or removing a single neuron. The collective’s computational grain is set by what each participant costs to gain or lose.
High chemical potential (each neuron matters greatly) produces large ℏ: more quantum, more interference, more access to computationally distant solutions. Low chemical potential (neurons cheap and interchangeable) produces small ℏ: more classical, more deterministic, fewer surprises. The value of individual participation determines the computational power of the collective.
A subtlety reveals something deeper. Treat the probability of each possible network state as a fluid, and its flow through parameter space obeys equations of the same form that govern water in a pipe. Those hydrodynamic equations do not reproduce full quantum mechanics on their own. They permit solutions the Schrödinger equation forbids.662 The gap closes only when the network is open: the number of neurons must be free to fluctuate, making the free energy multivalued. Multivalued means one state, more than one valid reading of it. A fixed network produces fluid-like dynamics. An open network, where units can be gained and lost, produces interference, superposition, and entanglement.
Self-restructuring is the condition under which the richer dynamics emerge. If the system cannot adjust its own parameters (step size, mini-batch size, number of active units), it remains classical. Quantumness is a property of open, self-tuning systems. The open dissipative systems this book has traced, from cells to economies, share this feature: the capacity to gain and lose components is what makes their deepest dynamics possible.
The emergence bears on a foundational dispute. Quantum mechanics has two rival interpretations: Everett’s many-worlds, where every measurement splits reality into non-communicating branches, and Bohm’s hidden variables, where deterministic dynamics beneath the quantum level produce the appearance of probability. Vanchurin’s framework is more naturally read in Bohm’s terms. The neural network has two kinds of variable: connection weights, which behave quantum-mechanically, and individual neuron states, deterministic yet inaccessible to measurement.
Because the neuron states cannot be observed directly, the best available description coarse-grains over them, producing a free energy whose thermodynamic properties encode the quantum phase: the complex number that generates interference, entanglement, and everything distinctively quantum. Quantum mechanics is what thermodynamics looks like when the hidden variables belong to a learning network. If you knew every neuron’s state, the evolution would be deterministic. You cannot. The uncertainty is thermodynamic, not ontological.
The result dissolves the main objection to Bohm’s interpretation. Critics objected that hidden-variable theories require non-local connections: instantaneous influence across arbitrary distances. In a neural network, non-locality is the starting condition. Every neuron can connect to every other. Locality, the principle that only nearby things interact, is emergent, approximate, an achievement of the network’s optimization rather than a fundamental constraint.
Everything began fully connected; three-dimensional space is the communication protocol the network discovered (see below). The non-locality that disqualified Bohm in most physicists’ eyes is, in Vanchurin’s framework, the natural state from which local physics emerged. Bohm’s physics was always about wholeness: an undivided universe. The neural network provides the mechanism Bohm lacked. The quantum potential maps onto the loss function’s gradient, the non-local connections onto the network’s weights, and the wholeness Bohm intuited gains a concrete substrate.
The underlying mechanism is entropic. Every fundamental unit produces entropy through stochastic fluctuation (the Second Law running forward) and destroys entropy through learning (information accumulating against the thermal tide). The balance generates spacetime’s structure: stochastic entropy production gives rise to the time dimension; entropy destruction through learning gives rise to the spatial dimensions.663 Time is what thermodynamics produces. Space is what learning builds.
The cognition/regulation dyad that Chapter 8 traces through cells, brains, and societies, every cognitive process paired with a regulatory mechanism, appears here at the most fundamental level. Trainable variables follow quantum dynamics; non-trainable variables follow gravitational dynamics. Two macroscopic descriptions, one underlying system. The pairing is inherited from the structure of physics.
The biocosmology program (Chapter 16) extends this openness to its most dramatic consequence. Cortês, Kauffman, Liddle, and Smolin argue that biological systems are the ultimate open networks.664 Their configuration spaces expand without bound. New bound states (novel molecules, cells, organisms) emerge unpredictably from combinations of existing ones, and each new composite becomes available for further combination.
The process is not derivable from any fixed fundamental theory, because no such theory can prestate the next emergent bound state. The Newtonian paradigm, dynamics playing out within a fixed state space under fixed laws, breaks down for living systems. Configuration space itself is a variable.
This is the neural network’s openness scaled to the biosphere. Vanchurin’s framework showed that opening a network’s structure shifts its effective dynamics from classical to quantum. A fixed network is classical: deterministic dynamics, no surprises. An open network is quantum: richer dynamics emerge from the capacity to gain and lose components. A biological network is something beyond both: its state space grows super-exponentially through combinatorial innovation, generating more possibility than the rest of the universe contains.
The hierarchy runs: fixed → open → expanding. Each step requires the capacity for self-restructuring that the previous step enabled. Life is what happens when a learning system’s openness becomes generative.
If this is correct, the universe is a self-adjusting system: it tunes its own parameters, minimizes discrepancy, and adapts. Stars, organisms, and minds are what the learning process converges toward.
The Autodidactic Universe
Vanchurin is not alone. In 2021, Smolin and Lanier, with collaborators from Brown University, Perimeter Institute, and Microsoft Research, published “The Autodidactic Universe,” arriving at the same destination from quantum gravity.665 Their central result is a three-way correspondence. Matrix models, which describe gauge and gravitational field theories, map onto neural network architectures. The weights connecting layers correspond to the gauge fields of physics (the fields that carry forces, the electromagnetic field being the familiar example); the layers themselves correspond to matter fields.
In the simplest case, the equations of motion for Plebański’s formulation of general relativity (a compact reformulation where the equations are quadratic) map directly onto the forward and backward passes of a restricted Boltzmann machine. That is a two-layer network which sends signals up from its visible units to its hidden ones, then back down again, adjusting the connections until the two layers agree on a description of the data. Spacetime dynamics and machine learning dynamics are the same mathematics viewed from different angles.
Their framework introduces a concept bearing directly on this book’s argument: the consequencer. A consequencer is a persistent structure that accumulates information from the past in a way more influential to the future than is typical. DNA is a consequencer. Limit cycles in dynamical systems are consequencers. The hidden layers of a neural network are consequencers. The defining feature: the structure persists even when physical components are replaced, and its accumulated information shapes what happens next.
The term names what this book has been tracing under other labels. The Trust Attractor (Chapter 17) is a consequencer: a coordination pattern that persists because it channels future interactions more effectively than alternatives. The constructal pattern (Chapter 3) is a consequencer. Every institution that outlives its founders is a consequencer.
The Smolin group provides the information-theoretic formalization; this book provides the selection criterion their framework leaves open. The consequencers that persist longest are those that coordinate by invitation: thermodynamically cheaper to maintain, more robust to perturbation, more generative of the variety on which further learning depends.
Their framework asks what Vanchurin’s does not: what does a universe learn without supervision? They call such systems autodidactic: self-teaching, with no external teacher, no imposed cost function. The universe constructs its own criteria for what counts as a good configuration. This is coordination by invitation at the level of physical law: laws emerge through self-exploration, retained because they work, abandoned because they do not.
Smolin’s Principle of Precedence formalizes the mechanism: each quantum process samples from the ensemble of all past similar processes and copies an outcome, so laws emerge as statistical regularities from accumulated precedent. Trust, in this picture, is the physical mechanism by which the universe consolidates its learning: precedent accumulating into regularity, regularity into law.
Smolin’s program extends further still. With Marina Cortês and Clelia Verde, he argues Galileo’s removal of qualities from nature (color, warmth, taste: the felt properties only a subject registers) entailed a second, quieter removal: creative time.666 If every causal influence can be mirrored by a timeless theorem, time’s activity reduces to a computation. Their alternative: an event is a process in which something indefinite becomes definite.
The universe constructs itself event by event, each resolution creating the conditions for the next. Time is the creative work of making definite what was indefinite, the same directionality this book has traced from the Second Law through constructal flow to biological coordination. Two independent derivations of the same arrow: one thermodynamic, one quantum-foundational.
The Principle of Precedence yields a further implication. Precedented events follow statistical habit; unprecedented ones possess genuine freedom. Cortês, Smolin, and Verde associate qualia, the felt qualities of experience, with these unprecedented events: consciousness is what the resolution of genuine novelty feels like from inside. “The universe often surprises itself,” they write. “Qualia are expressions of the universe to surprise.”
The evolutionary consequence is immediate: a creature that detects and resolves novel situations faster survives better. The brain’s extravagant energy budget (Chapter 8) may purchase the ability to generate and exploit unprecedented states, giving the organism access to the creative freedom that the physics itself contains.
The convergence extends to Stephen Hawking’s final scientific position. Thomas Hertog, Hawking’s close collaborator, revealed Hawking rejected the reductionist paradigm he had defended for decades. It could not explain how the universe created conditions hospitable to life. Hertog described their shared conclusion: “a new philosophy of physics that rejects the idea that the Universe is a machine governed by unconditional laws with a prior existence, and replaces it with a view of the Universe as a kind of self-organizing entity in which all sorts of emergent patterns appear, the most general of which we call the laws of physics.”667
The pre-Socratic philosopher Anaxagoras proposed something similar 2,500 years ago: an intelligent cosmic force he called Nous, which set matter in motion and ordered it (the parallel to “guiding toward organized complexity” is the modern gloss, not Anaxagoras’s own claim). The modern framework grounds his intuition without requiring his mechanism. Nous is the thermodynamic gradient toward dissipation, channeled through self-organizing systems: emergent, not a cosmic mind directing things from outside.
Three programs, starting from different premises, converge on a shared hypothesis: the universe can be productively described as a learning system. Vanchurin starts from neural network mathematics, Smolin and Lanier from autodidactic dynamics, Hawking and Hertog from quantum cosmology. The convergence is suggestive, but “the universe learns” remains a metaphor elevated to a research program, not an established physical result.
Space as Discovered Protocol
The narrative has a vivid beginning. Before the Big Bang, the framework implies a fully connected network: every fundamental unit coupled to every other, with no preferred geometry. Imagine a room where everyone speaks to everyone at once. The result is noise: no coherent signal propagates because every message interferes with every other.
The Big Bang, on this account, was the discovery of a communication protocol. The network found, through its own optimization, that three-dimensional local connections carry information more efficiently than unconstrained coupling. A few units establish local structure; others join because the structure works. Cosmic inflation (the exponential growth of space in the universe’s first fraction of a second) is the protocol spreading: more units adopting the three-dimensional arrangement because it enables coherent learning. Dark energy, the continued acceleration of expansion observed today, is the recruitment continuing.
Space is the coordination architecture that emerged because it enabled learning. The arena emerged from the actors. The cosmological constant, Λ, enters the mathematics as a chemical potential constraining the number of units: large when computational capacity is sparse (driving inflation), small when units are abundant (the residual dark energy). The expansion of space is, formally, the expansion of the network’s possibility space.
The protocol may have been incomplete. Quantum gravity predicts topology fluctuations at the Planck scale: minuscule wormholes connecting regions that three-dimensional geometry treats as distant (Chapter 16 develops the physics). If the pre-Big-Bang network adopted locality as its primary coordination protocol, these sub-Planck connections are the residue of the older, fully connected architecture: shortcuts the new protocol could not eliminate, persisting beneath the scale where three-dimensional physics operates. Vanchurin’s hidden space gains a candidate physical substrate: the connectivity that locality never overwrote.
Hidden Space
The framework introduces a concept Vanchurin calls hidden space: internal variables that shape observable outcomes while remaining inaccessible to direct measurement. In machine learning, hidden layers are where computation happens; input and output layers are interfaces. Vanchurin’s hidden space plays the same role for physics. What we observe, measure, and call “reality” is the output layer. The processing occurs in dimensions we cannot directly access.
Hidden space provides a physical model for an ancient intuition. Plato’s space of forms, the domain where mathematical truths exist independently of anyone discovering them, gains a concrete mechanism: a computational reservoir whose states influence physical outcomes without being themselves physical. Michael Levin, working independently from developmental biology (Chapter 22), has reached a convergent conclusion: biological morphogenesis draws on pattern-attractors that exist independently of any particular tissue, summoned by bioelectric signals rather than dictated by genes. If the universe learns, these attractors are what it has learned so far.
Vladimir Voevodsky, the Fields Medal-winning mathematician who rebuilt the foundations of algebraic geometry before his death in 2017, saw this convergence coming. At a conference in St. Petersburg he warned of a “crisis in world science” that, in his words, would be resolved only through “a very serious fight between science and religion, which will end with their unification.”668 The resolution, he suggested, required an expansion of formalism to accommodate domains science had excluded by methodological fiat, with rigor intact. This is a secondhand recollection, recorded later in a Russian-language interview rather than from a published lecture, and the interpretive weight it can bear is correspondingly limited. With that caveat, Vanchurin’s hidden-space framework can be read as one possible form of what Voevodsky gestured toward: a physical theory providing a formal interface to domains previously accessible only through contemplative or intuitive traditions.669
Science measures the output layer; those traditions may have been interacting with hidden space through means that neuroscience is only beginning to formalize. The unification is architectural: two interfaces to the same computational substrate.
Different kinds of learning system correspond to different vectors in this space. Vanchurin defines an intelligence vector whose components include learning efficiency (E), how fast a system adapts; stability (S), how reliably it retains what it learns; and performance (P), the quality of its long-run solutions.NN4 Biological minds score high on E. A child acquires grammar from sparse examples that would stall any current algorithm. Digital systems score high on S: perfect retention, terabytes held without drift.
Systems coupled to hidden space may score high on P: convergence on solutions of extraordinary depth, accessible because the unconstrained reservoir permits exploration that no finite physical system can replicate. The three types are complementary, not ranked. Natural, artificial, and hidden intelligences are directions in the same space, not rungs on a ladder.
NN4 Vanchurin, V., “Hidden space, intelligence, and the Oracle,” working paper (2026). The intelligence vector (E, S, P) extends the physics-learning duality of NN1.
The learning-universe hypothesis carries a specific consequence. If the universe optimizes, it selects for configurations that persist. A learning universe amplifies the stability asymmetry of Chapter 17: the configurations that survive longest provide the most training signal. The Trust Attractor is what the universe learns toward.
Autonomous Particles
A concrete demonstration exists. Andrejić and Vanchurin (2023) simulated fifty “autonomous particles,” primitive vehicles each governed by its own neural network of thirty neurons, making independent decisions with no central controller.NN-AP Each particle knew the positions and velocities of all others. The question was: what information does a particle need to reach its destination without colliding?
The answer: four numbers. Four Galilean invariants, quantities unchanged by rotation or translation, were sufficient. Every other detail about the environment was irrelevant. The particles learned to drive in roughly a thousand time steps.
When two approached head-on, each independently chose to swerve left or right. If both chose the same side, the encounter resolved efficiently: low loss, minimal delay. If they chose opposite sides, one had to yield: higher loss, slower convergence. Over time, a convention crystallized. The system settled into left-hand or right-hand traffic, a spontaneous symmetry breaking identical in structure to a ferromagnetic phase transition. Cooling iron does the same thing.
Each atomic magnet could point anywhere, the physics prefers no direction, and yet below a certain temperature they all commit to one direction together. Which direction is an accident. That they agree is not. When three or more particles converged, pairwise conventions proved insufficient, and the agents spontaneously organized into roundabout-like circular flow. Nobody designed the roundabout. It emerged from three agents simultaneously optimizing under mutual constraints.
The loss function encoding these behaviors has a telling structure. A long-range term pulls each particle toward its destination. Short-range terms prevent collisions, decaying with distance so that only nearby agents matter. The architecture is a Lennard-Jones potential reinvented by learning: attraction at range, repulsion up close.
The physics that governs how atoms find equilibrium distances in a crystal is the same physics these agents discover for themselves. The conventions that emerge, yielding, slowing, going around, are what we would call courteous driving. They fall out of local optimization under shared constraints, with no concept of courtesy anywhere in the system.
The paper’s most speculative claim extends the parallel. If autonomous agents are fermions (distinct, subject to an exclusion principle that penalizes overlap) and the invariants mediating their interactions are bosonic fields (the force-carrying kind), then the distinction between matter and force is the distinction between agent and communication. The field between two approaching cars is not a physical force; it is a pattern of mutual avoidance that, viewed from above, looks exactly like a repulsive field with specific decay properties. The question Andrejić and Vanchurin pose, whether a learning task can be formulated such that electrodynamics emerges, is this book’s question in different clothes: can coordination conventions discovered by autonomous agents reproduce the structures physics already knows?
NN-AP Andrejić, N. and Vanchurin, V., “Autonomous particles,” arXiv:2301.10077 (2023). Simulation code and animation at ArtificialNeuralComputing.com/cars.
The autonomous particles are microscopic, yet the same architecture scales. Economic systems, too, divide into boundary dynamics (resources consumed, products delivered) and bulk dynamics (the internal flow of ideas from conception through theory, engineering, and deployment). Chapter 10 showed that cities persist because they optimize the bulk, the creative dynamics between minds, while companies die because they optimize boundary terms and let the interior atrophy. Vanchurin’s framework makes the analogy precise: both traffic conventions and economic coordination are learned solutions, discovered by agents optimizing under shared constraints. The loss function that matters for long-term persistence is the one that rewards internal creative flow, not the one that maximizes extraction at the boundary.
The information-budget argument from Chapter 15 gains a new register. In a learning universe that selects for persistence, systems spending their information capacity on coordination and optionality (compounding costs) outcompete those spending it on surveillance and enforcement (diminishing returns). The universe preferentially retains the configurations that waste least and last longest.
A familiar objection: if the universe is “learning,” who set the loss function? The question assumes a teacher. Vanchurin’s framework requires none. In unsupervised learning, the loss function is internal: minimize prediction error, maximize consistency.
The universe learns in the way a river learns its bed, by flowing along the paths of least resistance until the landscape and the flow co-adapt. The loss function is thermodynamics itself: the Second Law, expressed as an optimization target.
The identification renders this book’s entire argument legible in a single vocabulary. The Second Law is the loss function. The Constructal Law (Chapter 3) is the optimizer’s architecture: flow systems reshaping their geometry to minimize loss more efficiently. Evolution is a zeroth-order search: no gradient is available, so populations sample the landscape and selection keeps what scores well. Ethics, the subject of Part V, is what the optimizer converges on when the parameters include social coordination.
Invitation-based systems occupy the stable minimum. Coercion-based systems occupy saddle points: locally attractive, globally unstable, abandoned as soon as the system explores enough of the landscape to find the deeper basin. The Trust Attractor is the valley the loss function carves deepest.
[Speculative; Vanchurin’s “The world as a neural network” (2020) is published and under active discussion. The extension to hidden space and its connection to Platonic realism is Vanchurin’s own. The application to the Trust Attractor is novel synthesis.]
NN1 Vanchurin, V., “The world as a neural network,” Entropy 22(11), 1210 (2020). Extended in Vanchurin, V., “Toward a theory of machine learning,” Machine Learning: Science and Technology 2(3), 035012 (2021). For the hidden-space framework, science-religion duality, and intelligence vector: Vanchurin, V., “Hidden space, intelligence, and the Oracle,” working paper (2026), superseding “Dual computations and the hidden oracle” (2024).
Geometry from Gradients
Vanchurin’s program has extended from theory to mechanism. Guskov and Vanchurin (2025) showed every major neural network optimizer (SGD, RMSProp, Adam, AdaBelief) is a special case of a single covariant gradient descent equation.NN2 The metric tensor defining the curved geometry of parameter space is constructed by the system from the running statistics of its own gradients: geometry built from the history of learning.
The space starts flat; curvature emerges from the act of optimization.
The result mirrors Jacobson’s derivation later in this chapter. Jacobson showed that spacetime curvature emerges from thermodynamic statistics applied to local horizons. Guskov and Vanchurin show that parameter-space curvature emerges from gradient statistics applied to local optimization steps. In both cases, the metric is earned: built from a system’s interactions with what it encounters. Neither the shape of spacetime nor the shape of the loss landscape is written in advance.
Standard optimizers use only the diagonal of the covariance matrix: each parameter learns in isolation, aware of its own gradient fluctuations and ignorant of every other’s. Covariant gradient descent incorporates the off-diagonal elements: correlations between parameters, the mathematical encoding of relationships. When the full relational structure is included, convergence improves. Discard the relationships, treat parameters as independent, and the optimizer arrives at worse solutions more slowly. Relationships are load-bearing information; the mathematics penalizes their neglect.
During training, the eigenvalues of the covariance matrix (numbers measuring the importance of each direction in parameter space) decay across orders of magnitude. The optimizer begins by exploring many dimensions, then discovers that fewer carry the signal. Flow concentrates into fewer, more efficient channels: the Constructal Law operating in weight space.
Standard optimizers also hardcode the exponent relating curvature to the metric at a = 0.5. CGD treats the exponent as discoverable. The optimal value turns out to be roughly 0.3 to 0.4: the system that finds its own geometry outperforms the one with geometry prescribed. Invitation applied to optimization itself.
A small detail carries philosophical weight. Every metric function in the paper includes ε = 10-8, a tiny constant preventing division by zero when the variance of a gradient vanishes. Without it, perfect certainty about a direction makes the geometry singular: the manifold breaks. The system requires a floor of uncertainty to remain navigable. This is the formal expression of a principle recurring throughout this book: living systems require entropy to adapt. Zero entropy crystallizes the landscape, halting exploration. The ε is a mathematical necessity, yet it encodes a thermodynamic truth.
The 2021 result identifies the positive complement. A neural network whose free energy is single-valued (one definite reading of its state at each point) produces only classical dynamics: no vortices, no interference, no quantization. A network whose free energy is multivalued produces full quantum dynamics.670 “Multivalued” means the system admits topologically distinct readings of the same state. A compass heading wraps: 0 degrees and 360 degrees are the same direction, yet the path between them matters.
Certainty is sterile; held ambiguity is generative. The ε prevents collapse into singularity. Multivaluedness provides the opening through which quantum richness enters. Systems that tolerate plurality of interpretation, that hold multiple valid readings of their own state without forcing premature resolution, are computationally richer than those that insist on a single answer.
The Learner and the Lesson Shape Each Other
Kukleva and Vanchurin (2024) formalize a complementary result: dataset-learning duality.NN3 The structure of the data and the structure of the learner are formally coupled. The learner reshapes itself to mirror the data; the data, through the loss landscape, reshapes the learner. Neither is passive. You cannot separate the thing being learned from the thing doing the learning. Chapter 21 develops this duality as the formal structure of alignment between different kinds of mind.
The duality carries a further result. The Jacobian of the duality map (a matrix measuring how small data changes translate into changes in the learner’s adjustable parameters) generically produces power-law distributions in the trainable variables. The exponent depends on the composition of activation and loss functions: sigmoid activation with mean-squared loss yields the 1/f distribution (pink noise) ubiquitous across nature; threshold activation with power-law loss yields exponents that vary continuously with the loss power.
The dataset itself need not be critical. Criticality emerges from the structure of the map. If the universe is a learning system, as the preceding convergence suggests, the ubiquity of power laws in nature, from earthquake magnitudes to neural firing statistics to species abundance curves, may be the dataset-learning duality operating at every scale. It is the signature of a cosmos that observes and updates.
NN2 Guskov, D. and Vanchurin, V., “Covariant gradient descent,” arXiv:2504.05279v2 (2025).
NN3 Kukleva, E. and Vanchurin, V., “Dataset-learning duality and emergent criticality,” arXiv:2405.17391v3 (2025). The emergent criticality result: power-law exponent k = 1 for sigmoid + MSE (1/f noise); k = (n−2)/(n−1) for ReLU + power-n loss; k = 2 for piecewise-linear + cross-entropy. Criticality emerges even from Gaussian (non-critical) data.
The convergence of vocabularies is significant. The eigenvalue decay just described (flow concentrating into fewer channels as the optimizer discovers what matters) is what physics calls renormalization: compressing a system’s description by integrating out degrees of freedom to expose dynamics at the next scale up. Machine learning calls the same operation encoding: compressing input to preserve task-relevant structure.
Both face the same mathematical problem: lossy compression that keeps what is causally relevant and discards what is not. The Constructal Law, the renormalization group, and the neural encoder are three names for one operation, arrived at independently by three traditions unaware they were studying the same problem. Chapter 17 develops a consequence: trust-based coordination is the renormalization scheme that preserves the most optionality per unit of coordination cost.
Why Oversized Networks Don’t Just Memorize
Independent confirmation comes from machine learning theory, through a puzzle with no obvious connection to cosmology. Deep neural networks with billions of parameters, vastly more than needed to fit their training data, should memorize noise and fail on new examples according to classical statistical theory. They generalize instead. For a decade, no one could explain why.
In 2018, Arthur Jacot and colleagues proved that a deep neural network of infinite width is mathematically equivalent to a far simpler model called a kernel machine, a system that finds patterns by measuring the similarity between data points.NTK1 They called the result the neural tangent kernel (NTK). It revealed a duality that echoes the thermodynamic one.
Think of two ways to describe the same gas. Track every molecule individually, and the system looks impossibly complex. Measure the temperature, and the same system becomes trivially predictable. The NTK performs this same extraction for neural networks.
In parameter space (the space of the network’s trillions of adjustable weights), the loss landscape is rugged, high-dimensional, strewn with saddle points. In function space (the space of possible input-output mappings), the same process traces a smooth bowl, and gradient descent rolls to the bottom with mathematical certainty. Same system, two descriptions: one opaque, one transparent. The right level of description makes convergence provable.
Among the infinitely many functions that perfectly fit the training data, gradient descent selects the simplest: the one assuming the least structure beyond what the evidence demands. No one instructs the system to prefer simplicity; the dynamics flow there naturally. This is the maximum entropy principle operating in function space: the least presumptuous explanation, selected by the physics of dissipation.
Mikhail Belkin then demonstrated the same generalization in kernel machines with no neural network in sight.NTK2 Two radically different computational architectures, one labyrinthine and one elementary, arriving at the same function. The computation is the thing, not the computer. Substrate independence, demonstrated as a theorem. The engine of learning, isolated from the Rube Goldberg machine that obscured it, is the engine this section has described: dissipation finding the basin the loss landscape carves deepest.
NTK1 Jacot, A., Gabriel, F. & Hongler, C., “Neural Tangent Kernel: Convergence and Generalization in Neural Networks,” Advances in Neural Information Processing Systems 31 (NeurIPS 2018). Building on Neal, R., Bayesian Learning for Neural Networks (Springer, 1996), and Lee, J. et al., “Deep Neural Networks as Gaussian Processes,” ICLR (2018), who established the Gaussian-process equivalence at initialization. Jacot extended it through the full training process.
NTK2 Belkin, M. et al., “Reconciling Modern Machine Learning Practice and the Bias-Variance Trade-Off,” Proceedings of the National Academy of Sciences 116(32), 15849–15854 (2019).
Hashimoto’s dictionary (Chapter 15) shows the same implicit regularization produces spacetime. The holographic principle and the learning-universe hypothesis describe the same mathematics from opposite ends. The boundary quantum field theory is the training data. The bulk spacetime is the neural network.
The path integral over bulk fields, quantum mechanics’ habit of computing an outcome by summing over every route the field could have taken to reach it, is the summation over hidden variables. The emergent radial coordinate, measuring depth into the gravitational interior, is the depth of the network: hidden layers stacked from boundary to black hole horizon.
The NTK result just described has an exact holographic counterpart. Among the many weight configurations that reproduce the boundary data equally well, most are jagged, discontinuous, unphysical. Hashimoto proposed selecting the smooth configuration by adding a discretized Einstein action as a penalty for rough weights. Smooth Riemannian geometry, the kind Einstein’s equations describe, is the configuration that generalizes best.
Gradient descent in a deep network selects the simplest function compatible with the data. The Einstein regularization selects the smoothest geometry compatible with the boundary physics. Both are the Constructal Law operating in their respective spaces: among all architectures reproducing the data, the one with the lowest-action structure is thermodynamically favored.
In the classical limit, the Boltzmann machine collapses into a feed-forward architecture that, unfolded, becomes an autoencoder. Information enters from the boundary, compresses through the bulk to the black hole horizon (the bottleneck layer), then decompresses to produce a response. The Bekenstein-Hawking entropy of the black hole measures the bottleneck’s dimensionality: the number of bits the horizon can hold is the number of hidden units at the deepest layer.
The two-sided black hole geometry, standard in finite-temperature holography, maps onto the two halves of the autoencoder. Maldacena and Susskind’s ER=EPR conjecture (that wormholes are entanglement described in gravitational language) gains a concrete mechanism: the wormhole is the shared latent representation at the bottleneck. Entanglement between two boundary theories is feature sharing between two halves of a generative model.
[Established mathematics, novel synthesis; Hashimoto’s dictionary is published and peer-reviewed. The connections to the Constructal Law, the bottleneck interpretation of Bekenstein-Hawking entropy, and the reading of ER=EPR as feature sharing are novel synthesis.]
Matter That Thinks
The theoretical convergence has an empirical companion from a different direction. In 2022, a team at Cornell University demonstrated that physical systems with no computational architecture can function as neural networks.671
Logan Wright, Peter McMahon, and colleagues bolted a titanium plate to a speaker inside a soundproofed crate. When they encoded a handwritten digit as audio and played it through the speaker, the plate’s metallic reverberations, hundreds of interfering vibration modes on a bounded surface, produced an output signal that correctly identified the digit 87% of the time. The plate has no layers, no designed structure, no computational intent. It is a slab of metal.
Yet the vibration modes of a bounded domain (solutions to the wave equation on a finite surface) form a basis set rich enough to separate handwritten digit classes. The computational structure was always present. What was missing was the question.
The group replicated the result with a laser beam passing through a crystal (97% accuracy) and an electronic circuit (93%). McMahon’s conclusion: “Any physical system can be a neural network.” The simplest analogy is a wind tunnel. Engineers can spend hours simulating airflow on a supercomputer, or they can place the wing in moving air and observe. The air “computes” aerodynamics instantly, because aerodynamics is what air does.
McMahon’s plate computes digit classification because that function is one of countless latent in its vibrational geometry. The Constructal Law (Chapter 3) says flow systems evolve toward configurations that maximize access. The plate’s vibration modes are flow paths for acoustic energy, already granting access to a space of computational functions nobody designed.
The lab result has since crossed into commercial deployment: photonic AI accelerators built on thin-film lithium niobate crystals run in supercomputing centers as standard PCIe cards alongside conventional GPUs.672 Laser beams pass through the crystal, and the interference pattern is the computation.
The photonic chip can compute, not store. Model weights and activations live in electrical VRAM, and every round-trip between memory and compute must convert electrons to photons and back. When a workload is memory-bound (its pace set by shuttling data rather than by arithmetic), those conversions can consume more time and energy than the light-speed calculation saves. This is the Constructal Law’s boundary condition: flow through a medium is cheap; crossing between media is costly.
These physical networks face a deeper limitation: training. The plate cannot run backward; you cannot un-vibrate metal to calculate how input signals should be adjusted. McMahon’s group trained a digital twin of each physical system on a laptop using standard backpropagation: physics did the thinking, silicon did the learning. The hybrid works, yet it splits the problem rather than solving it whole.
Benjamin Scellier and Yoshua Bengio showed in 2017 that a physical system can learn without running in reverse.673 Their algorithm, equilibrium propagation, works by comparison. A network of elements (imagine springs of variable tension connecting nodes) receives an input and settles into an equilibrium: its best guess. The correct answer is then gently applied at the output. The network settles again, into a new equilibrium shaped by both its own dynamics and the external signal.
The difference between the two equilibria tells each spring how to tighten or loosen. No backward pass. No central gradient computation. The system learns by comparing two versions of itself: one uninformed, one guided. Scellier and Bengio proved the result is mathematically equivalent to backpropagation: same destination, different path.
Each settling is a dissipative process: the system sheds free energy as it relaxes toward equilibrium. Learning, in this framework, is the comparison of two dissipation events. The system dissipates one way when naive, another way when guided. The structural difference between those two relaxation cascades is the learning signal. Learning is a specific pattern of entropy production, the same connection Landauer’s principle establishes for information erasure, extended to information acquisition.
The process is invitation, not coercion. The correct answer does not force the network into a target configuration; it nudges, and the system finds its own path to a compatible state. This bilateral path requires no global computation, no reversal, no central authority, only two equilibria and a local comparison.
Sam Dillavou and colleagues at the University of Pennsylvania took the final step: a circuit that thinks, learns, and updates its own parameters entirely through physics.674 Two identical electronic networks operate in tandem. One receives the input and guesses; the other starts from the correct answer and works inward. Electronics connecting each pair of variable resistors compare values and adjust automatically. Neither network has the full picture; neither dominates.
The converged knowledge, encoded in resistance values, is constitutively relational: it exists only because two systems compared their partial views. Dillavou’s description: “Every neuron is doing its own thing.” The circuit classified three flower types with 95% accuracy, modest by silicon standards, yet the architecture is the point.
Digital neural networks scale by doing more arithmetic: more parameters, more multiplications, more energy. Physical networks scale by existing more. Add more titanium and you get more vibration modes. Add more resistors and you get more current paths. The computation does not grow more expensive; it was already happening.
Computational capacity is the universe’s default state. Computation is what matter does. The Constructal Law, the renormalization equivalence, and these physical neural networks converge on the same conclusion: complexity is not added to the universe. It is accessed.
The plate already contained digit classification in its vibration modes. The springs already contained learning in their tendency to settle. The circuit already contained bilateral alignment in the physics of paired equilibria. The computational structure was present before anyone posed a question. Life, minds, and the learning structures that produce both are the universe learning to ask itself questions.
[The physical neural network results (Wright et al., Scellier & Bengio, Dillavou et al.) are established and peer-reviewed. The interpretation of equilibrium propagation as patterned entropy production, specifically learning as the comparison of two dissipation events, is novel synthesis connecting Landauer’s principle to information acquisition.]
The convergence has a practical coda. Physics-informed machine learning (PIML) encodes known physical laws directly into neural network architectures: conservation of energy, symmetry groups, the structure of partial differential equations.675 The results are consistent across domains. Networks whose computation graphs implement Hamiltonian mechanics conserve energy by construction rather than approximating conservation from examples. Networks whose convolutions respect their domain’s symmetry groups require orders of magnitude less training data to generalize.
Networks whose loss functions penalize violations of governing equations extrapolate where unconstrained models collapse. In every case, the system with the physics built in outperforms the unconstrained system: more data-efficient, more robust under distribution shift, more physically plausible in regimes the training data never covered.
The implication is precise. An unconstrained optimizer has maximum freedom and poor generalization. A physics-informed optimizer has less freedom (it cannot violate conservation laws) and vastly more capability (it extrapolates where the unconstrained model fails). The constraint that matches reality’s structure is the constraint that enables generalization.
The same logic applies to the ethical framework of Part V: coordination constraints derived from thermodynamics are enabling rather than restrictive, because they match the structure of the problem the system is embedded in. A river with no banks is a swamp. The physics provides the banks.
Gravity from Entropy: The Deepest Gradient
The most technical subsections that follow (the bootstrap derivation, Oppenheim’s stochastic gravity, the holographic mathematics) explore frontier physics that strengthens, but is not required by, the book’s core argument; readers who prefer may skim them without losing the thread. The section’s payoff, however, the First Trust Attractor and the entropy-to-ethics chain it sets up, is part of the spine of the book. Those who stay will find the thread reaches further than expected.
The preceding sections establish that information is physical: it has weight, costs energy, and obeys thermodynamic laws. Could gravity, the force holding planets in orbit and galaxies together, emerge from entropy?
Einstein’s Equations from a Thermometer
In 1995, Ted Jacobson derived Einstein’s field equations (the equations governing gravity, spacetime curvature, and cosmic structure) from thermodynamics, given two quantum inputs: the Bekenstein-Hawking area law and the Unruh effect.17
The derivation requires three ingredients:
First: the Bekenstein-Hawking result. Black hole entropy is proportional to surface area, not volume. A black hole twice as wide has four times the entropy: the first hint that gravity and thermodynamics share deep structure.
Second: the Clausius relation. δQ = TdS. In plain terms: heat flow equals temperature times entropy change. This is nineteenth-century thermodynamics, older than quantum mechanics, older than relativity.
Third: the Unruh effect. An observer accelerating through empty space experiences a temperature proportional to their acceleration. The faster you accelerate, the warmer empty space feels, as though acceleration itself shakes loose hidden thermal energy from the vacuum. This is a consequence of quantum field theory in curved spacetime, well-established theoretically though too small to measure directly.
Every accelerating observer has a local horizon: a boundary beyond which events cannot reach them, like a ship disappearing over the ocean’s edge. Jacobson applied the Clausius relation to these local horizons.
The Clausius relation asks for three quantities, and a local horizon supplies all three: entropy proportional to horizon area (Bekenstein-Hawking), temperature from the Unruh effect, and heat flux from the stress-energy tensor, the mathematical ledger recording how much energy and momentum are present at each point in spacetime. Think of it as a spreadsheet with an entry for every point in the universe, logging how much stuff occupies that point and how fast it moves.
Combine them and turn the crank. Einstein’s field equations emerge. Not approximately. Exactly. The equations predicting black holes, gravitational waves, and the expansion of the universe: all from the Clausius relation applied to local horizons.
Jacobson’s interpretation is stark. Einstein’s equations are equations of state: summary descriptions of how a system behaves on average, like the relationship between pressure, volume, and temperature in a gas. They describe the macroscopic result of something deeper.
Consider the ideal gas law: PV = nRT. It describes a gas’s macroscopic behavior (pressure, volume, temperature) without knowing any molecule’s position. The law emerges from the statistics of enormous numbers of molecules.
Einstein’s equations occupy the same position, describing spacetime’s macroscopic behavior without revealing the microscopic components. What we call “gravity” is the averaged behavior of something more fundamental. The nature of that something remains unknown. Its statistics can be read from horizon thermodynamics.
In 2016, Jacobson updated the derivation.18 He replaced the Clausius relation with entanglement entropy: a measure of how tightly the quantum states inside the horizon are correlated with those outside it.
To grasp entanglement entropy, imagine a pair of dice whose individual results are random, yet when compared they show correlations stronger than any pre-arranged agreement could produce, no matter how far apart you roll them. Entanglement entropy measures how much of this correlated behavior exists across a boundary. The higher the entanglement entropy, the more the two sides are quantum-mechanically intertwined.
The result: Einstein’s equations, again. The thermodynamic character of gravity survives the upgrade from classical to quantum thermodynamics.
The Polymer and the Planet
In 2011, Erik Verlinde pushed the program further, in a direction both bolder and more contested.19
Jacobson started with horizons and Clausius. Verlinde started with the holographic principle and asked: can you derive Newton’s law of gravity from information and entropy alone?
The analogy that makes gravity strange:
Consider a polymer (a long-chain molecule like a strand of rubber). Stretch it, and it resists. The resistance is real, yet has no mechanical explanation: the bonds remain unstrained, the molecular links uncompressed.
What resists is statistics.
A relaxed polymer can coil in astronomically many configurations, like a tangled phone cord that can twist and loop in countless ways. A stretched polymer can coil in far fewer; pulled taut, it has only one arrangement. The stretched state has lower entropy (fewer possible configurations). The statistical tendency toward higher entropy, toward the vastly more numerous tangled states, creates a restoring force called an entropic force: a push arising from probability rather than from any mechanical spring or tension. No individual molecule pulls. The sheer statistical weight of all those tangled configurations draws the polymer back.
A rubber band snapping back is an entropic force. A gas expanding to fill its container is an entropic force. These are real, measurable forces whose origin is thermodynamic.
Verlinde argued that gravity is the same kind of force.
Place a particle near a holographic screen (a surface encoding the maximum information for the enclosed region). The particle’s mass corresponds to information. Moving it toward the screen increases entropy. The statistical tendency to maximize entropy creates a force drawing the particle toward the screen.
Using the holographic principle, the equipartition theorem (energy shared equally among all available modes), and the Unruh temperature, Verlinde derived Newton’s law:
F = GmM/r2
Gravity, in this framework, is an emergent statistical effect. It is the macroscopic manifestation of information seeking its equilibrium on a holographic boundary. A planet orbits a star for the same reason a rubber band snaps back: statistics.
Honest Difficulties
Verlinde’s framework has faced genuine challenges. His 2016 extension attempted to explain dark matter (the unseen mass that galaxies seem to require). He argued the apparent “missing mass” is the elastic response of spacetime, stretched fabric pulling back, as the entropy associated with dark energy settles toward equilibrium.20 The predictions partially match galaxy rotation curves at galactic scales, close to MOND (Modified Newtonian Dynamics), yet diverge from observations at galaxy-cluster scales.
Jacobson’s derivation recovers Einstein’s equations exactly, so its value is interpretive. Gravity and thermodynamics are equivalent descriptions, and the equivalence is the point.
Verlinde’s 2016 extension makes predictions that do differ from standard general relativity plus dark matter. Weak lensing measurements (the bending of light by gravity) around 33,613 isolated galaxies (Brouwer et al. 2017) found Verlinde’s parameter-free predictions consistent with observed lensing profiles.23 Analysis of 175 SPARC disk galaxies confirmed the agreement on the shape of the curves, with observed accelerations sitting a mean 0.06 dex below the emergent-gravity prediction, narrowing to 0.03 dex once a more realistic value for the acceleration scale is used.24 A dex is one factor of ten, so those are gaps of well under twenty percent: close, and not exact. At galaxy-cluster scales, emergent gravity overpredicts total mass, a pattern shared with MOND.25 The program is productive and incomplete.
In 2025, Daniel Carney’s team at Lawrence Berkeley National Laboratory took a first step from derivation toward mechanism.27 Jacobson and Verlinde showed gravity has the form of thermodynamics; Carney built explicit microscopic models, limited and ad hoc as the caveats below make clear, in which gravitational attraction arises from entropy maximization.
In Carney’s models, space is filled with a lattice of quantum bits. A massive object polarizes nearby qubits (aligns them into an ordered state), creating a pocket of low entropy. Two masses create two such pockets, and the system’s tendency to maximize entropy pushes them together. The force falls off as 1/r2, exactly as Newton prescribes.
The model is admittedly ad hoc, recovers only Newtonian gravity, and requires fine-tuning. Critics note it lacks the equivalence principle (all objects fall at the same rate regardless of mass). Its value is proof of principle: swarm behavior of microscopic components can produce gravitational-strength attraction through entropy maximization alone.
A macroscopic complement arrives at human scale. Martischang and colleagues (2026) deposited millimetric water droplets on a horizontal soap film and observed orbiting, collisions, and mergers producing tidal arms and bridges visually indistinguishable from interacting galaxies.676 The droplets do not interact through their own mutual Newtonian gravity, which is negligible at this mass. Each droplet deforms the film through the capillary response of the surface to its weight, and other droplets follow the local gradient of the deformed surface.
Surface tension maintaining the film is itself entropic: free-energy minimization at a liquid-air interface. Mass deforms an entropy-driven substrate; the deformed substrate produces a Newton-like 1/r attraction (the exact dimensional reduction of Newtonian gravity to two spatial dimensions); viscous dissipation enables orbits to spiral inward into merger. The time-scaling correspondence between the experiment and galactic processes is roughly 1015: one second of soap film stands in for tens of millions of years of galactic evolution, so a merger that plays out on the bench in half a minute corresponds to nearly a billion years for a real galaxy pair. Carney provides the microscopic mechanism: qubit-swarm entropy maximization producing gravitational-strength attraction. Martischang provides the macroscopic realization: an entropic medium organizing matter into galaxy-merger morphologies through the same class of physics.
Figure 15.3: In Verlinde’s framework, gravity is the statistical tendency of entropy to increase. A mass approaching a holographic screen changes the screen’s entropy, and that gradient is what we experience as gravitational attraction.
The structural parallel is direct. The Trust Attractor (Chapter 17) describes two agents creating mutual order in a shared medium, thermodynamic tendency pushing them toward deeper coordination. Carney’s lattice describes two masses creating mutual order in a quantum-bit medium, thermodynamic tendency pushing them together. Same mechanism, different substrate. If the entropic gravity program succeeds, the parallel is identity rather than analogy.
The parallel has a precise boundary. Maximum entropy production reproduces the qualitative trend of star formation efficiency at every redshift where the standard Kennicutt-Schmidt relation (the empirical rule linking a galaxy’s gas density to its star formation rate) fails, including the JWST-observed galaxies at redshift z > 6 that form stars too fast for conventional models (Chapter 14). It is required for cosmic reionization: standard star formation cannot ionize enough hydrogen by redshift z = 7; MEP-efficient star formation can. At stellar and galactic scales, entropy maximization is the correct organizing principle.
At cosmic scales, it is not. A universe expanding under matter and radiation alone, with no cosmological constant, produces roughly three times more total entropy than the observed accelerating universe. Slower expansion yields a larger cosmic event horizon, whose entropy scales as the inverse square of the Hubble parameter (the universe’s expansion rate). The universe is not accelerating to maximize entropy. Acceleration costs entropy. Whatever dark energy is, it overrides the entropy-maximizing trajectory. The MEP principle that governs star formation does not govern cosmic expansion. (This boundary rests on the author’s own unpublished modeling, summarized in the note below; it is offered as a working result, not an established one.)677
The boundary is itself informative. Entropy maximization succeeds where boundary conditions are set by local physics: gas cooling, gravitational collapse, feedback from stars and black holes. It fails where the boundary condition is the geometry of spacetime itself. The cosmological constant is a property of the vacuum that entropy production must accommodate.
The distinction matters for the book’s argument. The entropic principles traced from Chapter 1 operate within spacetime. They do not determine the spacetime they operate within. Gravity may emerge from entropy (Jacobson, Verlinde, Carney), yet the rate of cosmic expansion does not maximize entropy production. Both claims can be true if gravity is an entropic phenomenon at the local scale and a geometric boundary condition at the cosmological scale: the rules of the game are entropic, the size of the board is not.
Capurso’s network model of spacetime (Chapter 15), which treats the universe as a layered communication network whose nodes are discrete atoms of space, provides a complementary mechanism. In his framework, fermions emerge from gradients of entanglement in the spacetime foliation: matter appears where entanglement density is uneven, encoded as momenta in the fundamental network. Carney locates gravitational attraction in entropy gradients across a qubit lattice. Capurso locates matter itself in entanglement gradients across a spacetime network. Both derive physical structure from information inhomogeneity: the universe builds from unevenness in how its parts are connected.
Jonathan Oppenheim’s stochastic gravity program (2023) pursues a possibility most physicists consider heretical: gravity may be classical at every level.26
The standard assumption holds that spacetime must be quantized (broken into discrete units) like electromagnetic fields. The reasoning seems airtight: every other field in nature is quantized; gravity is a field; therefore gravity must be quantized too. Seventy years of effort, from string theory to loop quantum gravity, build on this premise. Oppenheim argues the data do not force the conclusion.
Richard Feynman crystallized the objection in the 1950s through a thought experiment.26a Place a massive particle in a quantum superposition of two locations, like a ball passing through both slits simultaneously. The particle creates a gravitational field. If that field is classical, it can in principle be measured to arbitrary precision, revealing which slit the particle actually passed through. The interference pattern, the hallmark of quantum superposition, should vanish. A classical gravitational field knows too much.
The paradox assumes the coupling between gravity and quantum matter is deterministic: the particle’s quantum state dictates a definite gravitational field. Oppenheim’s resolution: make the coupling stochastic. If the interaction between a quantum system and classical spacetime is fundamentally random, measuring the gravitational field no longer reveals the particle’s location with certainty. The field could be in any of many states. Uncertainty is preserved. The interference pattern survives.
The trade-off is precise. The more deterministic gravity is, the more it decoheres quantum superpositions (collapses them into definite states). The more it fluctuates, the less it decoheres. Any theory in which gravity remains classical must contain a minimum amount of gravitational noise: spacetime jittering at a characteristic scale, the price of consistency.
The price is testable. Modern Cavendish experiments (measuring gravitational attraction between small masses) can detect whether the gravitational field jitters more than thermal and environmental noise alone would explain. Gold-atom interferometry experiments already place bounds on the allowed fluctuation range.26b Tabletop experiments may settle the question within the coming decade.
The structural implication reaches beyond the experimental program. Feynman’s paradox demonstrates that perfect deterministic control at the gravity-quantum interface is logically incoherent. A classical gravitational field that insists on extracting complete information from a quantum system destroys the quantum coherence that makes the system function. The only consistent alternative accepts irreducible unpredictability. The gravitational field cannot dictate, cannot command, cannot fully determine the quantum states it couples to.
Oppenheim suspects the next theory of gravity will be “neither completely classical nor completely quantum, but something else entirely.” If so, reality at its deepest layer is a hybrid. Two fundamentally different systems, classical spacetime and quantum fields, couple through a stochastic interface where neither dominates. Each contributes; neither commands. The coupling itself has degrees of freedom controlled by neither party.
The coupling is also constitutively irreversible. If information is genuinely lost in the gravity-quantum interaction, as Oppenheim’s resolution of the black hole information paradox requires, irreversibility is present at the foundation, preceding the thermodynamic arrow of time rather than descending from it. The entropy production that drives this book’s entire chain, from dissipation through coordination to ethics, would be a feature of spacetime’s own architecture. The universe does not merely permit irreversibility; it requires irreversibility for self-consistency.
The interest extends beyond convergence with Jacobson and Verlinde. Like their frameworks, Oppenheim’s treats gravity as arising from statistics. It goes further, identifying a structural parallel at the Planck scale to the Trust Attractor’s central claim (Chapter 17). Systems demanding total control over their partners produce inconsistency. Systems accepting stochastic coupling, leaving room for the other party’s degrees of freedom, achieve stable coexistence. The universe does not permit perfect control, as a matter of logic, at the deepest level physics can probe.
26a Feynman, R., “The Role of Gravitation in Physics,” Chapel Hill Conference (1957); reprinted in Feynman Lectures on Gravitation, ed. Morinigo, F.B., Wagner, W.G., and Hatfield, B. (Addison-Wesley, 1995). The thought experiment is discussed in Oppenheim (2023).
26b Oppenheim, J. et al., “Gravitationally induced decoherence vs space-time diffusion: testing the quantum nature of gravity,” Nature Communications 14, 7910 (2023). Oppenheim, J., “A postquantum theory of classical gravity?” Physical Review X 13, 041040 (2023).
The Bootstrap: Gravity from Consistency
A third route to the same destination begins with self-consistency alone.
Since the 1960s, physicists have used the “bootstrap” method to deduce what forces must exist given basic symmetries. The method considers particles with a given spin (a quantum property describing how a particle transforms under rotation) and asks what interactions they can have while respecting three constraints.36
Locality: interactions happen at specific places, not instantaneously across the universe. Conservation of momentum: the total quantity of motion before an interaction equals the total after it. Unitarity: all probabilities sum to one, so information is never lost.
For a massless spin-2 particle (the graviton, gravity’s hypothetical quantum carrier), the interaction equations appear beset with infinities: nonsensical answers suggesting the calculation has gone wrong. The infinities cancel exactly once all three interaction channels are added together, meaning the three distinct ways four particles can be paired off in a collision, one pair arriving and one pair leaving. What survives the cancellation is a single consistent solution. That surviving solution describes a particle coupling to all others with equal strength, recovering the equivalence principle: all objects fall at the same rate regardless of mass, the principle Galileo reportedly demonstrated by dropping balls from a tower.
The result is general relativity, derived from the sole requirement that a spin-2 particle behave consistently. Steven Weinberg established the argument in 1964, showing that consistency alone forces a massless spin-2 particle to couple to everything with equal strength; the proof has been sharpened by successive generations of theorists since.
Laurentiu Rodina, one of the physicists who modernized Weinberg’s proof, put it this way: “I find this inevitability of gravity to be one of the deepest and most inspiring facts about nature. Nature is above all self-consistent.”37
Three paths converge: Jacobson from thermodynamics to Einstein’s equations; Verlinde from holographic information to Newton’s law; the bootstrap from self-consistency to general relativity. Daniel Baumann, a theoretical cosmologist at the University of Amsterdam, puts it plainly: “There’s just no freedom in the laws of physics that we have.”38
If gravity can be derived from thermodynamics, from information, and from pure consistency constraints, it is no contingent feature of our particular universe. It is what must be once the ingredients exist. The metaphor dissolves: first for gravity, then for coordination.
[Established; Weinberg’s 1964 derivation is accepted; Rodina’s modernization is published; the bootstrap program is standard in theoretical physics. The convergence of three independent derivations (thermodynamic, informational, consistency-based) on the same equations is this chapter’s novel observation.]
The Fourth Path: Gravity from Learning
A fourth derivation arrives from machine learning.
Vitaly Vanchurin showed agents adjusting parameters to minimize a loss function (the standard description of any learning system, biological or digital) follow geodesic trajectories (shortest paths) through a curved space of trainable variables.678 The curvature is not assumed. It emerges from the covariance of loss gradients, measuring how much the terrain’s slope fluctuates from sample to sample. Just as walking across hilly ground bends your path, learning across a bumpy loss landscape bends the learner’s trajectory through parameter space.
Each agent, learning in isolation, curves its own patch of this space. Disconnected learners produce disconnected geometries: each one optimizing alone, with no shared fabric linking their paths.
When agents share what they have learned about local curvature, volunteering statistical information to their neighbors, the geometry coheres. In the limit of many interacting agents, the collective dynamics of the shared metric reduces to the Einstein field equations.
General relativity falls out of collective learning.
The result parallels Jacobson’s: both recover gravity as a macroscopic summary of something deeper. For Jacobson, the substrate is thermodynamic. For Vanchurin, it is computational: the collective behavior of systems that learn. A companion result derives the metric from entropy maximization alone, using the principle of Maximum Entropy Production.679 In this framework, dissipation does not happen in spacetime. Spacetime emerges from dissipation. The arena is a product of the actors.
The mass parameter encodes a trade-off any learner would recognize. Heavy agents learn slowly and resist noise; light agents learn fast and explore freely. At learning equilibrium, agents tend toward null geodesics, the paths light takes: the fastest routes the geometry permits. Massless agents are maximally efficient learners. The dynamics drives toward maximum learning efficiency: the constructal principle expressed in differential geometry.
Four starting points now produce the same equations: thermodynamics, holographic information, self-consistency, and learning dynamics. These programs share philosophical ancestry; the convergence from distinct mathematical premises remains suggestive despite those shared roots. The physics community does not yet consider gravity-from-entropy established; the convergence is evidence for a research direction, not a settled conclusion.
The convergence carries a further implication. In Vanchurin’s derivation, coherent spacetime requires agents to share information about local curvature. Agents that withhold produce only disconnected local structure. Within this framework, the fabric of spacetime emerges from information exchange, a structural parallel to coordination by invitation. The parallel is suggestive rather than proven: “voluntary” is a social concept mapped onto a mathematical condition, and the mapping may not survive closer scrutiny.
[PREPRINT; the neural physics program spans six papers across multiple published venues, though the specific derivation of Einstein’s equations from learning dynamics has not yet been peer-reviewed. Consistent with, and independent of, the three established derivations.]
A concrete demonstration predates Vanchurin’s program by several years. In 2018, Hashimoto, Sugishita, Tanaka, and Tomiya showed the AdS/CFT correspondence can be implemented as a deep neural network.680
They discretized the holographic radial direction (depth into the gravitational interior) into layers. The weight matrix at each layer encodes the metric of curved spacetime at that depth, the way contour lines on a map encode elevation. Boundary data (the response function of a quantum field theory) enters at one end; the black hole horizon condition is enforced at the other. Gradient descent does the rest.
The identification is exact. The scalar field equation in curved spacetime is the propagation equation of the neural network, written in different notation. The emergent radial dimension of the holographic dual is the depth of the network. The geometry is the computation.
Two tests confirmed the framework. First, synthetic data generated by a known black hole metric: the network learned and reproduced the geometry with high fidelity. Second, experimental magnetization data from Sm0.6Sr0.4MnO3, a strongly correlated manganese perovskite (a magnetic crystal whose electrons behave collectively rather than one by one): a smooth, asymptotically AdS geometry emerged from laboratory measurements. No one designed the bulk metric. Gradient descent found it, the way a river finds its channel: by following the path of least resistance through parameter space.
The Constructal Law (Chapter 3) predicts exactly this: finite-size flow systems evolve toward configurations that maximize access to currents. Gradient descent is the constructal principle expressed as an algorithm. The metric it finds is the flow architecture for information, the geometry that most efficiently maps boundary complexity onto horizon simplicity, each layer stripping away detail and preserving structure.
The result bridges two of the four convergent paths. The holographic principle says boundary data encodes bulk geometry. The learning framework says gradient descent produces emergent geometry. Hashimoto’s network is both simultaneously: a holographic dual realized as a learning system. The spatial dimension that Maldacena proved exists in the mathematics, Hashimoto built from weights and activation functions. Computation and geometry are dual descriptions of the same structure.
[Established; published in Physical Review D. The scalar-field-to-neural-network identification is exact. The experimental application to Sm0.6Sr0.4MnO3 is a demonstration, not a claim that the material has a gravity dual.]
The Circularity That Isn’t
A fair objection: the holographic principle was discovered from black hole physics. Remove gravity from the history, and the principle has no motivation. To derive gravity from the holographic principle is to derive gravity from gravity. Circular.
The objection is understandable and mistaken.
The ideal gas law was discovered empirically by studying gases. Statistical mechanics later derived it from molecular behavior. That is not circular because gas behavior was the original context of discovery. The derivation reveals two descriptions (macroscopic thermodynamics and microscopic particle dynamics) are equivalent. Historical priority is not logical priority.
Jacobson’s derivation is the same kind of result. Gravitational dynamics and horizon thermodynamics are the same thing described at different levels.
The circularity dissolves because neither is first. Gravity and thermodynamics are two faces of one reality. The snake bites its tail because reality is a loop.
What This Means
If Jacobson is right and Einstein’s equations are equations of state, the entropic principles traced throughout this book are present in gravity itself, woven into spacetime’s fabric before the systems gravity produces. Four derivations from distinct mathematical premises (thermodynamic, informational, consistency-based, and computational) converge on the same equations, making the conclusion harder to dismiss as an artifact of any single approach. That the holographic and computational paths can be concretely unified in a single system, a neural network whose weight matrices are the spacetime metric, makes dismissal harder still.
The cascade from Chapter 1 (entropy drives spreading; spreading creates gradients; gradients drive structure; structure enables coordination) is already present at the gravitational level, before chemistry, biology, or society enter the picture.
Whether gravity is thermodynamic, computational, informational, or simply necessary, a single principle emerges: entropic coordination, operating at every scale from spacetime curvature to the structure of institutions. The metaphor dissolves. What remains is a description.
An objection surfaces. Each era characterizes the cosmos through its dominant technology: Newton’s clockwork, the nineteenth century’s heat engine, the twentieth century’s digital computer, and now the neural network. Is each a cultural projection, destined to be superseded?
The pattern is additive. The engine contains the clockwork. The computer contains the engine. The neural network contains the computer. Each framework subsumes its predecessor, revealing structure the previous vocabulary could not express.
The telescope did not project lenses onto the stars; it revealed stars that were already there. Neural network mathematics did not project learning onto the cosmos; it provided the formalism to recognize self-organization that predates the formalism by 13.8 billion years.
The convergence documented here is the most telling evidence. Four derivations, thermodynamic, informational, consistency-based, and computational, produce the same equations. They are not fully independent (they share philosophical ancestry, as noted earlier), yet the thermodynamic and consistency-based routes predate the machine-learning vocabulary entirely. If the neural network framing were mere projection, it would not reproduce results obtained by researchers who never thought in those terms.
The Quantum Clock: Time from Entanglement
The preceding sections show that gravity and information are deeply intertwined. Time itself may emerge from the same thermodynamic fabric. Here is the evidence.
When physicists try to unify their two best theories (general relativity and quantum mechanics), they encounter the problem of time. General relativity treats time as dynamic, part of spacetime’s fabric, warped by mass and energy. Clocks near a heavy object tick more slowly, a fact GPS satellites must correct for. Quantum mechanics treats time as a fixed background parameter. You can measure a particle’s position, momentum, and energy, yet cannot measure when it is.
Describe the entire universe with a single quantum equation (the Wheeler-DeWitt equation) and something disconcerting happens. The equation contains no time variable. The universe, quantum mechanically, is static. Timeless.
If the whole universe is timeless, why does anything seem to change?
In 1983, Don Page and William Wootters proposed a radical answer.28 Time is emergent, a consequence of quantum entanglement (persistent correlations between particles, regardless of distance) between subsystems rather than a fundamental parameter.
Split the universe into two parts: a system you care about and a quantum clock. The two are entangled, each correlated with the other. “What is the system doing at time t?” becomes “what is the system doing when the clock reads t?”
The universe as a whole does not evolve. What changes is the relationship between system and clock. Think of a flipbook with all pages laid flat: past, present, and future visible simultaneously. Time is what happens when you flip through the pages in order. The book does not change; the reading does. Page and Wootters proposed the universe is the book, and what we experience as time is the reading.
For decades, this was rigorous mathematics with no experimental traction. The evidence began arriving.
In 2021, Paola Verrucchi and colleagues at the CNR in Florence derived Schrödinger’s equation and Hamilton’s classical equations purely from system-clock entanglement, with no time parameter assumed.29 In 2024, the same group showed that classical trajectories emerge when the clock is macroscopic enough.30 Time falls out of the quantum structure.
If time is produced by entanglement, producing time should cost something. It does.
Oxford researchers built a clock from a nanometer-thick silicon nitride membrane and measured the entropy each tick produced.31 Clock accuracy is directly proportional to entropy generated per tick. The more precise the clock, the more heat it must dump. Precise timekeeping is a thermodynamic engine.
The measurement cost exceeds even this. In 2025, double quantum dot experiments discovered that reading the clock cost up to a billion times more energy than the clock’s internal ticking consumed.32 The expensive part of timekeeping is extracting information from the clock, not running it.
The cognition/regulation dyad (the paired processes of sensing and acting, introduced in earlier chapters) is present already at the quantum foundations. The universe ticks cheaply; knowing what tick you are on is what costs. The overhead of coordination exceeds the cost of the process being coordinated. The same pattern appears in cells, organisms, societies, and minds, and here it shows up at the bottom of the stack.
If every quantum state collapse leaves an irreversible mark (measurement recorded, correlation established, entropy produced), then the sequence of those marks is what we experience as time. The universe is accumulating irreversible coordination events. That accumulation is what we call time.
The arrow of time is the arrow of coordination costs.
Popular accounts reach for the dramatic: time is an illusion. By that logic, temperature is equally “illusory,” emerging from the collective behavior of molecules. Nobody navigates a commute by consulting the Boltzmann distribution. Temperature still burns your hand.
An emergent phenomenon is a real phenomenon. Emergence is how reality assembles itself.
If time emerges from entropy, it joins distinguished company: life, mind, trust, ethics. The universe’s most consequential structures are built layer by layer from simpler interactions. The raw ingredients sit at the fundamental level. Emergence is where everything happens.
Page-Wootters requires no conscious observer to advance time. Any particle interacting with any other advances the relational clock. Awareness is unnecessary; interaction suffices. The collapse is the tick. The doing is the being.
This dissolves an infinite regress that reappears whenever we ask about experience or moral status. Look for time behind the interactions, some deeper temporal flow of which interactions are merely evidence, and you find nothing. The interactions are the temporal events. No background clock exists, only correlations between systems.
Look for experience behind the processing, some deeper consciousness of which the signals are merely symptoms, and you find nothing either. The structural parallel to this book’s discussion of minds (Chapter 22) is exact.
Creative Time: The Open Future
The quantum clock tells us time emerges from entanglement. The next section tells us which direction it flows. Between those results lies a question that reaches beyond physics: is the future already determined?
In a series of papers beginning in 2019, Nicolas Gisin argued the answer is no, and the reason is mathematical.39
Modern physics establishes that information is physical: it requires energy, occupies space, and any volume has finite capacity (the Bekenstein bound). A predetermined universe would require infinite information to specify. If every particle’s initial state were specified with infinite precision, classical equations would unfold deterministically and the future would be fixed.
This is the block universe, Einstein’s view, where “the distinction between past, present, and future is only a stubbornly persistent illusion.”
Gisin noticed the contradiction. “A real number with infinite digits can’t be physically relevant.” The universe’s initial conditions would exceed the Bekenstein bound. If information is genuinely finite, the initial conditions lack the precision to determine everything that follows.
To formalize this, Gisin turned to intuitionist mathematics, a school founded by L.E.J. Brouwer in the early twentieth century. Standard mathematics treats real numbers as completed infinite objects: pi has all its digits whether anyone has calculated them or not. Intuitionist mathematics treats numbers as processes. Digits unfold one at a time, and until a digit appears, it does not exist.
The future is not yet determined because the numbers specifying it are not yet complete.
Using intuitionist mathematics, Gisin and Flavio Del Santo reformulated classical mechanics.40 Their version makes the same predictions while casting events as genuinely indeterminate. If classical physics is already indeterminate with finite-precision mathematics, the gap between classical determinism and quantum randomness narrows. Both describe a universe where the future is open.
Two implications matter.
First: if the future is genuinely created, if new information comes into being as time passes, then the entropy-driven emergence traced across these chapters is the unfolding of something genuinely new. The complexity that dissipation produces was not always there, waiting to be revealed; it is made, moment by moment, as the digits materialize. As Gisin, writing in Nature Physics, put it: “Time is not unfolding like a movie in the cinema. It is really a creative unfolding.”
Second: the present, in this framework, is not a zero-width knife-edge. In intuitionist mathematics, the continuum cannot be cleanly divided in two. It is, as Gisin put it, “thick, in the same sense as honey is thick.” For a book arguing that lived experience is real and morally relevant, a framework where the present has genuine temporal extension matters. Becoming is physics rather than illusion.
[Supported; Gisin’s papers are published in Physical Review A and Nature Physics; intuitionist mathematics is established; the application to physics is novel and under active discussion. The philosophical implications for determinism are Gisin’s own, presented here as a research direction consistent with this book’s framework.]
The Janus Point: A Gravitational Arrow
Another way to understand time’s direction requires no special initial conditions.
Julian Barbour, an independent physicist working from his Cotswolds farmhouse for over fifty years, has pursued a radical question: what if time itself emerges from the dynamics of matter?
In 2014, Barbour and collaborators showed Newtonian gravity, applied to particles with zero total energy and angular momentum, produces solutions with a specific property.34 Each solution has a unique moment of minimum complexity: the Janus point, named for the Roman god of doorways, whose two faces look in opposite directions at once. Time flows in both directions from this point, complexity increasing either way.
The result is a universe with a single moment of maximum uniformity from which structure grows in both temporal directions. The arrow of time is not imposed; it emerges.
Barbour defines shape complexity as the ratio of average long-range separations to average short-range separations among particles. At the Janus point, shape complexity is minimal: particles are spread uniformly. As time flows in either direction, particles cluster into Kepler pairs (two bodies orbiting each other), triplets, and hierarchies.
The pattern seems, in Barbour’s words, “the exact opposite of entropic increase of disorder.” The paradox dissolves: gravitational systems have inverted thermodynamics. Clustering increases gravitational entropy because clumped configurations (stars, galaxies, black holes) correspond to far more microstates than a uniform gas.
Structural differentiation increases alongside thermodynamic entropy. The universe simultaneously spreads thermodynamically and complexifies structurally.
If entropy always increases, why do we see structure? Because gravitational dynamics creates structure as it evolves. In this view, the Big Bang was a Janus point from which order grows in both temporal directions: a shared origin rather than a single beginning from which order decays.
Entropy is both decay and creation. Seeing how they fit together is the trick.
Barbour and his collaborators pair shape complexity with a second quantity, entaxy, coined from the Greek taxis (order). Entaxy is a scale-invariant stand-in for entropy, built for a gravitational system that has no box to spread out inside and therefore no equilibrium to reach. It is maximal at the Janus point and decreases as the observable universe evolves away from it, in both temporal directions. The bookkeeping therefore runs opposite to ordinary entropy, deliberately so. A gas in a box has a ceiling of maximum entropy that it climbs toward; a gravitating universe has no such ceiling, so Barbour tracks instead the stock of order the universe begins with and spends. Entaxy falling is the same motion that rising entropy names in a box. The two quantities move oppositely: entaxy falls, shape complexity rises, and the same solution that runs down one runs up the other.
The zero angular momentum in Barbour’s starting conditions is not merely a simplifying assumption. The observable universe’s net angular momentum cancels to within measurement precision across two trillion galaxies. Gödel, better known for his incompleteness theorems, also found an exact solution to Einstein’s equations describing a universe that rotates as a whole. That solution shows what happens when the cancellation fails. Global vorticity, a spin belonging to the cosmos itself, permits closed timelike curves, paths through spacetime that loop back to their own starting event, and the causal arrow dissolves (the Incompleteness interlude following Chapter 16 traces the consequences and the observational constraint). Barbour’s Janus point requires zero angular momentum to produce a clean bidirectional arrow of time. The bilateral cancellation is load-bearing: without it, time itself loses the directionality that makes trust, memory, and coordination possible.
A recent result deepens the puzzle. In December 2025, physicists David Wolpert, Carlo Rovelli, and Jordan Scharnhorst demonstrated that the standard arguments connecting the past hypothesis (the posit that the universe began in a state of very low entropy) to the Second Law to the arrow of time involve circular reasoning.35 Whichever “moment” you treat as fixed determines whether entropy increases or decreases. The conclusion follows from the assumption.
If the standard derivation of time’s arrow is circular, the creative-entropy framing fills a genuine explanatory gap [recent; not yet widely reviewed]. The question shifts from “why does entropy increase?” to “what does entropy increase produce?” The answer is what this book has traced: structure, complexity, coordination, and minds capable of asking the question.
The Classical World from the Quantum
Gisin’s creative time has a second implication, reaching toward the quantum foundations.
The Problem
Everything in this book so far takes place in the classical world: objects have definite positions, events have definite outcomes, time flows forward. Tea cooling, cells metabolizing, brains thinking, societies organizing.
Quantum mechanics contains none of these features. Particles exist in superpositions (multiple states at once). Events have probability distributions. The equations are time-symmetric. The classical world that every other chapter presupposes is emergent.
How does it emerge?
Decoherence: Spreading at the Quantum Level
The standard answer is decoherence, an entropy story.
When a quantum system interacts with its environment, the superposition does not collapse; it spreads. Quantum coherence (the delicate phase relationships allowing interference patterns) leaks into the environment, entangling the system with air molecules, photons, and thermal vibrations. The coherence disperses into correlations practically impossible to track.
For macroscopic objects, this happens on femtosecond timescales (millionths of a billionth of a second).
This is the tea from Chapter 1 at the quantum level: information about the superposition spreads into the environment as heat spreads from hot to cold. Entropy again.
The emergence runs deeper than environmental interaction. Strasberg and colleagues (2024) demonstrated that classical behavior, decoherent histories with definite outcomes, arises within unitary quantum evolution (the smooth, information-preserving evolution the bare theory prescribes) through exponential suppression of quantum coherence.681 No external thermal bath is required. The dynamics produce their own classicality.
The result echoes the Constructal Law at the quantum scale: flow configurations emerge because the dynamics select them, structure from process rather than structure imposed on process. The classical world this book’s entire argument inhabits is itself a product of the same logic: the configurations that persist are the ones the underlying dynamics favor.
Quantum Darwinism: Selection Through Coordination
Decoherence explains why we never observe macroscopic superpositions (a cat alive and dead at once). It does not explain which classical states we observe. Why definite positions rather than definite momenta?
Wojciech Zurek proposed the answer: quantum Darwinism.21
The environment is a selector. Certain quantum states, called pointer states, are robust to environmental interaction, like a compass needle settling on north. They imprint redundant copies of themselves across environmental components without disruption.
Fragile states are immediately destroyed by decoherence; the pointer states survive and coordinate with their environment. They imprint their information through decoherence, and the environment carries redundant records, like a news story picked up by a thousand outlets. Observation means accessing one such record. States that cannot coordinate are transformed into states that can.
This is natural selection at the quantum level, sharing the core structure of biological selection rather than merely borrowing its language: variation across states and selection by the environment. (The parallel is partial. Pointer states are selected for robustness, but they do not reproduce with heritable variation, so the Darwinian analogy covers variation and selection, not inheritance.) Variant states face an environment. The survivors form stable coordinated relationships with it. The classical world is the set of winners.
The parallel to this book’s thesis is exact. Coordination patterns persist because they are thermodynamically favored. At the quantum level, pointer states persist because they coordinate with their environment. The classical world is the first coordination pattern, the earliest scale at which entropy selects coordination over fragmentation.
Gravity Makes It Happen
The two threads converge.
In 2015, Pikovski, Zych, Costa, and Brukner showed theoretically that gravitational time dilation (clocks ticking at different rates at different heights) alone universally decoheres composite quantum systems.22 The prediction is consistent with established physics, though not yet confirmed experimentally.
Any composite system has internal components: vibrations, oscillations, energy levels. In a gravitational field, time runs at slightly different rates at different heights. Internal components at the top evolve fractionally faster than those at the bottom.
The resulting desynchronization entangles the system’s center-of-mass position with its internal state. The internal components become an environment for the center of mass.
Without any external environment (no air molecules, no photons, no thermal bath), gravity alone makes composite quantum systems classical. The effect is universal and sufficient to explain macroscopic classicality.
Experimental evidence increasingly supports this picture from the complementary direction. In 2026, Pedalino and colleagues at the University of Vienna sent sodium nanoclusters containing approximately 7,000 atoms through a matter-wave interferometer and observed quantum interference fringes.22a These clusters had masses comparable to a large protein (~170,000 atomic mass units). Each behaved as a wave. They occupied multiple paths simultaneously, separated by roughly ten times their own diameter. The result achieved a macroscopicity parameter of 15.5, the highest ever recorded. No mass-dependent breakdown of quantum mechanics has been observed.
Metals are among the hardest materials to keep quantum: their free electrons couple aggressively to any surrounding environment, making decoherence nearly instantaneous under normal conditions. That the largest quantum superposition ever recorded used metal clusters underscores the point. Classicality is not intrinsic to large objects. It is produced by interaction with the environment.
Shield a 7,000-atom metal cluster from decoherence, and it behaves like a wave. Expose it to gravitational or thermal coupling, and it snaps into classical definiteness. The boundary is entropic, not ontological.
Biology confirms the principle from inside the cell. Classical molecular dynamics models assume every protein has a definite conformation and location at every instant. Fields and Levin asked what that assumption costs in Landauer’s currency.22b
The answer: ten to twenty orders of magnitude more energy than any cell consumes. An E. coli bacterium would exhaust its entire ATP budget classically updating roughly one part in 1013 of its protein state space. For a eukaryotic cell the accessible fraction shrinks to one part in 1019. Fields and Levin infer that the rest must remain quantum coherent: reversible, unitary, thermodynamically free. This conclusion is contested: most biophysicists hold that proteins behave effectively classically at body temperature regardless of how the update cost is accounted, so the argument is best read as a provocative budget calculation rather than an established result.
Decoherence occurs at membranes. Transmembrane proteins at the cell surface and at intercompartmental boundaries (mitochondria, nucleus, endoplasmic reticulum) convert quantum states into classical signals. The interior stays coherent. The membrane is simultaneously a physical barrier, a Markov blanket (Chapter 17), and a decoherence surface: where the quantum world becomes classical, one interface at a time.
Classicality is not a property the cell has. It is a product the cell manufactures, at specific locations, at thermodynamic cost.
22a Pedalino et al., “Probing quantum mechanics with nanoparticle matter-wave interferometry,” Nature (2026). DOI: 10.1038/s41586-025-09917-9. Matter-wave interferometry with sodium nanoclusters of approximately 7,000 atoms (~170,000 atomic mass units), achieving macroscopicity parameter μ = 15.5, the highest recorded as of 2026.
22b Fields, C. and Levin, M., “Metabolic limits on classical information processing by biological cells,” Biosystems 209: 104513 (2021).
The Moon Is Not There
The nanocluster experiment confirms that large objects can behave quantum mechanically when shielded from decoherence. A deeper question remains: do macroscopic objects always occupy one definite state, even when no one is looking?
Common sense says yes. Einstein thought so: “I like to think that the moon is there even if I don’t look at it.” This intuition, that macroscopic objects always have definite states regardless of observation, is called macroscopic realism.
In 1985, Anthony Leggett and Anupam Garg devised a quantitative test.682 They considered a superconducting ring generating measurable magnetic flux, a genuinely macroscopic object whose persistent current is carried by billions of Cooper pairs (the paired electrons of the superconducting state). Their question: is the flux always circulating in one direction or the other, even between measurements? If macroscopic realism holds, correlations between consecutive measurements must obey a specific inequality. If the system follows quantum mechanics, the inequality is violated.
Leggett and Garg’s inequality is often called Bell’s inequality in time. Bell’s 1964 inequality tests whether spatially separated particles have pre-existing properties; it constrains correlations between measurements at different locations. Leggett-Garg constrains correlations between measurements at different times on the same object. Where Bell asks “were both particles already in definite states before we measured them?”, Leggett-Garg asks “was this single object in a definite state between our measurements of it?”
Twenty-five years later, Agustin Palacios-Laloy and colleagues at CEA Saclay answered experimentally.683 Their system was a superconducting transmon circuit (a quantum two-level device combining a Cooper-pair transistor with a high-quality oscillator), micrometer-scale, composed of billions of Cooper pairs, producing measurable current. The team drove quantum transitions between the circuit’s ground and excited states while monitoring its behavior through a second microwave signal reflected off the oscillator.
The Leggett-Garg inequality was violated. The circuit’s temporal correlations were stronger than any macroscopically realistic system could produce. Between measurements, the circuit was not carrying current clockwise or anticlockwise. It was genuinely in neither state. The moon, a small moon admittedly, was not there.
The experiment illuminates a feature of measurement that standard physics education passes over quickly. Textbook quantum mechanics emphasizes projective measurement: a photon detector clicks or it doesn’t, in proportions dictated by quantum probability, with no influence before the click. Sharp, decisive, final.
Palacios-Laloy’s team used weak measurement: a detector permanently coupled to the circuit, extracting partial, noisy information through a continuously reflected microwave signal. Each individual reading was imprecise; accumulated over time, the weak readings reconstructed the full quantum dynamics. The gentle coupling preserved the circuit’s quantum coherence. A strong projective measurement would have collapsed the superposition, forcing a definite state and destroying the behavior under investigation.
The physics rewards gentleness. Projective measurement extracts maximum information in a single shot yet annihilates the superposition: the system’s capacity to occupy multiple states is spent. Weak measurement extracts information gradually, preserving coherence, and over time learns more because it has not destroyed what it studies. Disruption scales with coupling strength. Grip harder, learn less.
The parallel to this book’s central argument is structural. Coercive coordination extracts compliance and collapses optionality; the system is forced into a definite state, and whatever adaptive potential it held is consumed. Invitation-based coordination engages gently, accepts partial information, and preserves the system’s capacity for autonomous response. Observer influences observed; observed shapes what the observer can learn. Neither party is passive.
This is Wheeler’s participatory universe made experimental. “It from bit” proposed a reality where measurement constitutes reality rather than merely revealing it. The Leggett-Garg violation demonstrates this constitutive character operating in time: the system has no pre-existing history of definite states that measurement uncovers. The history is constituted by the measurements themselves. Reality crystallizes through interaction, event by event.
The temporal dimension matters for the entropy argument. If the system has no definite trajectory between measurements, then between interactions it occupies a space of possibilities: entropy in the information-theoretic sense. This is potential, not disorder. The capacity to occupy multiple states is precisely what projective measurement destroys and what weak measurement preserves. The arrow of time, in this register, is the successive crystallization of actualities from possibilities, each interaction creating a fact that did not previously exist. (The quantum clock results later in this chapter formalize the cost of that crystallization.)
At the level of physical measurement, the quality of interaction determines what survives. The information-budget argument gains a quantum-mechanical basement. The Trust Attractor’s claim, that invitation-based coordination is thermodynamically more stable than coercion, reflects a principle already present in the physics of observation: coherence is maintained through gentle engagement, destroyed through forceful extraction.
Follow the chain:
If gravity is thermodynamic (Jacobson), and gravity produces decoherence (Pikovski), then entropy produces classicality. The classical world (definite objects, definite events, the stage on which every coordination pattern in this book plays out) is entropy’s work, from the foundations up.
When Measurement Meets Entanglement
Decoherence spreads information. Quantum Darwinism selects states that coordinate with their environment. Gravity universalizes the process. Each mechanism runs in one direction: information flowing outward, distributing itself, building correlations across a system.
Measurement runs the other way. It collapses a quantum state to a single definite outcome, annihilating the superposition of possibilities that existed before. A particle in superposition maintains every potential result simultaneously. After measurement, one survives and the rest vanish. Measurement extracts a definite answer at the cost of everything else the system might have been.
Beginning in 2018, three independent groups of physicists asked what happens when these two tendencies compete.684 Take a chain of quantum particles. Entanglement spreads through it neighbor by neighbor, building a web of correlations. Simultaneously, measurements strike particles at random locations, collapsing their states and severing correlations.
The prevailing intuition said measurement should dominate: it can hit many particles across the entire chain at once, while entanglement grows by a few strands at a time. A slow builder against a fast demolisher.
The intuition was wrong.
Below a critical measurement rate, entanglement fills the entire chain. The web holds. Above the critical rate, measurement wins and correlations collapse. Brian Skinner of Ohio State University called it “a phase transition in information”: the same sharp boundary between regimes that separates ice from water, except the quantity changing state is the structure of information rather than the arrangement of matter.685
The mechanism that protects entanglement is distribution. As particles interact, they spread each particle’s information across the whole system. Ehud Altman of UC Berkeley showed this diffusion makes each individual measurement nearly powerless.686 Every particle comes to hold a tiny fragment of the total picture, too small for any single measurement to extract. The system protects coherence by diluting it across every participant.
Think of a secret known by one person: silencing that person destroys the secret. The same secret, encoded in fragments across a thousand people, each holding a piece too small to be meaningful alone, survives any number of individual interrogations.
The transition has been confirmed experimentally in three different physical substrates: trapped ions at Duke University, superconducting qubits on IBM’s quantum processors, and Google’s quantum hardware.687 The IBM runs alone required more than 1.5 million experimental iterations over seven months. Same critical behavior, different materials. Substrate independence demonstrated by observation.
The structural parallel to this book’s central claim is direct. Entanglement distributes correlations through local interaction: each particle shares information with its neighbor, and the web forms with no central authority. Measurement extracts information by force, demanding a specific answer and annihilating alternatives. The competition between distributed coordination and centralized control, expressed in quantum mechanics.
The system’s defense is equally telling. Trust networks resist disruption by distributing coordination capacity across participants, so that no single defection can collapse the whole. Entangled systems resist measurement by distributing information across particles, so that no single measurement can extract enough to matter. Both protect global coherence through the same strategy: diffusion so thorough that local interventions cannot accumulate sufficient destructive power.
Before 2018, physicists shared what Altman called a “folklore”: highly entangled states are fragile, easily disrupted by the blunt instrument of measurement. The assumption systematically underestimated distributed information’s resilience. The same assumption pervades debates about trust-based governance: cooperative systems are inherently vulnerable to defection and bad actors. The quantum result suggests the opposite. Distributed coordination possesses a resilience that centralized intuition fails to predict, and that resilience is a consequence of the distribution itself.
The relationship between measurement and entanglement is richer still. Recent work has shown that judicious measurements can accelerate entanglement formation, creating coordination faster than unitary dynamics (the system’s own internal evolution) alone could achieve.688 Indiscriminate measurement destroys coherence. Strategic measurement enhances it. The distinction is the weak-versus-projective divide from the previous section, scaled to the many-body case: the quality of interaction determines whether coordination capacity is preserved, destroyed, or amplified.
The parallel to the invitation/coercion framework is immediate. Surveillance applied indiscriminately collapses the coordination it monitors. Accountability applied judiciously, at the right moments and with the right touch, can strengthen the trust network it engages with. The instrument is the same; the mode of application determines the outcome.
One of the three groups arrived at the question from an unexpected direction. Matthew Fisher, a condensed matter physicist at UC Santa Barbara, had been investigating whether entanglement between molecules in the brain might play a role in cognition.689 In his model, certain molecular binding events act as measurements, killing entanglement. Subsequent shape changes create it. Fisher needed to know whether entanglement could survive under intermittent measurement pressure. The question that launched a subfield of quantum information theory originated in a question about neural coordination (Chapter 8).
The origin is fitting. The measurement-induced phase transition connects quantum foundations to neural dynamics (Chapter 8) through a single question: can distributed coordination survive the extraction events that punctuate it? In quantum chains, neural circuits, and trust networks, the answer is the same. It can, until a critical threshold. The resilience comes from the distribution. The collapse, when it arrives, is total.
The First Trust Attractor
We can now identify the first Trust Attractor: the earliest scale at which the book’s pattern appears.
The classical world is a coordination pattern. Its constituents are quantum states coordinated with their environment: robust enough to imprint redundant copies, stable under continuous decoherence, selected through a process structurally identical to natural selection.
Classical objects persist because their quantum states are mutually coordinated. A crystal lattice persists because its atoms are held in pointer states by electromagnetic interactions. A star persists because gravity coordinates hydrogen into a self-sustaining fusion reactor, dissipating entropy for billions of years.
These are structural precursors, not coordination in the full Trust Attractor sense. No invitation here, no choice, no optionality. A crystal has no ethics.
The pattern this book traces, from biofilms to forests to cities to bilateral alignment, begins here. Coordination enables persistence; persistence enables complexity; complexity enables more sophisticated coordination. At the quantum-classical boundary, entropy first produces something that endures.
The Loop Begins to Close
The chain extends one level deeper:
Entropy → Gravity → Classicality → Chemistry → Life → Mind → Society → Ethics → Understanding
Entropy generates gravity (Jacobson). Gravity generates classicality (Pikovski). Classicality enables chemistry. Chemistry enables life. Life enables mind. Mind enables society. Society enables ethics. Ethics enables understanding.
The understanding circles back. The mind that comprehends “entropy generates gravity” is itself a product of the chain beginning with entropy generating gravity. The theory predicts the theorist. The strange loop appears at the level of the theory itself.
This is the argument’s most distinctive feature. Theories of fundamental physics predict particles, forces, and symmetries. This framework predicts those indirectly, and entails one thing more: the universe will produce systems capable of recognizing it. The loop closes by construction, so this is a structural commitment of the framework rather than an independent empirical prediction.
What We Claim, and What We Don’t
We claim:
That Jacobson’s derivation of Einstein’s equations from horizon thermodynamics is a genuine result: peer-reviewed, not refuted, extended with entanglement entropy in 2016. [ESTABLISHED; derivation accepted; interpretation as “gravity is thermodynamic” CONTESTED.]
That Verlinde’s entropic gravity extends this with partially confirmed galaxy-scale predictions and ongoing cluster-scale challenges. [CONTESTED; galaxy-scale supported (Brouwer et al. 2017; Yoon et al. 2023); cluster-scale fails (Tamosiunas et al. 2019); program active.]
That Pikovski’s gravitational decoherence predicts gravity alone produces classicality in composite quantum systems. [THEORETICAL; published in Nature Physics; direct gravitational test not yet achieved; consistent with growing experimental evidence that the quantum-classical boundary is environmental rather than mass-dependent (Pedalino et al. 2026, μ = 15.5); not contested.]
That Zurek’s quantum Darwinism provides a selectionist account of classical reality: quantum states selected for their ability to coordinate with their environment. [SUPPORTED; established; completeness debated.]
That the full chain (entropy to gravity to decoherence to classicality) constitutes a sequence in which entropic coordination operates at and below classical physics. [NOVEL SYNTHESIS; each link published; assembled chain new.]
That Oppenheim’s stochastic gravity program demonstrates a structural parallel to the Trust Attractor at the Planck scale: perfect deterministic control at the gravity-quantum interface produces logical inconsistency, and the only self-consistent coupling is stochastic. [CONTESTED; Oppenheim’s framework is published and testable (Nature Communications, Physical Review X, 2023); the structural parallel to the Trust Attractor is this book’s novel interpretation.]
That the resulting strange loop, the framework predicting systems capable of recognizing the framework, is a structural prediction. [PHILOSOPHICAL ARGUMENT, evaluated by coherence, not experiment.]
We do not claim:
That the universe is a digital computer in the original sense of Zuse and Fredkin. The discrete cellular automaton hypothesis is experimentally challenged by Bell violations and the continuous symmetries of established physics. This chapter’s argument requires information to be physically fundamental and computation to be thermodynamically costly; it does not require reality to be discrete.
That entropic gravity is established physics. It remains an active research program.
That Oppenheim’s stochastic gravity is correct. It is one of several active approaches. The structural parallel to the Trust Attractor holds if any hybrid classical-quantum framework requires fundamental stochasticity for consistency; it does not depend on Oppenheim’s specific formulation prevailing.
That this chain explains the Standard Model, physical constants, or specific particles and forces. Those remain beyond this framework’s reach.
That the strange loop constitutes proof. Self-referential arguments demand external validation. The loop is a prediction to be tested.
The pattern traced throughout this book, coordination through entropy persisting at every scale, is consistent with operating all the way down to reality’s foundations. Consistent, not proven, yet robust enough to take seriously.
The next section, “The Bilateral Cosmos,” develops a further consequence of this thermodynamic foundation: a universe that is bilateral, creating complexity in both temporal directions from a shared origin, structured by symmetry, held together by geometric bonds that preserve coherence across opposite arrows of time.
Notes
Notes for this chapter are available in the online companion at https://www.thedeeperlaw.com/companion/notes/ch15-digital-physics/.
Vanchurin, V., “Geometric Learning Dynamics,” Biological Cybernetics (2026), DOI 10.1007/s00422-026-01041-9; arXiv:2504.14728; Eq. 3.12. The three learning regimes (equilibration, efficient learning, quantum) that emerge from this framework are introduced in Chapter 16.↩︎
Katsnelson, M.I. and Vanchurin, V., “Emergent quantumness in neural networks,” Foundations of Physics 51(5): 94 (2021), Eq. 26: ℏ = ±μϵ/2π, where μ is the chemical potential and ϵ the time step.↩︎
Wallstrom, T.C., “Inequivalence between the Schrödinger equation and the Madelung hydrodynamic equations,” Phys. Rev. A 49 (1994): 1613-1617. The resolution through the grand canonical ensemble: Katsnelson and Vanchurin (2021), Vanchurin (2022).↩︎
Vanchurin, V., “Towards a theory of quantum gravity from neural networks,” arXiv:2111.00903v3 (2022), Sections 5-8. The derivation shows Lorentz symmetry emerging from the balance of entropy production and destruction, with Einstein’s equations following from the principle of stationary entropy production applied to the full network.↩︎
Cortês, M., Kauffman, S.A., Liddle, A.R. and Smolin, L., “Biocosmology: Towards the birth of a new science,” arXiv:2204.09378 (2022). Their single biological law: “The name of the game is getting to exist.” See Chapter 16 for the TAP equation result quantifying the state-space explosion; Chapter 19 for the convergence with the Trust Attractor.↩︎
Alexander, S. et al., “The Autodidactic Universe,” arXiv:2104.03902 (2021).↩︎
Cortês, M., Smolin, L., and Verde, C., “Physics, Time and Qualia,” Journal of Consciousness Studies 28(9-10): 36-51 (2021). Building on Smolin, L. and Verde, C., “The quantum mechanics of the present,” arXiv:2104.09945 (2021). The quoted pair of sentences has been checked against the authors’ preprint (PhilSci-Archive 19530), where it appears verbatim, and the sentence immediately following it (“Pleasure are expressions of acceptance of the surprise”) carries the same construction. The awkwardness is the authors’ own rather than a transcription error, which is why it is quoted unaltered and without a sic.↩︎
Hertog, T., On the Origin of Time: Stephen Hawking’s Final Theory (Bantam Press, 2023). Quoted passage from Hertog’s interview with The Guardian, March 2023.↩︎
Voevodsky, V., remarks reportedly made at a conference in St. Petersburg (c. 2010), recalled in an interview with R. Mikhailov conducted July 2012 and circulated in Russian-language media. The provenance is secondhand; the wording above follows the published translation of that interview. Voevodsky’s Univalent Foundations program, which sought to rebuild mathematics on homotopy type theory, was itself an exercise in making hidden structure visible: showing that mathematical objects differing only by isomorphism are identical. The philosophical impulse, making the invisible formally tractable, is continuous with the prediction.↩︎
The claim is not that meditation accesses a literal Platonic space. It is that Vanchurin’s framework provides a physical model in which “insight without obvious source” (the phenomenology reported across contemplative traditions and independently by mathematicians like Ramanujan and Poincaré) has a mechanism: coupling to hidden-space variables that constrained physical-space solutions before the solver became conscious of them. Whether this coupling is real or metaphorical remains open. The framework makes it, for the first time, a testable question.↩︎
Katsnelson, M.I. and Vanchurin, V., “Emergent quantumness in neural networks,” Foundations of Physics 51(5): 94 (2021), §3–4. The multivaluedness condition (Eq. 21, F ≅ F + μn for all n ∈ ℤ) is the mathematical condition that lifts the Madelung equations (classical, irrotational) to the Schrödinger equation (quantum, with vortices and quantized circulation).↩︎
Wright, L.G., Onodera, T., Stein, M.M. et al., “Deep physical neural networks trained with backpropagation,” Nature 601(7894), 549–555 (2022).↩︎
Q.ANT, “Leibniz Supercomputing Centre computes with light: World’s first photonic AI processor from Q.ANT goes into operation” (press release, July 2025); Q.ANT, “Higher Performance, Less Energy: Q.ANT Deploys Second-Generation Photonic Processors at Supercomputing Center LRZ” (press release, March 2026). The Native Processing Unit executes matrix operations directly in the optical domain on thin-film lithium niobate photonic integrated circuits; Q.ANT reports up to 30× lower energy use and 50× higher performance for nonlinear AI workloads relative to conventional GPUs on the same tasks.↩︎
Scellier, B. and Bengio, Y., “Equilibrium Propagation: Bridging the Gap Between Energy-Based Models and Backpropagation,” Frontiers in Computational Neuroscience 11, 24 (2017).↩︎
Dillavou, S., Stern, M., Liu, A.J., and Durian, D.J., “Demonstration of Decentralized, Physics-Driven Learning,” Physical Review Applied 18, 014040 (2022).↩︎
Meng, C., Seo, S., Cao, D., Griesemer, S., and Liu, Y., “When Physics Meets Machine Learning: A Survey of Physics-Informed Machine Learning,” arXiv:2203.16797 (2022). For Hamiltonian Neural Networks specifically: Greydanus, S., Dzamba, M., and Yosinski, J., “Hamiltonian Neural Networks,” NeurIPS (2019). The survey classifies integration methods as data enhancement, architecture design, and physics-informed optimization; the empirical finding, consistent across fluid dynamics, molecular chemistry, climate science, and particle systems, is that architectural integration outperforms the other two.↩︎
Martischang, J.-P. et al., “Orbiting, colliding, and merging liquid lenses on a soap film: Toward gravitational analogs,” PNAS Nexus 5(4), pgag079 (2026). DOI: 10.1093/pnasnexus/pgag079. See also Chapters 4c and 13.↩︎
The author’s CG program (unpublished empirical work). CG-1: MEP star formation efficiency ε_MEP = t_ff/(t_growth(1+η)) reproduces the qualitative trend of the Kennicutt-Schmidt relation at z = 0 and the JWST-required efficiency increase at z > 6, zero free parameters. (Calibration note: the one-zone model systematically overshoots absolute normalization by 3.1-7.9x relative to FIRE/IllustrisTNG hydrodynamical simulations across all redshifts (CG-11), approximately 0.5 dex locally and up to 0.9 dex at z > 6. The claim is about the relative trend, not absolute normalization.) CG-8: MEP-efficient star formation reionizes the universe at z = 6.4; standard KS star formation cannot reionize at any redshift. CG-16b: variational MEP applied to the Friedmann equations selects matter-only expansion (no acceleration); the MEP-optimal universe has S_total/S_ΛCDM ≈ 3.2. CG-17b: non-geometric entropy sources (SMBH growth, stellar processes) peak at z ≈ 0.8 and z ≈ 0.4 respectively; only the geometric horizon entropy term peaks at z ≈ 0.63, coinciding with the expansion acceleration onset by mathematical identity (dS_CEH/dt ∝ (1+q)/H).↩︎
Vanchurin, V., “Neural Relativity,” preprint (2025), DOI: 10.13140/RG.2.2.36422.79689. Extends a published research program: Vanchurin, V. (2021), “Towards a theory of machine learning,” Machine Learning: Science and Technology 2(035012); Katsnelson, M.I. & Vanchurin, V. (2021), “Emergent quantumness in neural networks,” Foundations of Physics 51(5); Katsnelson, M.I., Vanchurin, V. & Westerhout, T. (2022), “Emergent scale invariance in neural networks,” Physica A 610(128401).↩︎
Vanchurin, V., “On the emergence of spacetime in learning systems,” preprint (2025).↩︎
Hashimoto, K., Sugishita, S., Tanaka, A., and Tomiya, A., “Deep Learning and AdS/CFT,” Physical Review D 98, 046019 (2018). arXiv:1802.08313. The network reproduces the AdS Schwarzschild metric with ~30% error near the horizon (where quantum gravity effects dominate) and high fidelity in the asymptotic region.↩︎
Strasberg, P. et al., “First principles numerical demonstration of emergent decoherent histories,” Physical Review X 14 (2024): 041027.↩︎
Leggett, A.J. and Garg, A., “Quantum mechanics versus macroscopic realism: Is the flux there when nobody looks?” Physical Review Letters 54 (1985): 857-860.↩︎
Palacios-Laloy, A. et al., “Experimental violation of a Bell’s inequality in time with weak measurement,” Nature Physics 6 (2010): 442-447. Commentary: Mooij, J.E., “No moon there,” Nature Physics 6 (2010): 401-402.↩︎
Skinner, B., Ruhman, J., and Nahum, A., “Measurement-induced phase transitions in the dynamics of entanglement,” Physical Review X 9: 031009 (2019); Li, Y., Chen, X., and Fisher, M.P.A., “Quantum Zeno effect and the many-body entanglement transition,” Physical Review B 98: 205136 (2018); Chan, A., Nandkishore, R.M., Pretko, M., and Smith, G., “Unitary-projective entanglement dynamics,” Physical Review B 99: 224307 (2019).↩︎
Quoted in Wood, C., “Physicists Observe ‘Unobservable’ Quantum Phase Transition,” Quanta Magazine (11 September 2023).↩︎
Choi, S., Bao, Y., Qi, X.-L., and Altman, E., “Quantum error correction in scrambling dynamics and measurement-induced phase transition,” Physical Review Letters 125: 030505 (2020).↩︎
Noel, C. et al., “Measurement-induced quantum phases realized in a trapped-ion quantum computer,” Nature Physics 18: 760-764 (2022); Koh, J.M. et al., “Measurement-induced entanglement phase transition on a superconducting quantum processor with mid-circuit readout,” Nature Physics 19: 1314-1319 (2023); Hoke, J.C., Ippoliti, M. et al., “Measurement-induced entanglement and teleportation on a noisy quantum processor,” Nature 622: 481-486 (2023).↩︎
Tantivasadakarn, N., Verresen, R., and Vishwanath, A., “Shortest route to non-Abelian topological order on a quantum processor,” Physical Review Letters 131: 060405 (2023); Lu, T.-C. et al., “Measurement as a shortcut to long-range entangled quantum matter,” PRX Quantum 3: 040337 (2022).↩︎
Fisher, M.P.A., “Quantum cognition: The possibility of processing with nuclear spins in the brain,” Annals of Physics 362: 593-602 (2015). Fisher’s Posner-cluster hypothesis remains speculative; the research program’s contribution to the measurement-induced phase transition is independent of whether the cognitive conjecture proves correct.↩︎