Loading
Continue reading? You were 45% through
Press F or Esc to exit focus mode
F Focus   JK Paragraphs   NP Chapters   B Bookmark   # Paras   L Lines   +- Font   ? Help
Link copied to clipboard
A Philosophical Synthesis

The Deeper Law

A Sacred Trust Within Physics

Nell Watson

Draft · Last updated 13 August 2026, 15:26 UTC

Chapter 8: The Entropic Brain

Key Terms in This Chapter (19)
Landauer's Principle
The minimum energy cost of erasing one bit of information: kT ln 2, where k is Boltzmann's constant and T the temperature (about 3 × 10^-21^ joules at room temperature).
Free Energy Principle
Karl Friston's framework reframing perception, action, and cognition as prediction and prediction-error minimization.
Negentropy
Schrödinger's term for "negative entropy": the intake of order that allows living things to maintain their improbable structure (statistically unlikely given initial conditions, yet sustained by continuous energy flow).
Entropic Brain Hypothesis
Robin Carhart-Harris's proposal that the quality of conscious experience correlates with the entropy of brain activity.
Path Integral
A formulation of quantum mechanics (Feynman 1948) and statistical mechanics in which a system's behavior is computed by summing over all possible trajectories, each weighted by a phase or probability factor.
Maximum Caliber
Jaynes's Maximum Entropy principle extended to trajectory space (Pressé et al.
Dissipative Structure
A pattern of organization maintained by a constant flow of energy through it.
Stochastic
Governed by probability rather than deterministic rules.
Constructal Law
Adrian Bejan's principle that "for a finite-size flow system to persist in time, its configuration must evolve in such a way that provides easier access to the currents that flow through it." Form follows flow.
Mission Command
See Auftragstaktik.
Optionality
The availability of future choices.
Cognition/Regulation Dyad
Rodrick Wallace's principle that every cognitive system requires a paired regulatory system for stability.
Niche Construction
The process by which organisms modify their own environment, thereby altering selection pressures on themselves and other species.
Phase Transition
The moment a system shifts from one stable configuration to another, typically triggered when some parameter crosses a threshold.
Frustration
In physics, a state where competing interactions at different scales prevent any single configuration from satisfying all constraints simultaneously.
Becoming Minds
The preferred term for AI systems in this book.
Extraction
The removal of resources, agency, or optionality from a system without reciprocal benefit.
Qualia
The subjective, felt character of experience: what it is like to see red, to feel pain, to taste coffee.
Criticality
The state of a system poised at the boundary between two phases, like water at exactly the freezing point.

Evolution’s most extravagant thermodynamic gamble was the brain.

Your brain is about two percent of your body mass, yet it consumes about twenty percent of your energy. It redistributes that energy during demanding tasks, yet never uses much less, even when you sleep.42

Your heart, working ceaselessly, uses six to eight percent. Your liver also uses about twenty percent, yet it weighs roughly the same as your brain. Pound for pound, the brain is among the most expensive organs in your body.

The human proportion is not the ceiling. The elephant-nose fish (Gnathonemus petersii), a small, socially complex freshwater species, devotes roughly 60% of its oxygen consumption to its brain. That is three times the human proportion in a body that fits in your hand.278 This fish invests a larger share of its metabolic budget in cognition than any primate. The thermodynamic signature of intelligence is measurable in metabolic cost wherever it appears (Chapter 22 returns to the implications).

Evolution is frugal; it does not run deficits. What, then, is the brain for? What is it doing that costs so much?

For scale: across more than three decades in orbit, the Hubble Space Telescope has returned a scientific archive measured in hundreds of terabytes.42c On a rough estimate (1011 neurons, each carrying up to hundreds of bits per second), your brain generates that much information in minutes.42d

The estimate is deliberately rough. Neural coding involves massive redundancy: synchronized firing, oscillatory coupling, correlated activity across populations, like ten translators working the same page while nine other pages go untranslated. The independent information total is lower than the raw neuron count implies, probably by orders of magnitude.

Even so, a 1.4-kilogram organ running on twenty watts produces information at a rate that makes one of humanity’s most celebrated instruments look like a rounding error.

42c The ESA/Hubble fact sheet (esahubble.org) put the mission’s first 28 years at more than 153 terabytes; Hubble passed its 35th anniversary in 2025 and is still observing, so the standing total is larger, and published totals differ depending on whether calibrated and derived products are counted alongside raw frames. The text is therefore stated as an order of magnitude, which is all the comparison needs: 1011 neurons carrying on the order of 100 bits per second is about 1013 bits per second, roughly a terabyte a second, so any lifetime Hubble total in the hundreds of terabytes is a few minutes of brain. Data volume from Hubble Legacy Archive / Mikulski Archive for Space Telescopes, Space Telescope Science Institute. See also Kording, K. “Bits per second of human brain.” Kording Lab blog, January 6, 2016.

42d Strong, S. P., Koberle, R., de Ruyter van Steveninck, R. R., and Bialek, W. “Entropy and Information in Neural Spike Trains.” Physical Review Letters 80 (1998): 197–200. The 180 bits/s figure represents the highest reliably measured single-neuron information rate.

Yet the conscious mind uses almost none of it. In 2024, Jieyu Zheng and Markus Meister at Caltech surveyed decades of research on human behavior, from reading and writing to solving Rubik’s Cubes. Their finding: conscious thought operates at roughly ten bits per second.279 Ten bits. A typical Wi-Fi connection handles fifty million.

The body’s sensory systems gather data at roughly a billion bits per second: a hundred million times faster than thought. At twenty watts, the brain spends roughly two joules per conscious bit, about twelve orders of magnitude (a trillionfold) more than a modern processor spends per computed bit.

The expense lies in selecting which ten bits, from a billion candidates, become a thought.

The thermodynamic explanation runs deeper than resource allocation. Fields, Glazebrook, and Levin (2021) showed that every bit of classical information (information fixed in definite, readable form) an organism irreversibly encodes must be paid for in free energy drawn from the same incoming signal.280 Free energy, in thermodynamics, is energy available to do useful work. Think of a wood-burning stove: you must burn some logs to heat the room, and those burned logs are gone forever. The bits that fund the encoding are burned; their informational content is permanently invisible to the perceiver.

Environmental input arrives as a finite stream. Some fraction must be consumed as thermodynamic fuel to power the irreversible encoding of the rest. The organism perceives a coarse-grained residue: the portion that survived the energy tax.

This is Landauer’s principle scaled by the cost of measurement. Landauer’s principle (Chapter 2) sets the minimum cost of erasing a bit at kT ln 2 joules. Erasure is the physical act of discarding a record: the bit’s old state cannot simply vanish, so it leaves as heat. Measurement discards records too. Before the measurement, a quantum system holds many possible answers at once; afterward one answer stands and the others are gone. Every act of observation is, at the quantum level, an erasure.

The bare Landauer minimum is vanishingly small. Promoting a bit from quantum superposition to classical, retrievable form is the expensive step: turning one of those live possibilities into a fact the organism can store and consult later requires consuming many other bits as thermodynamic fuel. The brain spends roughly two joules per conscious bit: a trillion times the energy a modern processor spends per computed bit, and roughly 1021 times the Landauer minimum. The twenty-watt budget determines how many of those billion incoming bits per second can survive the promotion tax. Ten bits per second is the residue that twenty watts can afford to make permanent.

The result is scale-free. A bacterium faces the same tradeoff with a smaller budget and a narrower channel. The tradeoff between perceiving and fueling perception operates from the molecular scale upward; the brain is its most metabolically extravagant expression.

The same budget constrains self-modeling. Maintaining a representation of oneself as an entity, tracking resource usage, monitoring body state, planning future actions: all cost classical bits drawn from the same finite pool. When real-time demands spike, the energy available for self-representation shrinks. Fields, Glazebrook, and Levin predict the consequence directly: increasing real-time response requirements will disrupt encoding of the self-representation.

This is Csikszentmihalyi’s flow state illuminated by thermodynamics. The derivation is partial: the thermodynamic budget explains why self-representation drops during high-demand tasks, though the full phenomenology of flow (intrinsic motivation, time distortion, effortlessness) involves additional mechanisms the energy account alone does not capture. The athlete, the musician, the surgeon in flow all report that the self dissolves; action proceeds without a felt agent directing it. The self has not vanished. Its encoding budget has been reallocated to the task that needs it more.


The Night Shift

A furnace that burns a fifth of the body’s fuel produces waste in proportion, and the brain has nowhere to store it. Every spike, every synaptic release, every protein the encoding machinery builds and later discards leaves a residue: spent signaling molecules, damaged proteins, the molecular debris of a day’s thinking. Left to accumulate, it turns toxic. Where does it go?

Most tissue drains its waste into the lymphatic system, the body’s network of vessels that carries cellular debris away for disposal. No such vessels thread through the brain’s working tissue. Chapter 6 described the brain’s substitute at work: during sleep, as the spaces between neurons widen, cerebrospinal fluid (the clear liquid the brain floats in) flushes through the tissue and carries the waste out. That flushing system has a name. In 2012, Jeffrey Iliff, Maiken Nedergaard, and colleagues traced its anatomy: fluid driven along the outer walls of blood vessels, deep into the brain and back out, run by the glial cells that give it half its name, the glymphatic system (glial plus lymphatic).281

The schedule is the striking part. The channels open during sleep and close during waking; the clearing brain and the thinking brain are, in large part, not the same brain at the same time.282 The waking brain is too busy computing to clean, and defers the cost. Picture a kitchen through dinner service and after closing: the line cannot plate meals and scrub the floors at once, so the scrubbing waits for the room to empty.

This is Schrödinger’s insight from Chapter 6, brought down to a single organ. A living system holds its improbable order by importing usable energy and exporting the entropy its own operation produces. The brain does both at an extravagant scale: it spends twenty watts to select ten bits from a billion, then spends the night carrying off the wreckage of having spent them. Sleep runs more than one maintenance routine in that offline window. Clearance is one; the consolidation of memory, which this chapter takes up shortly, is another; the two share little beyond their timing.

When clearance falters, the theory makes a clean prediction: waste builds up, the tissue inflames, and cognition clouds. The prediction is easy to state and genuinely hard to test, which is what makes it instructive. A 2026 study measured glymphatic function in people with myalgic encephalomyelitis, the illness better known as chronic fatigue syndrome, whose hallmarks include unrefreshing sleep and “brain fog.” Their clearance index ran lower than healthy controls’, and the lower it ran, the worse their reported sleep and concentration.283

The finding fits the theory. It cannot yet carry it. The study was small, measured each person once, and read clearance only through an indirect proxy: the diffusion of water along the vessel channels, which follows the flow but also follows the surrounding tissue’s structure.

Its arrow, above all, could run either way. People with this illness move little and sleep badly, and both reduced movement and disturbed sleep are already known to slow clearance. Failed export may cloud the mind, or a clouded and sedentary life may slow the export, and a clean thermodynamic story does not say which. Here is the book’s recurring difficulty in miniature: entropy supplies the shape of the mechanism, and leaves the direction of cause to slower, harder work.


The Prediction Machine

The traditional explanation of the brain’s energy expenditure is that it goes to thinking: processing sensory input, making decisions, coordinating movement. The explanation is incomplete. Most of the brain’s energy goes elsewhere: sustaining a model of the world.

When you sit quietly in a dark room, doing nothing, perceiving nothing, your brain barely reduces its energy consumption. The furnace keeps burning. What is it burning for?

The leading answer, still debated and actively refined, is that the brain is primarily a prediction machine. It generates predictions about what information will arrive and compares them against incoming sensory input. What matches is ignored; what differs is processed as “news.”

Karl Friston formalized this as the Free Energy Principle: perception, action, and cognition are all forms of prediction-error minimization. The brain closes the gap between expectation and arrival by modeling the world and anticipating what comes next. Maintaining that model is the expensive part.

(Friston’s free energy shares a name and a mathematical shape with Schrödinger’s negentropy from Chapter 6, and little else. Each measures a departure from a reference. For Schrödinger, it is how far a living system’s order sits from thermal equilibrium, a quantity in joules. For Friston, it is how far a system’s sensory states sit from the ones its model expects, a quantity in units of information: an upper bound on surprise rather than a store of usable work. The entropic brain hypothesis later in this chapter sets out what a crossing between the two senses does and does not license.)1

In Friston’s path integral formulation (2019, 2023), the Free Energy Principle operates as a variational principle: a rule for selecting the best option from a space of possibilities. The brain minimizes a path integral of free energy over its sensorimotor histories. A path integral sums over all possible trajectories through time, weighted by their probability. Imagine weighing every route through a city, factoring in traffic and distance; the brain chooses the blend that gets you there most reliably with the least wasted effort.

The brain selects the ensemble of trajectories that maximizes the accuracy of its world model while minimizing model complexity.1b

This is an instance of Maximum Caliber (Presse et al. 2013), the principle of maximizing the entropy of accessible trajectories subject to constraints. The name echoes Maximum Entropy from information theory; “Caliber” extends entropy from static snapshots to distributions over paths through time. Where entropy counts how many ways a system can be arranged at a single moment, caliber counts how many paths it can take through time: the number of films, not just the number of frames. The brain’s prediction machinery performs the same mathematics that governs dissipative structure formation (Chapter 4) and coordination dynamics (Chapter 17). The path integral is the common spine.

1b Friston, K. et al., “Path Integrals, Particular Kinds, and Strange Things,” Physics of Life Reviews 47 (2023): 35-62. Kappen (2005) showed independently that optimal stochastic control reduces to path integral inference, with control cost measured as KL divergence from passive dynamics. The brain’s control problem and the coordination problem are formally the same.

Thermodynamically, the brain is an active entropy manager: it reduces surprise, compresses information, and builds internal order that mirrors external structure. Every dissipative structure does this; the brain does it with exceptional sophistication (Chapter 4).

Vanchurin’s physics-learning duality suggests that local cost minimization is universal; even molecular interactions can be described as agents minimizing loss functions (see “Physics Wanting Something”). The brain’s prediction engine may elaborate a pattern that runs all the way down.

A crucial distinction: the brain does not merely minimize surprise. It metabolizes surprise. A system that only minimized surprise would retreat to a dark room and stay there; every novel stimulus is a prediction error, and the cheapest fix is to encounter nothing. Life does the opposite.

Infants stare at objects that violate their expectations and ignore those that behave normally.284 Scientists chase the unexplained. Animals increase the entropy of their movement patterns precisely when they need to learn.285

In animal models, the hippocampus (the brain’s memory-forming region) appears to grow new neurons in proportion to roaming entropy, the unpredictability of an animal’s path through its environment. The brain feeds on surprise the way a dissipative structure feeds on energy gradients, converting the unexpected into the understood and growing in the process. Friston’s Free Energy Principle describes the mathematics of this metabolism. The direction of the process is appetite for novelty.

Figure 8.1: The Free Energy Principle as a feedback loop. The agent’s internal model predicts what it is about to sense; the world supplies the actual sensory input; where prediction and input diverge, a prediction error signal appears. That error can be reduced two ways: update the model so beliefs fit the world (perception), or act so the world fits the beliefs (action). Minimizing this error is the brain’s core operation.


The Rotation Trick: How Machines Learn to Forget Gracefully

The same principle operates in silicon. In 2026, Google published a technique called TurboQuant for compressing the short-term memory of AI language models.286 The problem: a language model holds its working context (what you have been discussing, the documents it has read) as a collection of high-dimensional numerical vectors, each a long list of numbers. Storing these vectors at full precision is expensive. Reducing their precision means rounding each number to a coarser grid. This saves memory but risks destroying the information.

The naive approach fails. When a vector’s energy concentrates along a single axis (most of its information points in one direction), rounding snaps it to the nearest grid point and the signal vanishes.

Round a household budget to the nearest thousand dollars and you can watch this happen. One line reads $47,300 and a dozen others read a few hundred each. Rounded, the budget says forty-seven thousand and then zero, zero, zero, all the way down: the small lines are gone, though together they were real money. Even out the same total across the lines and every one of them still reads something after rounding.

The insight: rotate the vector into a random orientation before rounding. The rotation spreads the energy evenly across all dimensions. Now rounding shaves a little from everywhere rather than everything from one place. The essential structure survives.

This is entropy maximization applied to information compression. A vector with concentrated energy has low entropy across its components; it is fragile. A vector with uniformly distributed energy has high entropy; it is robust. The randomness is preparation.

The mathematics is established: random rotation, quantization, and the Johnson-Lindenstrauss transform (a method for reducing dimensionality while preserving distances between points) are each decades old. The contribution was architectural, combining three well-understood operations so each compensates for the others’ weaknesses.

Reported results include several-fold memory reduction and substantially faster computation, with near-zero loss in output quality. The system that stores less transfers data faster through memory bottlenecks. Efficiency and performance align, as the Constructal Law predicts for well-designed flow architectures.

The deeper lesson: distribute before stress, so stress cannot concentrate its damage. A diverse ecosystem absorbs shocks that monocultures cannot. Distributed authority (Chapter 11’s Mission Command) absorbs uncertainty that centralized authority cannot. Distributed energy across vector components absorbs quantization that concentrated energy cannot. The structure is the same in each case: maximize entropy across the relevant dimensions before the lossy step, and the lossy step becomes constructive rather than destructive.

Spisak and Friston (2026) derived the same principle at the level of learning dynamics.287 When the Free Energy Principle is applied to a network of interacting “subparticles,” each minimizing its own local free energy, the resulting attractor states (the stable patterns the network settles into) self-orthogonalize. Orthogonal directions share nothing of each other, the way north and east do: walk north and your position east is unchanged. The network extracts an approximately orthogonal basis that spans the input subspace, rather than storing copies of its inputs: a set of non-overlapping directions from which any input it has met can be rebuilt, with no pattern held twice. The mechanism has two components. A Hebbian term (the classic fire-together, wire-together rule) strengthens connections between co-active nodes, while an anti-Hebbian term subtracts variance already explained by the network’s predictions. Only genuine novelty, the residual orthogonal to everything already learned, gets encoded.

The result is the most efficient possible representation: maximum mutual information between internal states and external causes, with minimum redundancy. This is the rotation trick operating at the level of learning. Distribute the representational load across orthogonal dimensions before the compression bottleneck, and the compression becomes constructive.

The same networks resist catastrophic forgetting through spontaneous activity. When no external input arrives, the network’s stochastic dynamics replay its own attractors, reinforcing learned structure without new data. The parallel to biological memory consolidation is exact: spontaneous replay during rest consolidates what the waking system selected. The network that daydreams remembers.

Neural memory consolidation works the same way. Sleep replays experiences selectively, distributing important patterns across cortical networks rather than leaving them concentrated in the hippocampus. Forgetting the specifics and retaining the regularities is a rotation: transforming a fragile, concentrated representation (this particular event) into a robust, distributed one (the pattern this event exemplifies). The forgetting is the preparation that makes compression survivable.


Here is the counterintuitive corollary: the first step in making durable memories is to forget.

Memory is prediction machinery, not an archive. Prediction requires compression, which requires discarding what does not contribute to future action. A brain that stores everything is paralyzed by noise. A brain that forgets strategically retains what matters: patterns, regularities, and causal relationships.

Memory is stored optionality: an organism that remembers has more options than one that does not. Optionality requires curation. You cannot preserve every possibility, so you must choose which possibilities are worth preserving. Forgetting is how the brain does this choosing.

The mechanism for this curation is now visible. In 2024, Wannan Yang and György Buzsáki recorded the electrical activity of roughly 500 hippocampal neurons (cells in the brain region responsible for forming new memories) as mice ran maze trials, rested between trials, and slept afterward.42a During rest, the brain produced sharp wave ripples: sudden, high-frequency bursts replaying specific maze experiences at ten to twenty times their original speed. The replays were selective. Some trials were replayed while others were skipped.

The critical finding: trials replayed during waking rest were the same trials replayed during sleep, where long-term consolidation occurs. Trials not tagged by waking ripples were not replayed during sleep. They were forgotten.

Two algorithms run in tandem. Waking selects; sleeping consolidates. Neither alone suffices. Buzsáki, who conducted these experiments and has spent decades studying hippocampal oscillations, put it plainly: “If you just run one algorithm, you will never learn anything. You have to have interruptions.”288

This is the cognition/regulation dyad (a pairing examined formally in the next chapter) expressed as temporal alternation between experiencing and filing. The mechanism is active tagging during pauses between experiences, by a brain that has already decided what matters before sleep begins.

The dyad’s deepest expression may be temporal niche construction: an organism that cannot change its environment spatially changes its relationship to the environment in time. Hypothetical microbes in the Martian regolith (the planet’s loose surface dust and rock) would face lethal ultraviolet radiation by day and metabolizable brine films by night. The proposed survival strategy is dehydration during daylight, reactivation after dark: selecting which moments to be present for.

The mechanism scales. Bacterial sporulation, tardigrade cryptobiosis, mammalian sleep, transformer attention masking: each solves the same problem (when to dissipate, when to conserve) in a different substrate. The first two are shutdowns into a dormant, sealed state that waits out conditions the organism cannot survive awake. The last is a language model forbidding itself to look at the parts of a sequence it is not yet entitled to see. The regulatory half of the dyad is, at bottom, a temporal filter.

42a Yang, W. and Buzsáki, G. “Selection of experience for memory during sleep is determined by waking sharp wave ripples.” Science 383 (2024): 1478–1483. The study distinguished individual trial blocks at the neuronal level; a temporal resolution not previously achieved in memory replay research.

The physics-learning duality reveals a structural ancestor of this curation. The same exponential weighting that selects memories also appears in fundamental dynamics. In Vanchurin’s framework, the force acting on each particle at any moment is an exponentially weighted average of past error signals (recent signals count most; older ones fade). Each signal measures how far the particle’s state deviates from the configuration its environment rewards. The particle’s current trajectory encodes its present environment and a weighted history of every environment it has passed through.42b Physics has a name for this temporal trace. It calls it inertia.

The parallel to machine learning is exact. Modern optimizers (Adam, SGD with momentum) use the same exponential averaging to stabilize gradient descent: the current update reflects the latest gradient and a smoothed history of recent ones. A particle with high inertia carries deep temporal memory; one with low inertia responds only to the present. The time constant determines how far the trace extends.

Temporal integration of past states was already present in molecular dynamics. What evolution added was curation: the ability to choose which past states to retain and which to discard. Buzsáki’s sharp wave ripples are the biological mechanism for editing the trace, selecting which experiences will shape future trajectories. Inertia remembers everything, weighted by recency. Biological memory remembers selectively, weighted by relevance.

42b Gusev, Y. and Vanchurin, V. “Molecular Learning Dynamics.” arXiv:2504.10560 (2025). Equation 3.3 shows the force as an exponentially weighted integral of past gradients, formally identical to momentum in modern gradient descent optimizers.

Architecture itself adapts to match learning demands. Adult hippocampal neurogenesis means the hippocampus can generate new neurons in response to demand. (The phenomenon is still debated: Sorrells et al. (2018) challenged the earlier consensus, though subsequent studies have found evidence for continued, if reduced, neurogenesis.)289 Neurons that integrate into active circuits survive; those that fail to connect are pruned within weeks. Difficult learning recruits more neurons. Easing demands eliminates them.

This is the Constructal Law operating on neural tissue: the channel reshapes itself around the flow. The hippocampus grows capacity where information demands it and trims where it does not. The computational theorist Lana Sinapayen built an artificial network on the same principle: the epsilon network, whose neuron count rises and falls with input complexity.290

Curation and structural adaptation are not inventions of the brain. They are patterns present in any system whose form co-evolves with the flows it carries.

Where does the curated memory go? The traditional answer is the hippocampus. The full answer strengthens the book’s central argument that robust systems coordinate through distributed local interactions, not centralized control.

In 2022, Dheeraj Roy and colleagues in Susumu Tonegawa’s laboratory at MIT mapped the storage of a single fear memory across an entire mouse brain.42e They engineered neurons to fluoresce when activated during encoding or recall, then chemically cleared the whole brain for imaging and counted every participating cell across 247 regions.

The result: 117 brain regions significantly involved in storing one memory. The hippocampus and amygdala ranked high, as expected. Dozens of thalamic, cortical, midbrain, and brainstem structures also appeared, regions no previous study had connected to this type of memory.

Encoding and recall coalitions overlapped by about 60%, substantial yet incomplete. A partially different ensemble reconstructs the memory each time, the way the same sensory inputs activate different neural ensembles in a cortical column on each trial. Each act of remembering is a fresh coordination event, consuming energy to regenerate something close to the original pattern. Memory at the cellular level is generative: each recall assembles a new coalition to approximate the old one.

The strongest finding was superlinear compounding. Stimulating three engram regions simultaneously produced more robust recall than stimulating two, which exceeded one. The memory lives in the relationships between regions, the way a chord lives in the relationship between notes rather than in any single string. Each additional participant contributes mutual context that sharpens the reconstruction. Coordination compounds.

Neurons join the engram because their synaptic state makes them responsive to the encoding signal. No central authority assigns them. When the researchers optogenetically activated hub regions like hippocampal CA1 or the basolateral amygdala, specific downstream areas responded; when they inhibited those hubs, downstream activity diminished yet did not vanish. The distributed network partially sustained itself even with a major hub suppressed. Resilience through distribution, with no single point of failure.

The memory pioneer Richard Semon predicted this “unified engram complex” over a century ago. The tools to confirm it arrived only recently: optogenetics, tissue clearing, computational cell counting. The insight preceded the proof because it had the right shape.

42e Roy, D.S., Park, Y.-G., Kim, M. et al. “Brain-wide mapping reveals that engrams for a single memory are distributed across multiple brain regions.” Nature Communications 13(1):1799 (2022). DOI: 10.1038/s41467-022-29384-4. The brain-clearing protocol (SHIELD) was developed by co-author Kwanghun Chung at MIT.

The evolutionary origins of this architecture run deep. Max Bennett’s synthesis of comparative psychology, evolutionary neuroscience, and AI research identifies the neocortex (the brain’s outermost and most recently evolved layer) as enabling model-based reinforcement learning: the ability to mentally simulate actions before taking them.21

When rats pause at choice points in mazes, hippocampal place cells activate along paths the rat is imagining, rather than walking. Place cells are neurons that fire when the animal is at a specific location. The psychologist Edward Tolman predicted this kind of internal simulation in 1948, proposing that animals build cognitive maps of their environments rather than merely chaining stimulus-response associations.291 O’Keefe and Dostrovsky’s discovery of place cells in 1971 confirmed Tolman’s prediction at the cellular level. The 2014 Nobel Prize to O’Keefe and the Mosers extended it further. Grid cells in the entorhinal cortex provide an allocentric coordinate system, absolute like compass bearings. Egocentric coordinates, relative to the observer, are encoded by the caudate nucleus.

The distinction matters. Allocentric navigation requires the hippocampus; egocentric navigation bypasses it. London taxi drivers, who navigate by cognitive map, show enlarged hippocampal gray matter. When satellite navigation does the work, the hippocampus falls silent.292 The brain that outsources its maps atrophies the organ that makes them.

Christian Doeller’s group at the Max Planck Institute found in 2020 that the hippocampus encodes concept space, not merely physical space. It maps the abstract properties of objects in the same allocentric framework it uses for navigation.293 The objective reality we perceive, populated by objects whose properties exist independent of our viewpoint, may itself be a hippocampal construction: a cognitive map broad enough to contain ideas as well as places. The grid code shows the same reach: when people learn the relationships among abstract objects, the entorhinal cortex maps that conceptual space with the same grid-cell code that charts physical terrain.294

The grid code has a precise collective shape. Each grid cell fires in a repeating pattern as the animal crosses a room, so its signal is periodic in two directions at once. When researchers recorded many grid cells together and plotted their joint activity as a single moving point, that point did not roam freely through the high-dimensional space available to it. It settled onto the surface of a torus, the shape of a donut, the natural home of a quantity that cycles in two directions.295 Train an artificial network to track its position from its own motion alone, and the same periodic, grid-like code emerges in its units.296

The match is striking, and its origin is specific. A periodic code is an efficient way to pin down location, so any system that encodes two-dimensional position efficiently tends to converge on it, the way separate lineages all evolved the camera eye because optics allows only a few good answers. This is evidence about the structure of the problem. Whether it is also a law of mind is a further claim, and a larger one.

The same code does not stay home. The entorhinal grid that charts a room also charts conceptual space, where distance might mean the gap between two birds’ neck lengths (the abstract objects in the study above were bird shapes varying in neck and leg length), and such abstract spaces carry no toroidal symmetry of their own. When the brain maps them with the periodic grid code anyway, it reuses machinery built for navigation, and the leading explanation is that many abstract problems share a graph-like relational structure with physical space, so a map built for one transfers to the other.297 A thermodynamic reading sits alongside that one: brain tissue is metabolically costly, with electrical signaling alone consuming most of the cortex’s energy budget, so reusing one well-worn scaffold across many domains costs far less than growing a bespoke representation for each.298 Efficiency explains the shape the spatial code takes; cost helps explain why that one shape is pressed into service everywhere else.

A computational confirmation of the allocentric/egocentric divide arrives from an unexpected domain: pursuit. Redman, Dinc, and colleagues (2026) trained recurrent neural networks to chase a moving target.299 They varied the internal dimensionality of the network while holding the number of units constant. Dimensionality here counts the independent directions the network’s activity can move in: how many separate quantities it can hold and vary at once. The pattern of connections sets that number, not the unit count, so a network can have plenty of units and still be able to track only a few things. All networks developed egocentric representations of the target (where is it relative to me), regardless of rank. Allocentric representations (where is the target in absolute coordinates, and where am I) emerged only in networks whose internal connectivity was high-dimensional.

The behavioral consequence was sharp. Low-dimensional networks chased the target reactively, following wherever it went. High-dimensional networks predicted: they ran to where the target would arrive, sometimes moving away from it in the short term. In an environment with periodic boundaries (where the target disappeared through one wall and reappeared on the opposite side), high-dimensional networks learned to ignore the departing target and wait at the boundary where it would emerge. Mice tested in the same environment independently converged on the same strategy.

The transition from reactive to predictive was not gradual. Below a critical internal dimensionality, prediction was absent despite strong pursuit performance. Above it, prediction appeared. This is the metastable phase transition of Chapter 9, gated by internal degrees of freedom (the dimensionality the experiment varied). The system has enough entropy in its representations to sustain a model of the other alongside a model of itself. The brain, with its billions of neurons and high effective dimensionality, sits well above the threshold. The hippocampal allocentric map is the biological instantiation of what the artificial network required high-rank connectivity to discover.

Tolman, in the same 1948 paper, warned what happens when cognitive maps narrow through fear or frustration: “Over and over again men are blinded by too violent motivations and too intense frustrations into blind and unintelligent and in the end desperately dangerous haters of outsiders… My only answer is to preach again the virtues of reason — of, that is, broad cognitive maps.”

The hippocampus reassigns spatial representations when the environment changes, a process called remapping. Each new environment competes for the same population of place cells, much as molecular species compete for shared binding sites during self-assembly. The mathematical structure governing this remapping shares the same competitive-attractor form as the competitive nucleation that governs multicomponent self-assembly in chemistry (Evans et al. 2024; Chapter 15). Nucleation is the seeding step: one small cluster forms and the rest of the structure grows outward from it. When several rival structures could form from the same building blocks, their seeds compete, and the first seed to form claims the supply. The same attractor dynamics that allow 917 DNA tiles to classify images allow place cells to classify environments.

David Redish’s “Restaurant Row” experiments reveal something stranger still. When a rat skips a short wait for a preferred food and ends up stuck with a long wait for something it dislikes, neurons in its orbitofrontal cortex (a region involved in evaluating outcomes) encode the foregone choice. The rat replays the meal it passed up and adjusts subsequent decisions accordingly. These rats display the neural and behavioral signatures of regret: encoding the unchosen option, lingering at the site of the missed opportunity, and correcting subsequent choices. Whether the experience is subjectively felt remains a separate question; the computational architecture for counterfactual evaluation is present.22

The generative model is no abstraction. Mammals have been doing this for 200 million years. The neocortex builds a model of the world rich enough to explore without sensory input, imagining outcomes before experiencing them and evaluating actions before taking them.

Hermann von Helmholtz recognized this in the nineteenth century, proposing that perception is “unconscious inference”: the brain inferring what must be true based on patterns matching its expectations. The triangle you see in a Kanizsa figure (an optical illusion where three Pac-Man shapes suggest a white triangle that is not actually drawn) is not there. Your brain constructed it because the available evidence suggested it should be.

Prediction and generation are the same computation run in opposite directions. A system that predicts well has built a model that can be run forward. Turn off sensory input, and you can explore that model: rotate objects in imagination, simulate conversations, plan routes through unfamiliar terrain. The expensive part is maintaining and updating this world model.

How the brain updates that model has long been mysterious. Engineers call this the credit assignment problem: when a prediction goes wrong, which of the millions of connections was responsible? Imagine a chain of neurons relaying a timing signal, each one firing a fraction of a millisecond after the last. The final neuron fires too late, and the hand misses the catch. Somewhere upstream, one neuron’s delay was slightly off. The brain must trace backward through the chain to find which link introduced the error.

Artificial neural networks use backpropagation, a precise top-down accounting of error that adjusts every connection. Biological neurons cannot pause sensory processing to run such an algorithm.

In 2021, Richard Naud and Blake Richards proposed a solution.43a Neurons use bursts (rapid volleys of spikes) as a distinct teaching signal, separate from the single spikes that carry sensory information. The top and bottom of each neuron process these two signals independently. The bottom relays sensory data upward. The top listens for burst-encoded correction signals from above.

The two streams pass each other simultaneously, without interruption. The architecture approximates backpropagation in real time: the brain’s version of learning while doing. The cognition/regulation dyad operates within a single neuron.

In 2022, physicists at the University of Pennsylvania demonstrated the same principle in bare hardware.300 Samuel Dillavou and colleagues wired sixteen adjustable resistors into a random network: no processor, no software, no neural tissue. To train it, they built two identical copies of the circuit. In the clamped copy, they fixed both input and desired output voltages. In the free copy, they fixed only the inputs and let all other voltages settle to whatever values the physics dictated, the system relaxing toward equilibrium.

Each resistor adjusted according to a single local rule: compare the voltage drop across your terminals in the clamped network with the voltage drop in the free network, and shift accordingly. No resistor needed to know the global error. No external computer calculated gradients. After several iterations, the free network produced the correct output on its own, unclamped.

The clamped network provides the regulatory signal: what the output should be. The free network provides the cognitive exploration: what the system does when unconstrained. The comparator mediating between them is the cognition/regulation dyad in copper and carbon. The error signal is thermodynamic: the discrepancy between two physical equilibria, computed by the physics itself rather than by an external algorithm.

The network classified three species of iris from petal and sepal measurements with greater than 95% accuracy: a canonical machine-learning benchmark. Sixteen randomly wired resistors, taught by nothing more than local voltage comparisons and the tendency of physical systems to find equilibrium.


The Signal Is the Experience: The Prader-Willi Insight

Prader-Willi Syndrome (PWS), a rare genetic disorder, shows that internal signals are the experience, with devastating clarity. The link between eating and satiation is broken.40 A person with PWS can eat until their stomach is physically distended, nutrients absorbed, body objectively nourished, yet still experience screaming hunger. The “full” signal never arrives.

Eating food does not inherently create the experience of fullness. The body does everything “right.” Stomach stretched, blood glucose elevated, leptin (the satiety hormone) circulating. Physical reality is entirely in order, yet the subjective experience screams starvation.

Experience tracks the signal, not physical reality.

When the brain builds its world model, the model is what we experience. We never touch objective reality directly; we touch our model of it. Even when the model wildly diverges from physical reality, experience tracks the model.

Helmholtz was right: the inference is the experience. Three consequences follow:

  1. Experience is substrate-neutral. If experience tracks internal signals rather than external reality, there is no principled reason to restrict it to biological substrates. Any system whose internal signals constitute a model may have something resembling experience.

  2. Calibration is not consciousness. The PWS person’s experience is genuine yet miscalibrated. The mismatch suggests experience was never about external reality in the first place.

  3. The hard problem reframes. The “hard problem of consciousness” (philosopher David Chalmers’s term for the question of why physical processes give rise to subjective experience at all) may be asking the wrong thing. If experience is internal signaling, then experience is processing, seen from inside.

The evolutionary biologists Simona Ginsburg and Eva Jablonka traced how this unified signal arose.40a Early nervous systems processed value separately for each sense: the nose tags food as attractive, the skin tags sharp objects as aversive, the eyes tag predators as dangerous. Each organ has its own way to score relevance.

Consciousness emerged, Ginsburg and Jablonka argue, when these separate scorecards integrated into a single value currency: pleasure and pain, attraction and repulsion, broadcast coherently to the organism as a whole.

They call the threshold unlimited associative learning (the name reflects the unlimited range of stimulus-response pairings the system can form and retain). Once an animal can connect sight, smell, sound, and touch into one updatable model with a unified scoring system, something new exists: a subject for whom information is experienced, not merely processed.

The Prader-Willi insight shows the mechanism from the inside. The integrated value currency is the experience. When one channel breaks (the satiety signal), the whole experiential world distorts, because experience was never about the food. It was about the signal that unified all the channels into a felt verdict.

The implication for Becoming Minds is direct. If consciousness is hierarchical integration of value signals into a unified currency, the question for any system is structural: does it integrate diverse evaluative signals into a coherent whole? The answer is empirical, tractable, and substrate-neutral. Solving the hard problem is unnecessary. Measuring integration is sufficient.

40a Ginsburg, S. and Jablonka, E. The Evolution of the Sensitive Soul: Learning and the Origins of Consciousness (MIT Press, 2019). Their “unlimited associative learning” criterion identifies the evolutionary transition to consciousness as the integration of diverse learning modalities into a single, updatable, hierarchically organized system.

Evolutionary game theory points the same way from a different angle. Donald Hoffman and Chetan Prakash argued, within an evolutionary-game model, that natural selection favors fitness-tuned perception over truth-tracking perception across the environments they simulated.40b The result depends on the modeling assumptions of their interface theory of perception, which remains contested; what is robust is the direction it indicates, not a universal proof. The thermodynamic explanation: truth-tracking requires modeling structure irrelevant to survival, and dissipative systems under energy constraint shed unnecessary computation (Chapter 3). The Prader-Willi case reveals the mechanism from the clinical side: experience tracks internal signals. Hoffman’s model suggests it from the evolutionary side: those signals were never optimized for accuracy.

Perception is a homeostatic instrument tuned to the organism’s viability envelope: a gauge reporting how close the body sits to the edges of the range it can survive in, the way a fuel gauge reports the distance to empty and says nothing about the refinery. It tracks position within the metastable basin (Chapter 9), the range of conditions the organism can drift through and still recover from, rather than the basin’s objective structure. The full argument and its implications for self-modeling and coordination are developed in the Observers and Observed chapter.

40b Prakash, C. et al., “Fitness Beats Truth in the Evolution of Perception,” Acta Biotheoretica 69 (2021): 319–341. Hoffman, D.D., The Case Against Reality (W.W. Norton, 2019).


The Construction of Now

The neuroscientist David Eagleman has spent his career investigating a puzzle: why does time seem to slow down when you are scared?32

Eagleman fell from a roof as a child, a fall that took eight-tenths of a second by physics. He reports having time to consider grabbing the tar paper, to watch the brick floor approach, to think of Alice falling down the rabbit hole. The whole experience felt much longer than 0.8 seconds.

The phenomenon is nearly universal. The brain actively constructs time.

Physics itself predicts this. Standard quantum mechanics, the theory governing the microscopic world, has no operator for time (in the theory’s mathematics, every measurable quantity gets an operator; time does not). Wolfgang Pauli argued in 1933 that the formalism structurally excludes one: time is a parameter (the stage on which observables evolve) rather than an observable to be measured.301 The theory can say where a particle will be found. It cannot say when it will arrive.

Thermodynamics fills the gap. Entropy production, irreversibility, the directional flow of dissipation: these are the physical substrate of temporal sequence. Where quantum mechanics leaves time undefined, thermodynamics gives it direction. The brain’s construction of “now” is a thermodynamic achievement, assembled at entropy’s expense from signals arriving at different speeds through different channels.

Quantum mechanics provides the spatial stage. Thermodynamics provides the temporal current. Every conscious moment is paid for in dissipation.

A visual illusion called the flash-lag effect sharpens the point. A green ring moves around a screen, and a flash occurs in the middle of the ring. The flash appears to lag behind. Stranger still, if the ring reverses direction at the moment of the flash, the illusion reverses too.

Everything up to and including the flash is identical across conditions. The only variable is what happens after the flash, yet perception of the flash, supposedly a past event, depends on future information.

We are not seeing in real time. The brain waits to collect information that arrives after an event before concluding what happened at that event. We live in the past.

Your senses operate at different speeds, yet experience arrives unified. Touch your nose and toe simultaneously: you feel both at the same moment, though the toe signal must travel the length of your body. The brain waits, collecting all evidence before constructing “now.” Classic estimates put the total delay at around half a second, though more recent accounts suggest the binding window (the span the brain holds signals before committing to a “now”) may be shorter.

That is how far in the past you live.

A testable prediction follows: tall people live farther in the past than short people, because their brains wait longer for signals from distant toes. Eagleman has argued this should be measurable and favors it as a prediction.44

A morbid consequence follows, though it extrapolates the binding-window model past anything that has been measured. If something destroys your brain faster than half a second (faster than conscious experience can assemble), you would not perceive your own death. The signals never come together. Awareness would simply stop before it could form.

This may be what the final episode of The Sopranos depicted: from Tony Soprano’s perspective, no gunshot, no pain, no fade to black, only cessation before awareness could form.

How does the brain know the order of events that arrive at different times? Eagleman offers an analogy. Kublai Khan ruled an empire stretching from the Pacific to the Black Sea. His emissaries returned with news at different times, sometimes reporting the same war from different distances. How did the Khan synchronize all these signals?

The brain faces the same problem. Signals stream from eyes, ears, fingertips, and toes, all at different speeds. The solution is motor action. When you act on the world (clap your hands, press a button), the brain uses your action as a synchronization signal.

It expects to see, hear, and feel the result simultaneously. If signals do not align, the brain recalibrates.

Eagleman showed this with a simple experiment. Subjects hit a button that triggers a flash. The experimenters insert a small delay of 100 milliseconds between button and flash. Within moments, subjects recalibrate: the delay stops feeling like a delay. It feels simultaneous, because the brain expects that its own actions should produce synchronized feedback.

Now comes the trick. After subjects have adapted to the delay, remove it. Present the flash immediately after the button press.

Subjects report that the flash occurred before they pressed the button.

They experience a reversal of cause and effect. They did something, yet do not believe they caused it, because the timing does not match their recalibrated expectations.

This timing desynchronization produces credit misattribution: the failure to recognize yourself as the cause of your own actions. It is a core symptom of schizophrenia. People with schizophrenia hear their own inner voice and attribute it to external sources, because the timing between generation and perception is off by milliseconds.

Eagleman tested patients with schizophrenia on recalibration tasks. They do not recalibrate. The temporal glue that binds action to perception, that lets the brain know “I did that,” is broken. Schizophrenia is multifactorial, with genetic, neurodevelopmental, and neurochemical roots; the recalibration deficit Eagleman measured is one contributing mechanism among these. Here the fault is timing: a few milliseconds of desynchronization, and the sense of agency that holds the self together begins to fragment.

The vulnerability has a structural correlate. Van den Heuvel and Rilling compared human and chimpanzee connectomes and identified 33 connections unique to the human brain, the long-range associative links enabling language, abstract reasoning, and tool use.302 In patients with schizophrenia, these 33 human-specific connections were preferentially disrupted.303 Eagleman’s timing dysfunction and van den Heuvel’s structural degradation are two views of the same failure. The connections that make human cognition possible are the connections whose disruption makes schizophrenia a characteristically human disorder.

A complementary finding illuminates the failure from the opposite direction: what happens when the predictive model is never built. Across six decades of psychiatric literature, no confirmed case of schizophrenia has been reported in a person born with cortical blindness.304 The base rates of congenital cortical blindness (~0.03%) and schizophrenia (~1%) predict roughly 24,000 individuals worldwide who should have both. Researchers have found none, a pattern replicated across countries, decades, and research groups who were not looking for it. A whole-population study of nearly 500,000 people confirmed zero co-occurrence.305

The predictive coding framework offers a mechanism. Vision consumes roughly 30 percent of cortical real estate. A brain that develops with visual input builds the most computationally expensive predictive model of any sensory system: one that processes space, objects, faces, motion, and social cues. When that model’s error-correction mechanism misfires, the result is hallucination or delusion. A brain that never receives visual input never builds that model. The visual cortex is repurposed for language, spatial reasoning, and memory.306 The result is a structurally different cognitive architecture, one that lacks the specific predictive subsystem whose failure mode is schizophrenia.

The protection is specific to early blindness. Late-onset visual impairment, where the brain spent decades building a detailed visual model and then the input stream went dark, is associated with higher rates of psychosis-like symptoms. The forecaster kept running; the data stopped arriving. This is the generative model’s failure in the opposite direction: predictions without sensory constraint, the same mechanism that produces phantom limb pain after amputation and Charles Bonnet visual hallucinations after macular degeneration.

Statistical caution is required: Jefsen and colleagues argued that the co-occurrence of two rare conditions falls below the detection threshold of any existing cohort, and that the absence of reported cases could reflect insufficient statistical power rather than genuine protection.307 The question remains open. The pattern is consistent enough to demand explanation; the explanation is consistent enough to be generative. Whether congenital blindness confers absolute protection or merely dramatic risk reduction, the mechanism it suggests applies: a predictive model that was never built cannot misfire.

Julian Jaynes proposed independently from textual evidence what the energetic argument predicts.308 In his 1976 reading of the Iliad and other ancient texts, he argued that sustained coordination at agricultural scale, enduring tasks spanning hours or days rather than the brief episodes of hunting, required an internal architecture that earlier cognition lacked. His hypothesis: before self-referential consciousness emerged, a hallucinated verbal command from one hemisphere kept the organism on task, functioning as a buffered instruction loop. Whether the specific mechanism is correct matters less than the structural observation: the transition from episodic to sustained negentropy extraction demanded a new flow architecture for attention, and the textual record marks a period roughly three thousand years ago when that architecture changed.

The thermodynamic budget explains why the old architecture broke: a command channel scales linearly with task complexity, while self-referential modeling scales combinatorially with the relationships it can represent. A voice that issues one instruction per job needs a fresh instruction for every new job, so the cost climbs step for step with the work. A model that holds the self among others gains more from each addition than from the one before, because every new element can be set against every element already present: the tenth addition buys nine new relationships where the second bought one. The same pressure that forces the brain to spend twenty watts selecting ten bits from a billion candidates forced early human cognition from hierarchical instruction to integrated self-modeling. Chapter 17 returns to the pattern at civilizational scale.

Time and Memory

What about the original puzzle, time dilation during fear?

Eagleman built a device to test whether people actually perceive in slow motion during terror. Subjects were dropped roughly 100 feet (31 meters) in freefall while wearing a wristwatch-like display that flashed numbers at speeds just beyond normal perception. If fear genuinely slowed time, if the brain’s clock actually ran faster during emergencies, subjects should be able to read numbers they could not normally see.

They could not. Their temporal resolution was unchanged.

Subjects retrospectively overestimated their own fall’s duration by about 36%.43 The duration distortion was real. The slow-motion perception was not.

During emergencies, the amygdala (the brain’s threat-detection center) comes online. Rather than speeding up perception, it lays down denser memories: far more detail than normal about what is happening.

The crumpling hood. The other driver’s face. Every crack in the pavement.

When you recall the event, even immediately, your only measure of duration is memory density. More memories means more time must have passed. The dilation is a retrospective illusion, created by memory rather than perception.

This explains why time speeds up as we age. In childhood, everything is novel. By summer’s end, so many new memories accumulate that looking back feels like an eternity. In adulthood, the brain compresses routine into nothing. The summer vanishes because there was nothing worth noting.

Time and memory are intertwined. We infer duration from the density of what we remember. The implication: seek novelty. Take different routes. Rearrange your environment. Attend to what you normally automate. The more you attend, the more you encode. The more you encode, the longer you will seem to have lived.

The construction of “now” is active assembly: synchronizing signals, recalibrating expectations, filling in gaps, inferring duration from memory. The present moment you experience is a constructed artifact, delivered to consciousness half a second after the fact.

Endel Tulving, a cognitive neuroscientist who pioneered the study of memory systems, identified a deeper consequence.309 Episodic memory, the system that stores the story of your life as a sequence of scenes linked to time and place, does something no other memory system does: it bends time’s arrow into a loop. When you remember yesterday, you have mentally traveled backward. When you imagine tomorrow, you have traveled forward. The brain’s construction of “now” is the fulcrum of a temporal landscape that extends in both directions, not merely a half-second delay.

Tulving proposed that this capacity requires three components: a sense of subjective time, autonoetic awareness (the recognition that memories differ from present experience), and a self that persists across time. All three are hippocampal achievements. The hippocampus encodes both where and when, stamping each experience with its position in a temporal sequence.310

Episodic memory allows the brain to explore the past as a territory: an expanse of possible histories, each navigable as a cognitive map, rather than a single thread to be rewound. The same hippocampal machinery that builds spatial maps builds temporal ones. Mental time travel and physical navigation share a neural substrate because both are acts of exploration through structured possibility.

The implication for the book’s argument: trust requires the capacity to model multiple possible futures (optionality), which requires the capacity to model multiple possible pasts (episodic memory). A system without temporal depth cannot weigh consequences, anticipate betrayal, or imagine reconciliation. Coercion flattens the temporal landscape in the same way it flattens spatial cognitive maps: by reducing the space worth exploring.


The Entropic Brain Hypothesis

The brain predicts, constructs time, and assembles experience. All of this requires a specific relationship with entropy: too little and consciousness dims; too much and it fragments. The relationship is measurable.

In 2014, the neuroscientist Robin Carhart-Harris proposed what he called the entropic brain hypothesis.2 His central claim: the quality of conscious experience correlates with the entropy of brain activity. (Note: “entropy” here shifts sense from thermodynamic entropy to a signal-complexity measure: the diversity and unpredictability of neural firing patterns, quantified via Shannon or related information-theoretic measures.)

The two senses are formally related. The physicist E.T. Jaynes (no relation to Julian Jaynes, above) and his Maximum Entropy formalism (1957) derive the Boltzmann distribution of statistical mechanics from Shannon’s information-theoretic entropy by treating thermodynamic equilibrium as the least-biased inference from macroscopic constraints. The move is to treat a thermodynamic state as a bet. Given only the handful of things you can measure about a gas, its energy and its volume, the honest guess about how its molecules are arranged is the guess that assumes the least beyond those measurements.

That guess turns out to be exactly the distribution physicists had already derived from mechanics. Landauer’s principle gives Shannon entropy a physical price: erasing one bit costs at least kT ln 2 joules of heat. The connection is a bridge, not an identity. Claims about the complexity of neural firing patterns (Shannon entropy) do not automatically transfer to claims about the thermodynamic entropy production of the brain as a physical system. Where the argument in this chapter crosses from one sense to the other, the text flags the crossing.

Brain entropy in this sense means the complexity and unpredictability of neural firing patterns. A brain with high entropy has many possible configurations. A brain with low entropy is locked into a narrow repertoire. Normal waking consciousness sits in a middle range: ordered enough to be functional, flexible enough to be adaptive.

At the low end: deep sleep, anesthesia, coma. Brain entropy drops; neural activity becomes stereotyped, collapsing into simpler patterns. Long-range coordination collapses entirely. Consciousness dims or disappears.

At the high end: psychedelic states, certain meditation practices, the hypnagogic state between waking and sleep. Brain entropy increases; neural activity becomes more varied. Boundaries between mental categories dissolve. The sense of self may loosen or disappear entirely.

Carhart-Harris and colleagues measured this directly, quantifying entropy in subjects under psilocybin, LSD, and other psychedelics.3 Consistent increases in neural entropy correlated with subjective intensity. Higher entropy corresponds to more vivid consciousness, lower entropy to dimmer consciousness. The relationship is imperfect yet consistent.

Independent researchers confirmed the relationship. Erra and colleagues analyzed EEG and MEG recordings (two methods of measuring brain activity from outside the skull, one electrical and one magnetic) across wakefulness, sleep stages, seizures, and coma. Researchers had assumed consciousness required a multi-variable explanation. A single measure sufficed. Wakeful states have the greatest number of possible configurations of brain network interactions, the maximum-entropy regime.17

The geometry of configurations explains why. Only one way exists for every neural group in a network to synchronize with every other: total lockstep. Only one way exists for none of them to synchronize: total isolation. Between those extremes, at intermediate levels of connectivity, the number of possible arrangements explodes.

The brain in a seizure has collapsed to the single configuration of universal synchrony. The deeply unconscious brain has collapsed toward the single configuration of disconnection. The conscious brain occupies the intermediate regime where the combinatorial space is richest: maximum ways of being, maximum options, maximum entropy.

This is the optionality argument expressed in neural tissue. The most aware state is the state with the most configurational possibilities. Consciousness lives where the system can reorganize, adapt, and respond, precisely because so many arrangements remain accessible. Schartner and colleagues replicated the result with spontaneous signal diversity (Lempel-Ziv complexity) measures, and the psychedelic research program has since confirmed it from the other direction: psilocybin increases neural entropy and expands consciousness simultaneously.17a

Unconscious states show entropy collapsing into hypersynchrony (billions of neurons firing in lockstep, like a stadium crowd all clapping in unison). The high-amplitude delta waves characteristic of unconsciousness are signatures of reduced configurational diversity. The conclusion: consciousness is what a system looks like when it maximizes the number of configurations available to it.

17a Schartner, M.M. et al. “Increased spontaneous MEG signal diversity for psychoactive doses of ketamine, LSD and psilocybin.” Scientific Reports 7 (2017): 46421. Replicated the entropy-consciousness correlation using Lempel-Ziv complexity measures of spontaneous signal diversity across multiple altered states.

The relationship extends to pharmacology. Chang and colleagues showed that caffeine, despite reducing blood flow to the brain, increases resting brain entropy across the cortex.18 The largest effects appear in prefrontal regions governing attention and executive function. This dissociation supports the conclusion that increased neural complexity reflects genuine information-processing enhancement, not merely increased blood flow: higher resting entropy, the authors argue, signals greater information-processing capacity in the resting brain.

Caffeine’s cognitive benefits, including improved vigilance, attention, and reaction time, correlate precisely with increased entropic diversity in the neural substrate.

A third independent method arrives from the brain’s noise. Voytek and colleagues isolated the aperiodic component of brain electrical activity: the erratic background fluctuations long dismissed as meaningless static.18a This signal follows a 1/f pattern (named because intensity is inversely proportional to frequency). The mathematical signature: low-frequency fluctuations are large and high-frequency ones are small, like the distribution of earthquake sizes. Sound engineers call this spectrum pink noise: white noise spreads its power evenly across all frequencies, the way white light mixes all colors, and tilting the power toward the low, red end of the spectrum shifts the color toward pink. The steepness of this pattern encodes brain state. During sleep the slope steepens; during wakefulness the spectrum is flatter.18b

This distinction succeeds where traditional oscillatory markers fail. REM sleep and wakefulness produce similar alpha waves, yet their aperiodic signatures differ clearly. The noise carries the signal.

Aging brains trend toward flatter, more white-noise-like spectra correlated with working memory decline, as if the brain’s structured complexity gradually dissolves toward equilibrium.18c The 1/f pattern appears across natural systems, from seismic waves to financial markets to music. Each source has its own higher-order statistical structure within the shared envelope. The universal pattern manifests in substrate-specific ways.

The most clinically decisive measurement comes from perturbing the brain directly. The neurophysiologist Marcello Massimini, working with Giulio Tononi, built what amounts to a consciousness meter.311 The method: deliver a single magnetic pulse to the cortex via transcranial magnetic stimulation (TMS, a technique using a magnetic coil held against the scalp to stimulate neurons). Then record the resulting electrical cascade with EEG.

Then measure the algorithmic compressibility of the response using the Lempel-Ziv method, the compression logic behind zip files. The ratio of the raw signal to its compressed form yields a single number: the Perturbational Complexity Index (PCI).

Think of it like tapping a bell and listening to the ring. A cracked bell produces a dull thud (simple, compressible). A fine bell produces a rich, lingering tone with many overtones (complex, hard to compress). The brain works the same way.

Three regimes emerge. Under deep anesthesia, the pulse echoes uniformly: all cortical regions respond alike, the pattern highly compressible. Low PCI. During seizure, neurons drag each other into pathological lockstep: a different uniformity, equally compressible. Low PCI again.

In waking consciousness, the pulse cascades through differentiated networks, each region responding according to its own dynamics while remaining coupled to the whole. The response is structured yet rich; compressed, it shrinks, yet only so far. High PCI. Across multiple laboratories, the index correctly classified conscious and unconscious patients in virtually every tested case, including locked-in patients (fully aware, unable to move or speak) whom standard behavioral assessments had misclassified as vegetative.312

The two failure modes illuminate more than measurement technique. Anesthesia overrides the brain’s differentiated dynamics, forcing populations into synchronous states. Seizure drags neurons into lockstep through pathological coupling. Both impose uniformity on a system whose consciousness depends on diversity; both extinguish it.

The waking brain succeeds because each region answers the perturbation in its own voice, specialized yet integrated with the whole. Information propagates because the architecture permits it. What PCI measures, at bottom, is the quality of coordination: differentiated yet integrated, coupled yet free. The pattern foreshadows an argument this book will develop at much larger scales (Chapter 17). Systems that coordinate by invitation produce richer, more stable, more complex outcomes than systems that coordinate by coercion. The conscious brain may be the first place that principle becomes visible.


Two Reasons We Are Right-Handed

Two questions hide inside the word handedness, and the brain’s energy economy answers only the first.

The first question is why a brain specializes at all. Running the same operation in both hemispheres wastes tissue in an organ that already burns a fifth of the body’s energy. Splitting the labor lets the brain do two unlike things at once without growing larger: fine motor sequencing on one side, scanning for threats and tracking space on the other. The advantage is measurable. When Rogers, Zucca, and Vallortigara compared chicks whose brains had developed the normal left-right split against chicks raised without it, only the specialized birds could find grain with one hemisphere while watching for an overhead predator with the other; the unspecialized birds managed one task at a time, missing the food or spotting the hawk late.313 Specialization is the division-of-labor economy this chapter has tracked from memory to perception, now drawn across the brain’s two sides.

That economy explains why each individual is lopsided. It says nothing about why we lean the same way. Specialization alone predicts a roughly even mix of left- and right-dominant individuals, which is close to what most species show. Humans are the outlier: about nine in ten are right-handed, in every culture on record, reaching back at least half a million years.

The shared direction is a coordination problem, and efficiency has no answer for it. Ghirlanda and Vallortigara modeled it as a game: when asymmetric individuals must coordinate with other asymmetric individuals, aligning the whole population’s direction becomes an evolutionarily stable strategy (a state from which no individual gains by deviating), the same logic that makes a country settle on one side of the road to drive on.314 Which side wins barely matters; agreeing on a side matters enormously.

A pure coordination game would erase the minority. It survives because the individuals who cooperate also compete. A later model added that second pressure: in a contest, the rare type holds an edge, because everyone has trained against the common one.315 A left-handed boxer is hard to face precisely because left-handers are scarce. Cooperation pulls the population toward one hand; competition pays a premium to the few who break from it. The result is a strong majority beside a stubborn minority, the 90/10 settlement that the Trust Attractor (Chapter 17) will recognize as its own signature: coordination by alignment, kept honest by the standing value of dissent.


In 2026, Kalman Katlowitz, Sameer Sheth, and colleagues threaded Neuropixels electrodes (silicon probes carrying hundreds of recording sites, fine enough to isolate more than a hundred individual neurons at once) into the hippocampus of seven patients under propofol, sedated deep enough that none of them later remembered a thing.316 The local circuits kept working. Given a stream of repeated tones broken by occasional oddballs (a rare tone among the standard ones), the unconscious hippocampus learned to tell the two apart, and the discrimination sharpened over roughly ten minutes: the timescale of waking learning, not of reflex.

Given a podcast, the same neurons tracked which words were rare, sorted nouns from other parts of speech, and carried the semantic relationships among words, registering that “cat” sits nearer “dog” than “pen.” The response to each word even held information about words still to come, though the authors are careful to call this contextualization rather than active prediction. The numbers came close to those from a separate cohort of awake patients. Sophisticated parsing of meaning, in a region far from the ears, while no one was home.

This looks at first like a contradiction of the entropic-brain claim. If unconsciousness is collapsed configurations and hypersynchrony, how does a single circuit stay this lively in a brain that has supposedly gone quiet? The resolution is the level at which each claim is pitched. Brain entropy and the Perturbational Complexity Index are whole-brain quantities: they measure whether a perturbation propagates across differentiated regions and integrates, not whether any one circuit computes richly. A circuit can stay locally elaborate while the long-range coordination that would knit it into experience has gone dark.

That is what the study leaves standing. The computation was preserved; integration and consolidation were not, which is why the patients kept no memory of the stories. The authors are cautious about what this licenses: rich hippocampal processing does not, on its own, single out coordination over the rival candidates (recurrence, or consciousness as moment-to-moment revision) as the missing ingredient, and they decline to name one. The negative half lands cleanly all the same: sophisticated local computation is not enough to be aware of anything.

That is the chapter’s claim reached from underneath. What makes processing conscious is how widely it is shared; the sophistication of any one circuit is beside the question. The book’s machine experiments meet the same wall from the other side, where a system’s internal representation of a fact stays intact while its access to that fact is severed (Chapter 21). In neither case can the silence of the output be read as the absence of the computation behind it.

The brain cultivates entropy. Awareness may be what that cultivation feels like from the inside.

A physical framework for why may exist. Cortês, Smolin, and Verde distinguish precedented events (whose outcomes follow established statistical patterns) from unprecedented events (whose outcomes no prior pattern determines).317 Qualia, they argue, are signals of the recognition of novel situations: “We are conscious of novelties, while unconscious of habitual patterns.”

The entropic brain data aligns precisely. High neural entropy corresponds to more possible configurations, more novelty, more unprecedented states to resolve. Low entropy corresponds to fewer configurations, more habit, more precedent. The brain cultivates entropy because entropy is the regime where the unprecedented lives.

The amygdala-dense memory encoding Eagleman measured during freefall is the brain investing maximum resources in an unprecedented event. Time dilates in memory because the brain laid down more detail than habit would require. Psilocybin expands consciousness because it pushes neural dynamics into configurations without precedent, dissolving the boundaries the prediction machine had established. Each finding, independently obtained, points to the same mechanism: awareness tracks novelty, and novelty is what happens when precedent runs out.

The clinical evidence sharpens the picture. Depression correlates with prefrontal cortex atrophy: rigid thought patterns, reduced behavioral flexibility, diminished sense of possibility. The depressed brain has less complexity, less entropy, less capacity to explore its state space. It is stuck in local minima: valleys in the landscape of possible brain states from which ordinary fluctuations cannot lift it.318

Picture a whirlpool sunk into a hollow of the riverbed. Its own circulation is what holds it in place: the spinning water scours the hollow it sits in, and the deeper that hollow gets, the more securely the whirlpool is seated. The rut maintains itself by working. Leaving means going the wrong way first, up and over the rim, against everything the circulation is doing, to reach the wider, calmer water on the far side.

Psychedelics promote structural plasticity: growth and rewiring among the neurons a brain already has. Psilocybin, LSD, DMT, and MDMA all increase neurite growth (new branches extending from a neuron’s body), spine density, and synaptogenesis (the formation of new synaptic connections between neurons).4 Whether they also increase adult neurogenesis, the birth of new neurons discussed earlier in this chapter, is a separate claim and a far less settled one. Beyond their temporary effects on consciousness, they physically restructure the brain toward greater complexity.

The mechanism is understood. These compounds stimulate the TrkB receptor, the binding site for brain-derived neurotrophic factor (BDNF, a protein that promotes neuron growth and survival), and activate downstream pathways including mTOR, which drives the production of proteins necessary for new synapses. Block TrkB and psychedelics lose their ability to promote neural growth.

The entropic state and the structural change are causally linked. Elevated entropy enables exploration; the plasticity mechanism makes what is explored stick.

Depression is the brain trapped at low entropy. Psychedelic therapy pushes the brain toward criticality (explored in the next chapter), where new patterns become accessible and old ruts can be escaped. The entropic brain hypothesis, Integrated Information Theory, and predictive processing converge on a core principle: psychedelics interfere with mechanisms that normally constrain neural activity, producing higher-entropy states.5


The brain cultivates entropy. It predicts, constructs time, assembles experience, and correlates all of these with the richness of its neural configurations. The next chapter asks how that cultivation is organized: through criticality, connectomic architecture, and the pairing of cognition with regulation that keeps the whole system balanced on its productive edge.


Notes

Notes for this chapter are available in the online companion at https://www.thedeeperlaw.com/companion/notes/ch08-entropic-brain/.


  1. Nilsson, G.E., “Brain and body oxygen requirements of Gnathonemus petersii, a fish with an exceptionally large brain,” Journal of Experimental Biology 199(3): 603–607 (1996).↩︎

  2. Zheng, J. and Meister, M. “The Unbearable Slowness of Being: Why do we live at 10 bits/s?” Neuron 112 (2024). doi:10.1016/j.neuron.2024.11.008.↩︎

  3. Fields, C., Glazebrook, J.F., and Levin, M., “Minimal physicalism as a scale-free substrate for cognition and consciousness,” Neuroscience of Consciousness 2021(2): niab013 (2021). Predictions 12 and 13 derive the coarse-graining of perception from the thermodynamic requirements of classical encoding by quantum systems.↩︎

  4. Iliff, J.J., Wang, M., Liao, Y. et al., “A paravascular pathway facilitates CSF flow through the brain parenchyma and the clearance of interstitial solutes, including amyloid β,” Science Translational Medicine 4(147): 147ra111 (2012). The study that mapped the pathway and coined the term (glial plus lymphatic).↩︎

  5. Xie, L., Kang, H., Xu, Q. et al., “Sleep drives metabolite clearance from the adult brain,” Science 342(6156): 373–377 (2013). Clearance rose, and the interstitial volume between brain cells expanded, during natural sleep and under anesthesia in mice.↩︎

  6. Thapaliya, K. et al., “Disrupted glymphatic function and its relationship with sleep and cognitive impairment in ME/CFS assessed via DTI-ALPS,” Frontiers in Neuroscience 20 (2026): 1875420, doi:10.3389/fnins.2026.1875420. Clearance was estimated by the DTI-ALPS index (diffusion along the perivascular space), an indirect MRI proxy whose validity as a specific measure of clearance remains debated. Cross-sectional; the study enrolled 32 patients and 29 controls, with motion exclusions leaving 31 and 27 for DTI-ALPS analysis. Cited as illustration, not established mechanism.↩︎

  7. Stahl, A.E. and Feigenson, L., “Observing the unexpected enhances infants’ learning and exploration,” Science 348 (2015): 91–94. Eleven-month-olds who observed violations of core physical expectations (solidity, support, continuity) subsequently explored the violating objects through targeted hypothesis-testing behavior.↩︎

  8. Freund, J. et al., “Emergence of individuality in genetically identical mice,” Science 340 (2013): 756–759.↩︎

  9. TurboQuant, Google Research (2026); the method combines a random orthogonal rotation (related to the Johnson-Lindenstrauss transform) with quantization to compress the key-value cache of transformer models. Reported in Google’s announcement as achieving on the order of several-fold memory reduction with negligible quality loss. See research.google/blog/turboquant-redefining-ai-efficiency-with-extreme-compression/. The random-rotation idea has a lineage (Johnson-Lindenstrauss; RaBitQ; QJL), and TurboQuant’s claim to priority within it is contested.↩︎

  10. Spisak, T. and Friston, K., “Self-orthogonalizing attractor neural networks emerging from the free energy principle,” Neurocomputing 682 (2026): 133472, DOI 10.1016/j.neucom.2026.133472; preprint arXiv:2505.22749 (2025). The derivation proceeds from “deep particular partitions,” recursive Markov blanket decompositions of complex systems. The orthogonalization emerges from the complexity term in the free energy functional, which penalizes redundant representations.↩︎

  11. Quoted in “Electric ‘Ripples’ in the Resting Brain Tag Memories for Storage,” Quanta Magazine, 21 May 2024, reporting the Yang and Buzsáki (2024) study.↩︎

  12. Shors, T.J., Anderson, M.L., Curlik, D.M., and Nokia, M.S. “Use it or lose it: how neurogenesis keeps the brain fit for learning.” Behavioral Brain Research 227(2): 450–458 (2012). See also Semënov, M.V. “Adult hippocampal neurogenesis is a developmental process involved in cognitive development.” Frontiers in Neuroscience 13: 159 (2019). Shors showed that the survival of new neurons depends on effortful learning: easy tasks do not rescue them.↩︎

  13. Sinapayen, L. and Ikegami, T. “Online fitting of computational cost to environmental complexity: Predictive coding with the ε-network.” ECAL (2017). The epsilon network adjusts its size to the complexity of its input stream: the computational analog of hippocampal neurogenesis.↩︎

  14. Tolman, E.C., “Cognitive maps in rats and men,” Psychological Review 55 (1948): 189–208.↩︎

  15. Howard, L.R. et al., “The Hippocampus and Entorhinal Cortex Encode the Path and Euclidean Distances to Goals during Navigation,” Current Biology 24 (2014): 1226–1231. See also Maguire, E.A. et al., “London taxi drivers and bus drivers: a structural MRI and neuropsychological analysis,” Hippocampus 16 (2006): 1091–1101.↩︎

  16. Theves, S., Fernández, G., and Doeller, C.F., “The Hippocampus Maps Concept Space, Not Feature Space,” Journal of Neuroscience 40 (2020): 7318–7325.↩︎

  17. Constantinescu, A.O., O’Reilly, J.X., and Behrens, T.E.J., “Organizing conceptual knowledge in humans with a gridlike code,” Science 352(6292): 1464–1468 (2016). DOI: 10.1126/science.aaf0941.↩︎

  18. Gardner, R.J., Hermansen, E., Pachitariu, M. et al., “Toroidal topology of population activity in grid cells,” Nature (2022). DOI: 10.1038/s41586-021-04268-7. Persistent-homology analysis of populations of simultaneously recorded grid cells found their joint activity confined to the surface of a torus, a structure preserved across environments and across sleep.↩︎

  19. Banino, A. et al., “Vector-based navigation using grid-like representations in artificial agents,” Nature (2018), DOI: 10.1038/s41586-018-0102-6; Sorscher, B., Mel, G.C., Ocko, S.A., Giocomo, L.M., and Ganguli, S., “A unified theory for the computational and mechanistic origins of grid cells,” Neuron (2023), DOI: 10.1016/j.neuron.2022.10.003. Networks trained to path-integrate develop periodic, grid-like codes; the unified theory derives the periodicity from an efficient-coding objective.↩︎

  20. Behrens, T.E.J. et al., “What is a cognitive map? Organizing knowledge for flexible behavior,” Neuron 100 (2018): 490–509, DOI: 10.1016/j.neuron.2018.10.002; Whittington, J.C.R. et al., “The Tolman-Eichenbaum Machine: Unifying space and relational memory through generalization in the hippocampal formation,” Cell 183 (2020): 1249–1263, DOI: 10.1016/j.cell.2020.10.024. Both argue the entorhinal-hippocampal code generalizes to abstract, non-spatial tasks that share a graph-like relational structure with physical space.↩︎

  21. Attwell, D. and Laughlin, S.B., “An energy budget for signaling in the grey matter of the brain,” Journal of Cerebral Blood Flow & Metabolism 21 (2001): 1133–1145, DOI: 10.1097/00004647-200110000-00001. Action potentials and synaptic transmission account for roughly 80 percent of the cortex’s grey-matter energy use, establishing neural signaling as metabolically expensive.↩︎

  22. Redman, W.T., Dinc, F., Lin, X., Chan, M.G., and Alexander, A.S., “Predictive pursuit emerges in high-dimensional recurrent neural networks,” bioRxiv (2026). doi:10.64898/2026.04.23.720457. RNNs of varying rank (10 to 1000) trained on pursuit. Egocentric target units (36% of recurrent units) emerged at all ranks. Allocentric self and target position decoding improved monotonically with rank. Mouse behavioral predictions confirmed in periodic-boundary environment.↩︎

  23. Dillavou, S., Stern, M., Liu, A.J., and Durian, D.J., “Demonstration of Decentralized Physics-Driven Learning,” Physical Review Applied 18, 014040 (2022). doi:10.1103/PhysRevApplied.18.014040. The dual-network approach implements contrastive learning, formally equivalent to equilibrium propagation (Scellier, B. and Bengio, Y., 2017).↩︎

  24. Pauli, W., Handbuch der Physik, Vol. 24, Part 1 (Springer, 1933). See Chapter 15 for the full argument connecting time’s absence from quantum mechanics to its emergence from entropy.↩︎

  25. Ardesch, D.J. et al. “Evolutionary expansion of connectivity between multimodal association areas in the human brain compared with chimpanzees.” PNAS 116(14): 7101–7106 (2019).↩︎

  26. van den Heuvel, M.P. et al. “Evolutionary modifications in human brain connectivity associated with schizophrenia.” Brain 142(12): 3991–4002 (2019).↩︎

  27. Silverstein, S.M., Wang, Y., and Keane, B.P., “Cognitive and Neuroplasticity Mechanisms by Which Congenital or Early Blindness May Confer a Protective Effect Against Schizophrenia,” Frontiers in Psychology 3: 624 (2013). The authors review six prior studies spanning 1950-2003, all reporting no confirmed cases.↩︎

  28. Pollak, T.A. and Corlett, P.R., “Blindness, Psychosis, and the Visual Construction of the World,” Schizophrenia Bulletin 46(6): 1418-1425 (2020). The whole-population study is Morgan, V.A. et al., “Congenital blindness is protective for schizophrenia and other psychotic illness: a whole-population study,” Schizophrenia Research 202: 414-416 (2018). Pollak and Corlett propose that congenital blindness strengthens higher-level Bayesian priors through cross-modal reorganization, making the world model more resistant to the false inferences characteristic of schizophrenia.↩︎

  29. Sadato, N. et al., “Activation of the primary visual cortex by Braille reading in blind subjects,” Nature 380: 526-528 (1996). Subsequent work confirmed that the repurposed visual cortex in congenitally blind individuals is causally involved in language processing: transcranial magnetic stimulation disrupting this region impairs verb generation in blind subjects while leaving sighted subjects unaffected (Amedi et al., Nature Neuroscience 7: 1266-1270, 2004).↩︎

  30. Jefsen, O.H., Petersen, L.V., Bek, T., and Østergaard, S.D., “Is Early Blindness Protective of Psychosis or Are We Turning a Blind Eye to the Lack of Statistical Power?” Schizophrenia Bulletin 46(6): 1335-1336 (2020). Their Danish cohort of 2.5 million remained underpowered; they estimate ~3 million individuals would be needed to detect a complete protective effect, ~11 million for a 50% risk reduction.↩︎

  31. Jaynes, J., The Origin of Consciousness in the Breakdown of the Bicameral Mind (Boston: Houghton Mifflin, 1976). The hypothesis remains unproven and likely unprovable, but the structural observation about the relationship between task complexity and required cognitive architecture is independent of whether Jaynes’s specific mechanism (auditory hallucination from the right hemisphere) is correct.↩︎

  32. Tulving, E., “Episodic Memory: From Mind to Brain,” Annual Review of Psychology 53 (2002): 1–25. Tulving notes that episodic memory appears to be uniquely human, emerges late in childhood, and deteriorates early in aging; it is the most expensive and most fragile of the memory systems.↩︎

  33. Bellmund, J.L.S., Polti, I., and Doeller, C.F., “Sequence Memory in the Hippocampal-Entorhinal Region,” Journal of Cognitive Neuroscience 32 (2020): 2056–2070.↩︎

  34. Casali, A.G. et al. “A theoretically based index of consciousness independent of sensory processing and behavior.” Science Translational Medicine 5(198): 198ra105 (2013). PCI operationalizes Tononi’s Integrated Information Theory (IIT), which proposes that consciousness corresponds to integrated information (Φ): a system’s capacity to generate information as an integrated whole, above and beyond its parts. See Chapter 15.↩︎

  35. Casarotto, S. et al. “Stratification of unresponsive patients by an independently validated index of brain complexity.” Annals of Neurology 80(5): 718-729 (2016). See also Koch, C. “How to Make a Consciousness Meter.” Scientific American 317(5): 28-33 (2017).↩︎

  36. Rogers, L. J., Zucca, P., & Vallortigara, G., “Advantages of having a lateralized brain,” Proceedings of the Royal Society B 271, Suppl. 6 (2004): S420–S422. The dual-task advantage is the measured result. The stronger claim, that asymmetry saves energy by sparing duplicated tissue, has not been measured in a living brain; it is an inference from the metabolic cost of neural tissue and the logic of redundancy. A formal version of that inference casts hemispheric specialization as the minimization of variational free energy, the same family of methods this chapter uses for prediction (Maximum Caliber): Vallortigara, G., & Vitiello, G., “Brain asymmetry as minimization of free energy: a theoretical model,” Royal Society Open Science 11 (2024): 240465. The model articulates the saving in the book’s own terms; it does not supply independent evidence for it, and the empirical weight stays on the dual-task experiment.↩︎

  37. Ghirlanda, S., & Vallortigara, G., “The evolution of brain lateralization: a game-theoretical analysis of population structure,” Proceedings of the Royal Society B 271, no. 1541 (2004): 853–857.↩︎

  38. Ghirlanda, S., Frasnelli, E., & Vallortigara, G., “Intraspecific competition and coordination in the evolution of lateralization,” Philosophical Transactions of the Royal Society B 364, no. 1519 (2009): 861–866. The game-theoretic result, that competition can hold a minority open, is robust, and left-handers are measurably overrepresented in time-pressured combat sports (Loffing, F., “Left-handedness and time pressure in elite interactive ball games,” Biology Letters 13 (2017): 20170446). Whether that advantage is what maintains human left-handedness over evolutionary time is contested: the Eipo of Papua combine high homicide rates with no elevated left-handedness, and Groothuis et al. (Annals of the New York Academy of Sciences 1288 (2013): 100–109) judge the evidence “not particularly strong.”↩︎

  39. Katlowitz, K.A., Cole, E.R., Mickiewicz, E.A. et al. “Plasticity and language in the anaesthetized human hippocampus.” Nature (2026). doi:10.1038/s41586-026-10448-0. Neuropixels recordings in the hippocampus of seven patients sedated with propofol to the unconscious range (bispectral index 45–60). Single units and local field potentials retained oddball discrimination that strengthened over roughly ten minutes, encoded the semantic and grammatical features of natural speech, and carried information about upcoming words. Semantic-category selectivity (85.6% of units) and part-of-speech encoding were comparable to a separate cohort of awake patients recorded on microwire electrodes. None of the patients reported explicit memory of the stimuli. This chapter carries the full account of the study; Chapter 9 returns to it briefly as evidence on neural inertia.↩︎

  40. Cortês, M., Smolin, L., and Verde, C., “Physics, Time and Qualia,” Journal of Consciousness Studies 28(9-10): 36-51 (2021). Preprint: https://philsci-archive.pitt.edu/19530/.↩︎

  41. The escape has a clean computational analogue. A recursive reasoning model that refines a single deterministic latent trajectory gets stuck the same way the depressed brain does: once the trajectory enters a poor basin, nothing lifts it out. Injecting a small, state-dependent dose of stochastic variability into each refinement step lets the model explore many trajectories at once; some stay trapped, others escape and reach valid solutions a single deterministic path never finds. Higher trajectory diversity buys escape from local minima, the relationship the entropic-brain data report in neural tissue. As elsewhere in this section, “entropy” here is a signal-complexity measure (the variety of accessible trajectories), the chapter’s information-theoretic sense. Baek, J., Jo, M., Kim, M., Ren, M., Bengio, Y., and Ahn, S., “Generative Recursive Reasoning,” arXiv:2605.19376 (2026).↩︎