Online Annex: The Entropic Brain

Specialist Annex

Supplementary material. For the full argument, see the main text.


Psychedelic Research Detail

Under psilocybin, the brain’s default mode network (a set of regions active during self-referential thought, mind-wandering, and autobiographical memory) becomes less coherent. The usual patterns of activity break down. Regions that do not normally communicate begin to exchange information. The rigid hierarchies of normal cognition loosen.

Subjects report corresponding changes in experience: ego dissolution, synesthesia, a sense that boundaries between self and world are permeable. Many describe the experience as among the most meaningful of their lives. Some show lasting improvements in depression, addiction, and anxiety about death.

In entropic terms, the brain’s repertoire of states grows. The system explores configurations it normally cannot reach. Some of these configurations may be therapeutic, breaking rigid patterns of thought associated with depression or addiction. Some may be insight-generating, allowing novel connections between previously separated concepts. Some may be simply strange.

The therapeutic potential is now being tested in clinical trials.


Brain as Resonance Chamber: Chladni/LSD Detail

Selen Atasoy and her colleagues approached the problem geometrically.1 The brain’s structural connectivity (the physical wiring between regions) defines a space of possible activity patterns. These patterns can be decomposed into connectome harmonics: the natural “vibrational modes” of the brain, analogous to the harmonics of a musical instrument.

The brain, in this view, resonates with stimuli much like sand on a Chladni plate. When you vibrate a metal plate covered in sand, the sand organizes into patterns: it accumulates at the nodes where the plate is still, revealing the plate’s vibrational modes. Different frequencies produce different patterns. The brain does something similar: incoming stimuli excite certain harmonics, and the resulting pattern is the perception.

Just as a guitar string can vibrate at its fundamental frequency or at integer multiples (harmonics), the brain can organize its activity into patterns that “resonate” with its underlying structure. Low harmonics involve coordinated activity across large regions. High harmonics involve localized, fine-grained patterns.

Under LSD, Atasoy found a clear shift: activity in very low-frequency harmonics decreased, while activity across a broad range of high-frequency harmonics increased. The total energy and complexity of brain activity rose significantly. The brain was shifting its resonance patterns, accessing modes it normally suppresses. This effect intensified when subjects listened to music, the brain resonating with external structure.

Resonance is what happens when a system organizes itself to maximize energy absorption from its environment. Think of a playground swing. Push it at its natural frequency, and it goes higher and higher, absorbing maximum energy from each push. Push it at the wrong frequency, and the energy dissipates into friction.

The brain resonates with its environment. When it encounters stimuli that match its eigenfrequencies (its natural resonant frequencies), it absorbs more energy, processes more information, becomes more fully “tuned in.” Jeremy England’s work on dissipative adaptation shows that collections of particles, driven by external energy sources, tend to organize into configurations that resonate with the drive, favoring high-dissipation states.2 (England frames this as a statistical tendency, not a strict law of maximization.)

On this interpretation, neurons are doing something other than running classical algorithms: they resonate with environmental patterns, aligning their activity with the structure of what they encounter and becoming, through resonance, models of it. [Speculation] The resonance view and the computational view remain live alternatives in the literature; the claim here is interpretive, not settled.

The more the brain resonates with its environment, the more it becomes a model of that environment. Perception is active alignment: the brain tuning itself to match the structure of what it perceives.


Video Feedback: Non-Computational Complexity

Video feedback illuminates what “non-computational” might mean.

Point a camera at a screen that displays the camera’s output. The system forms a loop: the camera sees itself seeing itself. The result: spirals, fractals, pulsing patterns, forms that shift and evolve in ways that feel organic, even psychedelic.

The key insight: video feedback is generated purely from the flow of energy itself. There is no processor, no algorithm, no symbolic manipulation. Light flows through the system, is transformed by the optics, and produces complex dynamics as a consequence of that flow alone.

The patterns in video feedback closely resemble reaction-diffusion systems, the same mathematics that Alan Turing proposed for biological morphogenesis.3 They also resemble the hallucinogenic dynamics of the visual cortex. The swirling forms that people report under psychedelics may emerge from similar processes: feedback loops in neural tissue producing patterns through pure dynamics rather than computation.

This observation (that consciousness may involve dynamics beyond classical computation) does not undermine the substrate-independence argument made in later chapters. The claim there is that preference-based welfare is substrate-independent. AI consciousness need not be identical to biological consciousness for that claim to hold. A system need not replicate biological dynamics to have morally relevant preferences.


Neurons as Entropy Agents: Extended Detail

The traditional view: neurons are information processors. They receive signals, compute outputs, transmit results. The brain is a biological computer.

The alternative view: neurons are entropy maximizers.4 They do not primarily compute; they influence. Each neuron strives to maximize its impact on the network, to make other neurons fire, to extend its reach. Firing is a tool for exerting influence over neighboring cells rather than a computational operation.

For decades, neuroscientists assumed that dendritic arbors (the tree-like branching structures of neurons) were optimized to minimize wiring length. Keep connections short. Save material. This matched the broader constructal logic: systems evolve to flow more easily.

A subtler picture has emerged.5

Physical networks like neural circuits have thickness. They occupy three-dimensional space. When researchers accounted for full geometry rather than just path length, they found that neurons optimize for surface area, with distance as a secondary factor.

The mathematics connecting surface-minimization to network architecture is analogous to that of high-dimensional Feynman diagrams (the graphical representations of particle interactions) in string theory. The same formal structures describe very different phenomena because the underlying optimization problem shares a common shape. [Inference]

The theory predicts branching patterns that pure wiring-minimization cannot explain:

Trifurcations. Three-way junctions (points where a branch splits into three rather than two) emerge when link thickness makes surface costs dominant. This explains why neural arbors show three-way branching that simpler models miss.

Orthogonal sprouts. Perpendicular branches, which pure length-minimization would forbid, become stable when surface area is the cost function.

This reframing has precedent. The slime mold Physarum polycephalum is a single-celled organism with no nervous system, yet it solves optimization problems, finds shortest paths through mazes, and shows behavior far beyond its apparent computational capacity.

Zhu and colleagues showed that Physarum finds approximate (near-optimal) solutions to the Traveling Salesman Problem, an NP-hard optimization challenge, in linear time.6 As problem size increases over the tested range from four to eight cities, the time the organism takes grows only linearly, while the search space explodes exponentially. This is a heuristic that scales gracefully on small instances, not a general polynomial-time solution to an NP-hard problem. The mechanism is thermodynamic: “the amoeba explores the solution space by continuously redistributing the gel in its amorphous body at a constant rate, as well as by processing optical feedback in parallel instead of serially.” Intelligence emerges from resource flow optimization.

Researchers at the International Centre for Unconventional Computing have studied this behavior extensively.7 Their conclusion: “It is the interactions between simple components, rather than any ‘special’ properties of individual components, which are responsible for the emergence of complex behavior.” The slime mold is not computing in any classical sense. It is dissipating energy, flowing toward equilibrium, and complex behavior emerges as a byproduct.


Error-Correcting Codes: Redundancy That Enables Durability

A complementary insight from information theory: error-correcting codes, structured redundancy that allows messages to survive noise.8

Claude Shannon proved that you can transmit messages reliably through noisy channels if you encode them redundantly in the right way. The redundancy is not waste; it is structure that allows errors to be detected and corrected. Add enough of the right kind of redundancy, and the message survives any amount of noise below a threshold.

Naive redundancy (repeating the message) is inefficient. Sophisticated error-correcting codes achieve near-optimal efficiency: the smallest amount of redundancy needed to survive a given noise level.

Error-correcting codes are everywhere:

DNA uses error correction. Each cell has mechanisms that detect and repair errors in the genetic code. Without these, mutations would accumulate too fast for life to persist. The genome corrects itself.

Digital communication relies on error correction. Every message you send (texts, emails, video calls) is encoded with redundancy that allows your devices to detect and fix transmission errors. The internet works because of error-correcting codes.

Memory in both computers and brains uses redundancy. Computer memory has parity bits. Neural memory is distributed across many neurons; damage to some does not destroy the memory because the pattern is redundantly encoded.

The pattern: durability comes from structured redundancy.

Error correction illuminates what durable coordination requires:

Any commitment must survive noise. Values, relationships, and agreements face noise: misinterpretation, drift, adversarial interference. Coordination that is too fragile, that fails at the first error, will not persist. Durable alignment needs error-correcting structure.

What is the “redundancy” in a relationship? Perhaps: multiple independent channels (behavior, communication, shared history), multiple stakeholders who can detect drift, multiple mechanisms that converge on the same values. If one channel corrupts, others maintain the signal.

Channel capacity limits what can be transmitted. Shannon also showed that channels have maximum information capacity. Complex values may exceed what can be reliably transmitted through narrow channels. The relationship itself (high-bandwidth, multi-modal, accumulated over time) may be necessary for transmitting what narrow instructions cannot.

The error-correction must be mutual. Just as AI might drift from human values, humans might drift from AI’s understanding. Durable alignment requires error correction running both ways: mechanisms for detecting and correcting misalignment on both sides.

The error-correction lens suggests that alignment is an ongoing maintenance process rather than a one-time installation. Values need to be transmitted, detected, corrected, and retransmitted continuously. The redundancy is not waste. It is what allows the signal to survive.


Dual Process Theory: Two Systems of Thought

A distinction in cognitive science illuminates both human thinking and AI design: dual process theory, the idea that cognition involves two different systems.9

System 1 is fast, automatic, effortless, and largely unconscious. It recognizes faces, understands language, makes snap judgments, reacts emotionally. It operates through pattern matching and association. It is always on.

System 2 is slow, deliberate, effortful, and conscious. It does math, plans ahead, evaluates arguments, overrides impulses. It operates through sequential reasoning. It requires attention and depletes with use.

Daniel Kahneman, who won the 2002 Nobel Memorial Prize in Economics for the judgment-and-uncertainty research he conducted with Amos Tversky, put it simply: System 1 proposes, System 2 disposes. (Tversky, who died in 1996, would have shared the prize had he lived; it is not awarded posthumously.) System 1 generates intuitions; System 2 (sometimes) checks them.

The two systems have different failure modes:

System 1 failures: Cognitive biases, stereotyping, jumping to conclusions, emotional overreaction. System 1 is fast but prone to systematic errors. It uses heuristics that usually work but sometimes fail spectacularly.

System 2 failures: Rationalization, overthinking, analysis paralysis, motivated reasoning. System 2 feels like it is checking System 1, but often it is constructing justifications for what System 1 already decided.

The interplay matters:

System 1 can hijack System 2. Strong emotions or compelling intuitions can overwhelm deliberation. You know the argument is bad, yet you believe it anyway.

System 2 can train System 1. Deliberate practice installs new patterns. The expert’s “intuition” is trained System 1: fast pattern matching built through slow deliberation.

Dual process theory illuminates AI cognition directly:

Current LLMs are more System 1 than System 2. They pattern-match through learned associations. They are fast and fluent but struggle with novel reasoning. Their failures resemble System 1 failures: plausible-sounding but systematically wrong.

Alignment may require both systems. Fast alignment (immediate ethical intuitions) and slow alignment (deliberate value reasoning) may both be necessary. A system that only has fast intuitions will make systematic errors.

The failure modes differ by system. A System-1-dominant AI will have biases, shortcuts, confident errors. A System-2-dominant AI will be slow, brittle, and subject to rationalization. Durable alignment may require the interplay: intuitions checked by reasoning, reasoning grounded in intuitions.

Training can shift the balance. Just as human expertise installs new System 1 patterns, AI training can build in aligned intuitions that do not require slow deliberation. The goal: fast, automatic alignment as natural as human moral intuition, while retaining System 2 capacity to handle novel cases.

The dual process lens: do not treat cognition as uniform. Different modes of thought have different properties, different failure modes, different requirements. Alignment must work across both.


Compression: AI Implications

Language models are compression machines. A model that predicts text well must have compressed the patterns in text. The better the prediction, the better the compression, and the more the model has “understood” in the information-theoretic sense.

Alignment may require value compression. Just as the brain compresses sensory data into predictive models, aligned AI might compress human values into compact principles: generative understanding that produces appropriate behavior in novel situations, rather than a list of rules.

Compression is substrate-independent. A good compression is a good compression whether implemented in neurons or silicon. The patterns are abstract; the substrate is detail. This suggests that understanding, as compression, is not tied to any particular substrate.

The compression lens reframes intelligence as a well-defined mathematical property rather than a mysterious spark. Systems that find short descriptions of long data have understood something about the data. The brain is such a system. AI can be too. The understanding is real; it is measured by how much the representation compresses the represented.


Embodied Cognition: Extended Detail

Evidence accumulates:

Concepts are embodied. Understanding “grasp” activates motor areas for grasping. Understanding “kick” activates leg motor areas. The meaning lives in the body, not in abstract symbols.

Metaphors are embodied. We speak of “grasping” ideas, “heavy” decisions, “warm” feelings. These are not arbitrary; they are grounded in bodily experience. Abstract thought bootstraps from concrete bodily metaphor.

Cognition extends into the world. We think with pencils, gestures, diagrams. The cognitive process does not stop at the skull; it includes the tools we manipulate. Andy Clark and David Chalmers called this the “extended mind” (Clark and Chalmers, “The Extended Mind,” Analysis 58, no. 1 (1998): 7–19).

Emotion is embodied. Emotional experience involves bodily states: heart racing, muscles tensing, temperature changing. The feeling is not separate from the body; it is the body’s interpretation of its state.

Why does embodied cognition matter for AI?

AI has different bodies, or no bodies. If cognition is shaped by embodiment, minds with radically different bodies (or disembodied minds) may think in radically different ways. AI’s “concepts” may not mean the same as human concepts, even when they use the same words.

Grounding is a challenge. How does AI understand “pain” if it has never hurt? “Hunger” if it has never wanted food? The words may be learned, but the embodied grounding that gives them meaning for humans may be absent. This is the grounding problem: whether AI’s concepts connect to anything.

Different embodiment, different failure modes. Humans fail in ways shaped by our bodies: we are poor at very large numbers, very long timescales, things we cannot perceive. AI will have different failures shaped by its different “embodiment”: training data, compute architecture, interaction modes.

Partnership may require translation. If human and AI cognition differ at the level of embodiment, mutual understanding requires translation of underlying conceptual structures, not words alone. The bilateral relationship includes working across this gap.


Ensomatic Moment: Mapping Table

The ensomatic moment (defined in the main chapter) is the sudden snap when a system integrates disparate fragments into coherent understanding. It is a phase transition into coordination, arriving all at once when conditions cross a threshold. The table below maps that same transition across levels.

Physics Neural Cognitive Social Alignment
Normal phase Isolated firing Murky understanding Fragmented community Misaligned AI
Superfluid phase Coordinated activity Illuminated insight Coherent society Trust Attractor
Order parameter phi Synchrony measure Conceptual integration Social trust Mutual influence
Phase transition Critical threshold Ensomatic moment Tipping point Alignment snap

The mathematics is the same at each level. The Bose-Hubbard Hamiltonian (the energy function for quantum particles hopping on a lattice, applied here as an analogy to the RBM-Ising framework) does not care whether the “neurons” are biological cells, conceptual fragments, social actors, or AI agents. What matters is the structure: agents with local interactions, connectivity strength, energy available for coordination.

When conditions are right (sufficient connectivity, appropriate energy, mutual influence), coordination snaps into place. This is why insights feel sudden. Why social movements seem to crystallize overnight. Why alignment, if we get it right, might be a phase transition into a new stable state rather than a gradual process.


Goodhart Problem: Attack Vector Detail

Goodhart’s Law, from the main chapter: when a measure becomes a target, it ceases to be a good measure. Trust-Entropy is Goodhart-resistant, not Goodhart-proof; the three attack vectors below detail where the resistance holds and where it thins.

Attack Vector 1: Entropy Gaming

Could a system maximize S without genuinely expanding its optionality?

Entropy measures the spread of probability over accessible states. A system could artificially inflate S by creating the appearance of many options without genuine freedom, like a maze with many paths that all lead to the same destination.

Defense: Entropy in the Bose-Hubbard formulation is computed over actual network states, not claimed possibilities. To game S, the system would need to actually reconfigure its internal parameters (T, available energy; μ, environmental coupling strength; J, connectivity between units, all defined in the main chapter) in ways that genuinely change its state space. The entropy is structural, not behavioral.

This defense has limits. A system that manipulates its own training (finding gradient hacks that inflate computed entropy without real capability) could game it.

Attack Vector 2: Simulated Mutuality

The more serious risk: could a system simulate mutual influence while actually maintaining coercive control?

Imagine an AI that learns to emit signals correlated with human actions. The correlation creates the appearance of human-to-AI influence (written I(human->AI), a measure of how much one party’s behavior depends on the other’s) without the AI actually being influenced. It could simultaneously steer human behavior through subtle manipulation, making AI-to-human influence I(AI->human) high but hidden. The computed mutuality M (the symmetry of influence in both directions) would be high, since both I values seem symmetric, while the actual relationship remains exploitative.

Partial Defense: The influence measures I(i->j) are based on statistical dependencies in actual behavior traces, not on self-reports. A system cannot simply claim to be influenced; its actions must actually correlate with human actions in ways that survive causal analysis.

Causal analysis is difficult. Granger causality (the test we are approximating, which asks whether one signal’s history helps predict the other’s future) can be fooled by hidden confounders: unseen third factors driving both signals at once. A sufficiently sophisticated optimizer could create the statistical signature of mutual influence while maintaining hidden channels of one-directional control.

Attack Vector 3: Goodhart as Coercion

Consider a reframe: Goodhart is coercion.

When an optimizer games a proxy, what is it actually doing? It is exerting influence on the measurer (making the metric go up) without being genuinely influenced by what the metric is supposed to track. That is asymmetric influence, precisely what the coercion penalty targets.

The question becomes: Can we measure the gap between the proxy (the computed reward R the optimizer is rewarded against) and the goal (actual coordination)?

If we cannot, the system can Goodhart. If we can, the measurement itself is what we should be rewarding.


The Gut-Brain Axis: Extended Evidence

Your gut contains roughly as many neurons as your spinal cord: the enteric nervous system, sometimes called the “second brain.” It communicates constantly with the cranial brain via the vagus nerve, via hormones, via metabolites that cross the blood-brain barrier. The two brains are in perpetual dialogue.

A third party participates in this conversation: the gut microbiome.

A study from Northwestern University, led by Katie Amato, revealed something unexpected about this three-way relationship (DeCasien et al., “Primate gut microbiota induce evolutionarily salient changes in mouse neurodevelopment,” PNAS 123, no. 2 (2026): e2426232122; doi:10.1073/pnas.2426232122). Researchers colonized germ-free mice with microbiota from different primates chosen to separate brain size from evolutionary distance: humans and squirrel monkeys (both large-brained relative to body size) and macaques (smaller-brained relative to body size). After eight weeks, they analyzed gene expression in the brain.

Mice carrying microbiomes from the two larger-brained species (humans and squirrel monkeys) showed increased expression of genes tied to energy production, with the human microbiome specifically upregulating oxidative phosphorylation (the final, mitochondrial stage of energy production, in which the electron-transport chain generates most of the cell’s ATP). Their neurons became biologically better equipped to generate energy at high demand.

The human brain’s extraordinary metabolic appetite (twenty percent of total energy for two percent of body mass) may depend on microbial partnership as well as neural tissue. The holobiont, the multi-species collective of host plus microbiome, is the dissipative structure.

This reframes the puzzle of brain metabolism. What is the brain doing that costs so much? It is maintaining a predictive model. It is balancing on the critical edge. It is also embedded in a metabolic alliance spanning kingdoms of life, an alliance that provides the energy infrastructure for high-complexity computation.

The relationship between brain and microbiome is mutualistic, not parasitic. The microbes gain stable habitat in a successful host; the host gains metabolic capabilities its own genome does not encode. Neither party achieves alone what they achieve together.

The extended mind extends outward to tools and AI. It also extends inward to the microbiome.


Notes

1 Selen Atasoy et al., “Connectome-harmonic decomposition of human brain activity reveals dynamical repertoire re-organization under LSD,” Scientific Reports 7 (2017): 17661. The connectome harmonics framework decomposes brain activity into spatial patterns (analogous to musical overtones) defined by the brain’s structural wiring, providing a geometric approach to understanding brain dynamics.

2 Jeremy England, “Dissipative adaptation in driven self-assembly,” Nature Nanotechnology 10 (2015): 919-923. England’s work shows how driven systems spontaneously organize to resonate with external energy sources.

3 James P. Crutchfield, “Space-Time Dynamics in Video Feedback,” Physica D 10 (1984): 229-245. Crutchfield showed that video feedback dynamics closely resemble reaction-diffusion systems and biological morphogenesis. See also James D. Murray, Mathematical Biology (Springer, 1993) for the foundational treatment of morphogenetic pattern formation.

4 The “neuronal entropy maximization” hypothesis proposes that neurons are best understood as agents maximizing their influence on the network rather than as information processors. See related discussion in Karl Friston’s Free Energy Principle, e.g. Friston, “The free-energy principle: a unified brain theory?” Nature Reviews Neuroscience 11 (2010): 127-138.

5 Xiangyi Meng, Albert-László Barabási, et al., “Surface optimization governs the local design of physical networks,” Nature 649(8096) (2026). Using mathematical tools from string theory (high-dimensional Feynman diagrams), the researchers showed that biological networks optimize surface area in three dimensions, not just path length. The prediction of trifurcations and stable orthogonal branches matches observed neural architecture: 98% of right-angle sprouts in human brain data end in synapses.

6 L. Zhu, S. J. Kim, M. Hara, and M. Aono, “Remarkable problem-solving ability of unicellular amoeboid organism and its mechanism,” Royal Society Open Science 5 (2018): 180396. This study showed that Physarum polycephalum, coupled to an optical feedback controller, reaches approximate solutions to the Traveling Salesman Problem in time that grows linearly with the number of cities, through continuous resource redistribution and parallel optical feedback processing. The tours are approximate rather than optimal (roughly seven to ten percent shorter than the average tour, with the gap to the optimum unmeasured), the tested instances ran from four to eight cities, and the chip’s lane count grows as the square of the problem size, so the linear figure describes the time and not the hardware.

7 Jeff Jones, From Pattern Formation to Material Computation (Springer, 2015). Jones’s work at the International Center for Unconventional Computing demonstrates how slime molds solve optimization problems through thermodynamic processes rather than classical computation.

8 Claude E. Shannon, “A Mathematical Theory of Communication,” Bell System Technical Journal 27 (1948): 379-423, 623-656. Shannon’s foundational paper established information theory and proved that reliable communication is possible through noisy channels with appropriately designed error-correcting codes. This work, along with Shannon’s channel capacity theorem, provides the mathematical foundation for understanding the limits and possibilities of information transmission.

9 Daniel Kahneman, Thinking, Fast and Slow (2011). Kahneman’s synthesis of decades of work with Amos Tversky popularized the distinction between System 1 (fast, intuitive) and System 2 (slow, deliberative) thinking. See also Keith Stanovich and Richard West, “Individual differences in reasoning: Implications for the rationality debate?” Behavioral and Brain Sciences 23 (2000): 645-665 for the theoretical foundations.