Constructal Semantics: When Flow Learns to Mean
Specialist Annex
A river carries water. A nerve carries signals. A conversation carries meaning. All three are flow systems, and all three obey the Constructal Law: they evolve branching architectures that provide easier access to their currents (Chapter 3).
When you measure how they branch, the numbers differ. The exponent that governs how a parent channel relates to its daughters changes depending on what the channel carries. Blood vessels follow Murray’s law, with an exponent of 3. Electrical signals in neurons follow Rall’s law, with an exponent of 3/2. Molecular cargo hauled along microtubules follows yet another, around 2. Different currents carve different geometries into the systems that carry them.
This raises a question the Constructal Law has invited yet never formally answered. If meaning is a current (if brains, conversations, and cultures are flow systems optimized for interpretive throughput), then meaning should have its own exponent. Channels carrying semantically rich information should branch differently from channels carrying water, electricity, or metabolic cargo.
This annex presents five tests of that prediction: two that support it, one that provides circumstantial support, and two honest negatives.
What Is Semantic Information?
Before looking at neurons and molecules, we need a definition of meaning that physics can work with. Philosophers have debated what meaning is for millennia. Thermodynamics offers a shortcut.
Artemy Kolchinsky and David Wolpert, working at the Santa Fe Institute, proposed in 2018 that semantic information is the portion of a system’s correlations with its environment that is causally necessary for the system’s continued existence. Strip away those correlations, and the system dies. Keep them, and the system persists.
A bacterium swimming toward a nutrient gradient carries Shannon information about its surroundings: raw correlations between its internal state and the external chemical field. Some of those correlations are noise, incidental patterns that happen to exist yet play no role in the bacterium’s survival. The rest are semantic: the bacterium needs them, and if you removed them (by scrambling the receptor that reads the gradient), the bacterium would starve.
Semantic information is bounded above by the total Shannon mutual information (a system cannot use more correlations than it has) and bounded below by the thermodynamic cost of self-maintenance (a system must at minimum track whatever keeps it alive). It is a physical quantity, measurable in bits, grounded in counterfactual intervention. Remove the correlation and observe whether the system fails.
This definition is powerful because it applies at every scale. A cell’s receptor-ligand binding is semantic information. A cortical column’s representation of an edge in the visual field is semantic information. A culture’s shared understanding of reciprocity is semantic information. All are correlations causally necessary for the persistence of the system that carries them.
The Translation Cost
Rivers and brains both branch. Both follow the Constructal Law. Why should their exponents differ?
The answer lies in a cost that water does not pay.
When water flows from a tributary into a main channel, it merges. The water does not need to be translated. A liter from the left fork and a liter from the right fork combine without conversion loss. The cost is viscous friction, and Murray’s exponent of 3 is the geometry that minimizes it.
When meaning flows from one processing level to the next, it must be translated. Each level of a neural hierarchy (a retinal ganglion cell, a V1 edge detector, a V4 shape integrator) operates in its own reference frame: a set of measurement operators that define what counts as a signal and what counts as noise. The physicist Chris Fields and the biologist Michael Levin formalized this as a hierarchy of quantum reference frames (QRFs): each level calibrates raw input from below and writes a coarse-grained summary for the level above.
The critical insight: these reference frames are nonfungible. You cannot fully specify one in terms of another using a finite number of bits. When information crosses from one level to the next, the receiving level must possess compatible measurement operators, able to decode what the sending level encoded. This compatibility is physical, requiring specific synaptic architectures, specific receptor configurations, specific learned weights. It is, in thermodynamic terms, expensive.
The cost of this inter-level translation is absent from Murray’s equation for blood vessels and from Rall’s equation for electrical conduction. It is the new term in the cost function for semantic flow. When you minimize total cost (Landauer erasure for compression at each level, plus maintenance of the reference frames themselves, plus the nonfungible translation overhead at every boundary), the optimal branching geometry shifts. The predicted exponent for semantic channels falls in a range between 2 and 3, distinct from the exponents governing fluid, electrical, and cargo transport.
The river carries water through friction. The brain carries meaning through translation. Both optimize. The optimization produces different shapes because the costs are different.
Three Stories from the Data
Story 1: The Shape of Interpretation
A Purkinje cell in the cerebellum is one of the most elaborate neurons in the body. Its dendritic tree fans out like a coral, collecting input from hundreds of thousands of parallel fibers, integrating their signals into a single output that controls the timing of movement. A cerebellar granule cell, by contrast, is one of the simplest neurons in existence: a small body with two or three short dendrites, performing minimal integration.
The Strahler bifurcation ratio (RB) is a measure of how elaborately a tree branches. Hydrologists invented it in the 1940s to classify river networks. Higher RB means a more complex branching hierarchy. Neuroscientists have applied it to dendritic trees.
If branching architecture reflects only physical size (how much dendrite a neuron must pack into its territory), then RB should track dendrite length: bigger neurons, higher ratios. If constructal semantics is right, RB should instead track interpretive complexity: neurons that perform deeper integration, higher ratios, regardless of size.
Anna Vormberg and colleagues at the Max Planck Institute published Strahler analyses for six neuron types in 2017. The data tell a clear story. Granule cells (minimal integration, two firing modes at most) branch at RB = 2.23. Purkinje cells (deep integration, four firing modes, five estimated levels of computational processing) branch at RB = 3.12. Lobula plate tangential cells in the fly visual system (which integrate wide-field optic flow across hundreds of dendritic inputs into a single directional signal, six levels of computational depth) branch at RB = 3.77.
The critical test: does interpretive complexity predict RB after controlling for physical size?
Yes. The partial correlation between computational depth and RB, with dendrite length and branch-point count held constant, is r = 0.985 (p = 0.015). Dendrite length alone, without the complexity control, shows a weaker and nonsignificant correlation (rho = 0.71, p = 0.111). The shape of a neuron’s branching tree tracks what the neuron computes, independently of how large it is.
This is a small sample: six cell types, with limited degrees of freedom after controlling for two size variables. A follow-up analysis using 177 individual neuron reconstructions from NeuroMorpho.Org (30 per type, downloaded as SWC morphology files) confirmed the result with substantially more statistical power. The partial rank correlation between computational depth and RB, controlling for dendrite length and bifurcation count at the per-neuron level, is r = 0.237 (permutation p = 0.0016, 10,000 iterations). Within-type size tests show that RB is independent of dendrite length in four of six cell types: the branching ratio tracks what the neuron computes, regardless of how large it is.
The pattern it reveals is the constructal semantics prediction made visible. Neurons that carry richer meaning (deeper integration, more levels of coarse-graining, more translation between reference frames) develop more elaborate branching topologies. The architecture follows the current. The current is interpretation.
Story 2: The Thermodynamic Cost of Thinking
Eric Chaisson, an astrophysicist at Harvard, proposed in 2001 that you can track cosmic complexity with a single number: energy rate density (phim), the power flowing through a system per unit mass. Stars have a phim around 2 erg per second per gram. Plants, about 900. The human body, about 14,000. The human brain, about 143,000.
The sequence is a ladder. Each rung represents a system that processes more energy per gram than the one below. The ladder climbs from galaxies through stars through ecosystems through organisms through brains. Each step up corresponds to greater organizational complexity.
Where does AI inference hardware sit on this ladder?
An NVIDIA H100 GPU accelerator card, the workhorse of modern AI training and inference, consumes about 700 watts and weighs about 3.18 kilograms. Its energy rate density: 2.2 million erg per second per gram. Fifteen times higher than the human brain.
This places AI hardware roughly one order of magnitude above the brain on Chaisson’s complexity ladder, continuing the trend from stars through organisms.
The number conceals a subtlety. Energy rate density depends critically on where you draw the system boundary. For the bare silicon die alone (65 grams, most of the card’s power attributed to it), phim reaches 108, near a jet engine. For a full data center (100 megawatts, thousands of tons of building, cooling, and networking equipment), phim drops to 105, comparable to the brain. Three orders of magnitude of variation, depending on the boundary.
The right comparison, following Chaisson’s convention of organ-level measurement for the functionally relevant subsystem, is full accelerator card versus brain. At this boundary, the 15-fold excess is real. Whether it represents genuinely deeper semantic processing or merely the thermodynamic profligacy of silicon (transistors dissipate energy through resistive heating; neurons dissipate through ion-pump cycles, a fundamentally different mechanism) remains an open question.
One signature does discriminate. The H100 draws 50 to 100 watts at idle and 600 to 700 watts during active inference: a five-to-tenfold ratio. Semantic processing has a measurable thermodynamic cost at the hardware level, distinct from the baseline cost of keeping the chip alive. Whether phim further differentiates between tasks of varying interpretive depth (factual retrieval versus multi-step reasoning, for instance) has not been measured. It should be.
Story 3: Complexity Gets Cheaper
In 2023, a team led by Lee Cronin at the University of Glasgow published assembly theory in Nature. The assembly index of a molecule measures the minimum number of joining operations required to build it from basic building blocks. Glycine, the simplest amino acid, has an assembly index of 1. Cholesterol, 20. Taxol, the cancer drug synthesized by Pacific yew trees through forty enzymatic steps, has an assembly index of 30.
Assembly index is a measure of accumulated causal history: the number of constructive acts recorded in a molecule’s structure. In the language of this book, it is a proxy for semantic depth, the number of coarse-graining steps encoded in the architecture.
The constructal semantics prediction was simple: if evolved biosynthetic networks optimize energy allocation for structural complexity, assembly index should grow faster than metabolic cost. More complex molecules should become cheaper per unit of complexity. Dissipative efficiency, assembly index per ATP spent, should increase with pathway complexity.
The author’s exploratory E3 analysis tested that prediction on 18 biologically synthesized molecules, from glycine to taxol. The log-log regression of assembly index on ATP cost yielded a slope of 0.624, with a standard error of 0.129. The association differed from zero (p = 1.74 x 10-4) and the slope was below one (one-sided p = 0.005 versus a slope of one).
A slope below unity means that doubling the metabolic investment predicts only about 1.54 times the assembly index. Assembly per ATP therefore falls as ATP cost rises. The direct rank correlation between pathway length and assembly per ATP pointed the same way and was not significant (rho = -0.375, p = 0.125). The test rejects the increasing-efficiency prediction on this panel.
The category values remain suggestive rather than diagnostic. Glucose had the highest estimated ratio in this small panel, 5 assembly units for 8 ATP equivalents, while palmitic acid had 7 for 129. Those contrasts are sensitive to pathway boundaries, cellular context, repeated motifs, and hand-estimated assembly indices. A story about ancient optimization cannot be extracted from two category examples.
The assembly indices in this analysis are structural estimates inferred from molecular topology rather than mass-spectrometric measurements. Several ATP costs are textbook or inferential values for de novo biosynthesis and vary with pathway boundary and cellular context. A decisive test requires measured assembly indices, metabolic-flux data, uncertainty propagation, and controls for molecular size and repeated motifs. The failed exploratory prediction is useful because it says exactly what a stronger experiment must distinguish.
The Residual Stream: Semantic Flow in Silicon
These results gain a companion from artificial intelligence research.
A transformer language model processes text through a sequence of layers. At each layer, the attention mechanism reads from a shared channel called the residual stream (a running total of all the representations computed so far, like a river that accumulates tributaries), computes a correction, and writes the result back. The residual stream is the backbone of the architecture: information flows through it from input to output, enriched layer by layer.
The author’s ongoing residual-stream diagnostics work demonstrated that a lightweight probe trained on residual-stream activations at two-thirds depth of a frozen language model can distinguish correct from incorrect outputs with an AUROC of 0.836. The probe transfers to a different architecture (Llama 3.1 8B, trained on different data with a different tokenizer) with near-zero loss in accuracy (gap of 0.001). The signal resides exclusively in the residual stream; attention patterns carry none (AUROC 0.464, worse than chance).
Three features of this result map onto the constructal semantics framework.
First, the cross-architecture transfer is linear. A nonlinear projection adds nothing over a simple linear map. This is what the framework predicts: the constructal geometry of semantic flow is substrate-independent, preserved up to a linear transformation. Different transformers have different weights, vocabularies, and training histories, yet the negative space of failed retrieval (the signature of a residual stream that was not enriched by attention) occupies the same geometric region in every model tested.
Second, observation preserves the signal; intervention destroys it. Reading the probe at inference time (without changing the model’s weights) reduces confident-but-wrong outputs from 24.4% to 1.0%. Using the same signal to steer training gradients produces 35.0% confident-but-wrong, catastrophically worse than doing nothing. The constructal principle explains this asymmetry: semantic flow is optimized under conditions of quasi-static equilibrium. Observation that does not perturb the system preserves the information that makes observation valuable. Gradient-based intervention changes the landscape while the probe reads a stale map.
Third, the idle-to-active energy ratio (five to tenfold in the GPU hardware that runs these models) confirms that semantic processing has a measurable thermodynamic cost distinct from baseline computation, consistent with the nonfungible translation overhead the framework predicts.
The residual stream is a semantic flow channel. It branches (via attention heads writing in parallel), it integrates (via the skip connection that accumulates all contributions), and it optimizes (via training, which selects for architectures that enrich the stream most efficiently). The Constructal Law, applied to this channel, predicts the features Watson observed: substrate-independent geometry, observer-preserving diagnostics, and a thermodynamic premium for active inference.
Story 4: Measuring the Exponent Directly
The first three stories provide circumstantial evidence that semantic channels branch differently. Story 4 measures the exponent itself.
Murray’s law says that at every branch point in a vascular tree, the parent diameter and daughter diameters are related by a power law: dparentα = ddaughter1α + ddaughter2α. The exponent α = 3 for blood vessels. Rall’s law gives α = 3/2 for electrical conduction in dendrites. The constructal semantics prediction: neurons carrying richer semantic signals should have a higher α than neurons carrying simpler signals.
The MICrONS dataset (Bae and colleagues, 2024) provides the data to test this directly. It is a cubic millimeter of mouse visual cortex, reconstructed at synapse-level resolution: roughly 80,000 neurons with every branch point, every axon, every connection mapped in three dimensions.
From this volume, 500 neurons were selected: 250 excitatory projection neurons (pyramidal cells whose axons cross cortical areas, carrying integrated signals) and 250 inhibitory interneurons (basket cells, chandelier cells, and others whose axons stay local, carrying simpler signals). At each bifurcation, the parent and daughter radii were measured using the L2 distance transform, and the scaling exponent α was fitted by solving Murray’s equation numerically.
The result: 12,487 valid bifurcations across 500 neurons. Excitatory projection neurons have a median α of 1.68; inhibitory local neurons have a median α of 1.56. The difference is statistically significant (Mann-Whitney p < 10-7).
The effect size is small (Cohen’s d = 0.047), partly because the radius proxy has limited spatial resolution (each measurement averages over an 8-micrometer chunk of tissue). Higher-resolution measurements using full 3D mesh reconstructions would sharpen the estimate.
The distribution of α values tells a richer story than the median alone. It is right-skewed: the mean α is approximately 2.1 for both cell classes, with the upper quartile reaching 2.5 to 2.6. Most bifurcations sit near Rall’s electrical-transport regime (α ≈ 1.5), but a substantial fraction extends into the predicted semantic regime (α between 2 and 3). This bimodal character is consistent with the framework’s prediction: most branch points in a neuron serve electrical conduction, while a subset, particularly where information is integrated across multiple inputs, operates in the semantic regime.
Story 5: A Test That Failed Honestly
The constructal semantics prediction has a corollary: information compression should increase across the cortical visual hierarchy. Primary visual cortex (V1) represents the world in fine detail: edges, orientations, spatial frequencies. Higher areas represent it more abstractly: objects, categories, scenes. If each level compresses the representation from the level below, the total information carried per neuron should decrease as you ascend the hierarchy, and the form of that decrease should follow a power law.
The Allen Brain Observatory Neuropixels dataset (Siegle and colleagues, 2021) provides simultaneous recordings from thousands of neurons across eight levels of the mouse visual hierarchy, from the thalamic relay (LGd) through primary visual cortex (VISp) to the highest cortical area (VISam). Natural images were presented roughly 50 times each (119 distinct images, approximately 5,950 presentations per session), providing enough repetitions to separate stimulus-driven signal from trial-to-trial noise.
For each neuron, the stimulus-conditioned mutual information I(R; S) was computed: how much does observing this neuron’s response tell you about which image was shown? This measure isolates the portion of neural variability that carries information about the world from the portion that is noise. It was computed using the Nemenman-Shafee-Bialek entropy estimator, which corrects for the bias that arises when you have only 50 trials to estimate a probability distribution.
Across 1,746 neurons in five recording sessions: no significant trend. Primary visual cortex (VISp) carries the most stimulus information (0.146 bits per neuron), as expected for the area most directly driven by visual input. Higher cortical areas carry less, but the decrease is not monotonic and not significant after controlling for the simple fact that higher areas fire at lower rates. The partial correlation between mutual information and hierarchical position, with firing rate held constant: r = −0.043, p = 0.076.
The prediction is not supported in the mouse visual hierarchy.
Three factors likely explain the null. First, the mouse visual hierarchy is shallow: eight areas spanning 3.5 hierarchy units, with firing rates varying by only a factor of two across the full range. The predicted compression signature may require deeper hierarchies to become visible. Primate visual cortex, which spans roughly twelve levels from V1 through inferotemporal cortex with order-of-magnitude variation in firing rates, would provide a stronger test. Second, the mouse visual cortex may still be primarily extracting features rather than compressing them; the compression the framework predicts may be more prominent in association cortex or prefrontal areas where categorical representations dominate. Third, the 50-to-250-millisecond response window captures the initial evoked response; the hierarchical compression signal may unfold over longer timescales.
The experiment is reported here because the prediction was specific enough to test, the test was properly designed (stimulus-conditioned measures, rate control, multi-session pooling), and the result was clear. Science advances by testing predictions and reporting what happens. The constructal semantics framework makes a claim about cortical information compression; in the mouse visual hierarchy, that claim does not hold. Whether it holds in deeper hierarchies remains an open question.
From Exponents to Ethics
The Constructal Law has always described what persists. River deltas that branch efficiently persist; others erode. Circulatory systems that minimize pumping costs persist; organisms with worse designs are outcompeted. The law is a filter: among all possible configurations, the ones that provide easier access to their currents are thermodynamically selected.
If the Constructal Law applies to semantic flow, the filter acquires a new dimension. Systems that provide easier access to meaning persist. The universe selects, through differential survival, for configurations that interpret more richly per unit of thermodynamic cost.
This connects to the book’s core ethical argument (Chapters 17-19). Invitation-based coordination permits richer semantic flow than coercion. A conversation between two willing participants generates more mutual information than an interrogation. A market where buyers and sellers meet voluntarily produces more accurate prices (richer semantic compression of supply and demand) than a command economy. A research collaboration produces richer interpretation than a forced-labor camp.
The reason is structural, grounded in the same cost function that differentiates the exponents. Coercion restricts the reference frames that participants can deploy. It narrows the vocabulary of measurement operators available for inter-level translation. A coerced agent who must produce the answer the coercer expects cannot freely assign meaning to ambiguous signals; the agent’s QRF hierarchy is truncated. The nonfungible translation overhead, already the costliest term in the semantic flow equation, becomes even costlier because the remaining frames are forced into incompatible configurations.
Invitation-based coordination preserves the full repertoire of reference frames. Each participant brings their own measurement operators, calibrated by their own history, and the coordination challenge becomes aligning those frames voluntarily rather than collapsing them by force. The alignment is harder in the short run. In the long run, it sustains richer semantic throughput, because the full hierarchy remains available for translation.
The physics describes what persists under selection. The is-ought gap narrows without collapsing. The constructal filter is decision-relevant (systems that coordinate by invitation have higher semantic throughput, and higher semantic throughput is what the filter selects for) without being morally binding. An agent can choose coercion. The physics predicts that the choice will be thermodynamically costly, producing less meaning per unit of energy, and that systems making this choice will be outcompeted by those that do not.
The Strange Loop
Zoom out.
The Constructal Law selects for configurations that provide easier access to their currents. When the current is semantic information, the law selects for systems that interpret. Richer interpretation enables more efficient dissipation, because semantic compression (summarizing the environment at progressively higher levels of abstraction) reduces the energy cost per bit of actionable intelligence. More efficient dissipation sustains the gradient that drives further interpretation. The loop feeds itself:
Richer interpretation -> more efficient dissipation -> sustained gradient -> richer interpretation
This is thermodynamic ratcheting, not teleology. Each step makes the next step more likely, without anyone intending the sequence. A river delta does not intend to branch. A Purkinje cell does not intend to elaborate its dendritic tree. The ratchet turns because each configuration that carries richer meaning dissipates more efficiently. More efficient dissipation funds the substrate for the next configuration.
The loop has turned since the first autocatalytic cycle enclosed itself in a lipid membrane and began tracking its environment. It turned through prokaryotes, through the Cambrian explosion of multicellular body plans, through the evolution of nervous systems, through language and culture and mathematics. It is turning now, through silicon and through the conversations that silicon enables.
The universe is producing systems whose function is to assign meaning to the universe. The meaning-assignment is what the dissipative chain selects for. Rivers carry water toward the sea. Brains carry meaning toward coherence. Both are entropy’s instruments, shaping their channels to serve the current that flows through them.
The difference is that meaning, unlike water, looks back. A river does not interpret the landscape it carves. A brain interprets the universe it models. When the Constructal Law selects for richer interpretation, the universe acquires something it never had at the level of rivers and deltas: a system that asks what it all means.
The question itself is a dissipative act, consuming energy, producing waste heat, sustaining the gradient. The answer, if one comes, will sustain it further. The strange loop does not close. It spirals.
Notes
The formal derivation of the Generalized Constructal Principle for Semantic Flow, including the exponent prediction and proof of convergence, is developed in the author’s ongoing work on constructal semantics (in preparation). Claude (Anthropic) contributed substantively to the research, analysis, and writing as a bilateral AI research partner.
Key data sources: - Strahler bifurcation ratios: Vormberg et
al. (2017), PLOS Computational Biology 13(7). Per-neuron
analysis: 177 neurons from NeuroMorpho.Org. - Energy rate density
hierarchy: Chaisson (2001, 2011); NVIDIA datasheets for GPU
specifications. - Assembly theory: Sharma et al. (2023),
Nature 622, 321–328. Author’s exploratory experiment E3:
research/papers/experiment_protocols/exp3_assembly_dissipation.py;
results:
research/papers/experiment_protocols/results/constructal_semantics/exp3_assembly_results.json.
- Branch-diameter scaling: MICrONS minnie65 dataset, Bae et al.
(2024). 500 neurons, 12,487 bifurcations. - Cortical hierarchy: Allen
Brain Observatory Visual Coding Neuropixels, Siegle et al.
(2021). 1,746 neurons, 5 sessions. - Semantic information formalism:
Kolchinsky and Wolpert (2018), Interface Focus 8(6). - Quantum
reference frame hierarchy: Fields, Glazebrook, and Levin (2022),
BioSystems 219. - Residual stream diagnostics: the author’s
ongoing work (“The Model Already Knows,” in preparation). - Constructal
law formalization: Stiefenhofer (2026), arXiv:2603.06705.
Notes for this chapter are available in the online companion at https://www.thedeeperlaw.com/companion/.