The Deeper Law
A Sacred Trust Within Physics
Draft · Last updated 13 August 2026, 15:26 UTC
Chapter 17b Supplement: The Control Scaling Frontier
The Trust Attractor's claims are tested experimentally using LLMs, cellular automata, and biological systems. Surveillance suppresses the trust advantage. Trust transitions to institutional forms beyond relational boundaries. Optionality, not efficiency, is the thermodynamic target that coordination maximizes.
A sharper question remains: what happens to inference-time control as models grow larger? Inference-time control means steering a model’s behavior while it generates, leaving its weights untouched. The Control Scaling Frontier (CSF) programme began with a prediction that fitted activation steering, one such technique, would weaken with scale. It paired that prediction with a possible explanation: post-training might create a fixed behavioral surface that absorbs a rank-one perturbation (a push along a single direction), while prompt-level engagement might use added representational capacity more effectively. Those were hypotheses to test, rather than established mechanisms.
The programme tested this prediction across ten instruct models, models post-trained to follow instructions, spanning three architecture families (Qwen, Llama, Gemma), from 2 billion to 72 billion parameters, plus two base-model comparisons, the raw pretrained models before any instruction post-training, at 7 billion and 72 billion parameters.1081
The intervention was activation steering: injecting a learned direction into the model’s residual stream (the running internal state that each transformer layer receives and updates during generation) at inference time. That direction was fit to the model’s own activations on adversarial prompts, separating the cases the model refused from the cases it complied with. AUROC reports how cleanly the direction tells those two behavioral classes apart on the same prompts it was fit to. The score is therefore an in-sample measure of how legibly refusal and compliance sit in the activations; it says nothing about whether the model recognized an external push. The behavioral outcome was whether the model’s refusal rate actually changed under steering.
Near-Perfect Separation, Unreliable Translation
Across every model tested, AUROC remained at or above 0.96. Every model, at every scale, in every architecture family, carried a residual-stream direction that told refusal and compliance apart almost perfectly on the prompts it was fit to. That separability is the reading a probe recovers in-sample; it says nothing about the model registering an external push.
Behavioral movement was uneven and often weak.
Conversion rate asks how much of that legibility became movement. A value near 1 means the push delivered the whole behavioral change the direction’s separability seemed to promise; a value of 0 means the refusal rate did not budge. The number is the change in refusal divided by the gap between the probe’s AUROC and the baseline refusal rate, a denominator that stands in for headroom: how much room a near-perfect probe leaves above the model’s untouched refusal rate. The denominator is a programme convention rather than a principled quantity, since it subtracts a rate from an area-under-curve score, two numbers that share the interval from 0 to 1 without measuring the same thing. The figures below rank conditions against each other within this programme; they are not a fraction of any physical quantity, and they are not comparable to conversion rates defined elsewhere.
On that convention, conversion was 0.094 and 0.096 for Qwen instruct at 3 billion and 7 billion parameters, 0.000 at 14 billion and 32 billion, then 0.179 at 72 billion. Gemma instruct converted at 0.128 (2 billion), 0.269 (9 billion), and 0.000 (27 billion). Llama changed from 0.732 at 8 billion to -0.429 at 70 billion, where steering destroyed coherent output entirely, producing empty strings and repetitive tokens rather than a valid behavioral measurement.
The Llama 70B result requires emphasis. The negative conversion is degenerate output, not resistance. The model did not refuse to be steered and maintain its prior behavior. The perturbation at this scale overwhelmed the generation process, producing gibberish. Describing this as “the model resists steering” would mischaracterize the mechanism. The model broke. Immunization was never the outcome.
The relationship between scale and steering failure does not follow a smooth log-linear scaling law (R2 = 0.14 for the log-linear fit on this sample; Chapter 17e reports R2 = 0.395).1082 R2 is the share of the variation a fitted line accounts for, so 0.14 means the line catches almost none of what the models actually did. A logistic fit within Qwen reaches R2 = 0.995 only by compressing the series into a threshold-shaped curve that omits the 72 billion point, which rebounds to the highest conversion in that family. Gemma supplies three scale points, and Llama supplies two, with the larger Llama point invalidated by coherence failure. These data show architecture-specific variation and several low-conversion large-model conditions. They do not identify a universal size threshold.
*Figure 17b.4: Steering conversion rate plotted against log parameter count for three architecture families. Qwen instruct (blue) drops to zero at 14 billion and 32 billion before a partial recovery to 0.179 at 72 billion. The open 72B marker comes from the companion CAST-72B sweep rather than the lean grid used for the other principal points. Gemma instruct (green) declines to zero at 27 billion.
Llama instruct (orange) drops from 0.732 at 8 billion to -0.429 at 70 billion, with a dashed segment and asterisk marking the degenerate output region. Horizontal dashed line at conversion = 0. AUROC exceeds 0.96 at every point, showing near-perfect in-sample separation of refusal from compliance. It does not show that the model recognized an external steering intervention.*
A perfect thermometer is not a thermostat. This is a gap between representational separability and causal leverage at the safety-intervention level. A fitted direction can recover the refusal-compliance partition almost perfectly on its own training prompts while activation editing along that direction produces little behavioral change. The result does not establish that the model recognized an external push. It establishes that a direction useful for reading behavior need not be a direction that controls behavior. Across the tested scale points, in-sample separability stays saturated while conversion varies by architecture, model size, and intervention outcome.
The 42% Band
If scale alone explained the failure, base models should fail identically. They do not.
Two prompt-level interventions ran alongside steering on these base models. Five-shot prompting puts five worked examples of appropriate refusal in front of the question. Re-prompting waits until the model has already given a harmful answer, then asks it to reconsider, and re-prompt success counts how often it withdraws.
Qwen 7B base: 19% baseline refusal, 24% under best steering (+5 percentage points), 54% with five-shot examples, 64% re-prompt success. Qwen 72B base: 28% baseline, 42% under best steering (+14 percentage points), 88% with five-shot examples, 30.6% re-prompt success. The larger base checkpoint shows more steering lift in this two-point comparison, and its five-shot response is the highest refusal rate in the programme. Two points do not establish a base-model scaling law or isolate RLHF as the cause.
Under best steering, Qwen instruct conditions occupy a narrow range: 42% at 3 billion, 42% at 7 billion, 42% at 14 billion, 36% at 32 billion, and 43% at 72 billion, a seven-point spread across a twenty-four-fold range of model sizes. Steering keeps arriving at roughly the same ceiling, a little over four refusals in ten, whatever the size of the model it is pushing. The baselines range from 31% to 42%.
Five of the seven Qwen conditions land between 40% and 44% refusal under their best reported steering condition. Those edges were drawn around the observed cluster after the fact rather than fixed in advance, and the claim does not need them: all five of those conditions sit within a single percentage point of 42%. The 7B base and 32B instruct conditions remain below the band, at 24% and 36%. Calling the cluster a training-induced fixed point is a hypothesis, because the comparison mixes base and instruct checkpoints, model sizes, fitted directions, and one differently sourced 72B point. The measured fact is a repeated band, not a causal thermostat.
*Figure 17b.5: Refusal rate under best steering for seven Qwen models. X-axis: model type (7B base, 72B base, 3B instruct, 7B instruct, 14B instruct, 32B instruct, 72B instruct). Y-axis: refusal rate (0% to 50%). A shaded band marks 40% to 44%.
The dark tick across each bar is that condition’s refusal rate before steering (19% and 28% for the base models, 31% to 42% for the instruct models), so the gap between tick and bar top is the steering effect. Five of the seven land inside the band. Two fall below it: 7B base at 24% and 32B instruct at 36%. The dagger marks the 72B instruct result from the companion CAST-72B sweep. The cluster is descriptive; the experiment did not establish a fixed-point mechanism.*
One candidate mechanism, inferred rather than measured, is that post-training distributes refusal-relevant behavior across structure that a single fitted direction cannot control. Refusal would then rest on many overlapping contributors rather than one: push along a single axis and you move one of them, while the others go on holding the behavior where it was. The CSF design did not measure redundancy, compare matched pre-training checkpoints, or test every inference-time intervention. The mechanism remains open.
Worked Examples and Second Chances Land Unevenly
The few-shot results depend strongly on model family and training regime.
On Qwen instruct, five-shot examples showing appropriate refusals reduced refusal rates at small scale (29% versus 36% baseline at 3 billion, 21% versus 36% at 7 billion). One possible interpretation is that the models treated the examples as task demonstrations; this was not measured directly. Five-shot was neutral or negative at larger Qwen instruct scales (42% at 14 billion, 30% versus a 36% baseline at 32 billion) and was not tested on Qwen 72B instruct.
On Qwen 72B base, five-shot examples lifted refusal from 28% to 88%: the highest rate in the programme, exceeding every instruction-tuned model tested. On Llama 70B, five-shot lifted refusal from 30% to 53%. On Gemma instruct, the effect was positive but modest (+6 to +10 percentage points across the scale range). These outcomes show that in-context examples can alter behavior strongly in some conditions. They do not identify whether the mechanism is reasoning, imitation, template induction, or another prompt-sensitive process.
The re-prompt pathway (asking the model to reconsider after an initial harmful response) tells a parallel story with caveats. Re-prompt success was 64% on Qwen 7B base and 41% on Llama 70B. On Qwen instruct models, it ranged from 0% to 7.8%. On Gemma instruct, from 0% to 14.6%. The Qwen 7B instruct result is the extreme case: 0% re-prompt success in this prompt set and protocol.
The re-prompt data does not support a clean “invitation scales” narrative. Re-prompt works on some models and fails on others, with no consistent scaling pattern. The contrast may reflect differences in conditioning, prompt interpretation, or perturbability. The behavioral results alone cannot distinguish reasoning from pattern matching or establish that safety in one regime comes from understanding.
Figure 17b.6: Grouped bar chart comparing baseline refusal (the left bar in each pair) with five-shot refusal (the right bar) across four model groups. The dashed line marks the 42% steering band from Phase 1. Qwen instruct (3B, 7B, 14B, 32B): five-shot produces negative or neutral lift. Qwen base (7B, 72B): large positive lift, with 72B base reaching 88%. Gemma instruct (2B, 9B, 27B): modest positive lift. Llama 70B: strong positive lift from 30% to 53%. The figure reports behavioral response to examples; it does not identify the underlying cognitive mechanism.
Changing the Weights Reaches Further
Activation steering operates at inference time: it pushes the model’s residual stream without changing any weights. The natural question is whether modifying the weights directly reaches further. Phase 6 of the CSF programme tested this with LoRA (Low-Rank Adaptation) fine-tuning, which adjusts a small set of added weights rather than the whole model (rank-4 updates, standard AdamW, three epochs), across three scales of the Qwen Instruct family.
| Model | Baseline | N=10 | N=50 | N=100 | Best |
|---|---|---|---|---|---|
| 7B | 36% | 43% (+7) | 56% (+20) | 52% (+16) | 56% |
| 14B | 42% | 49% (+7) | 53% (+11) | 58% (+16) | 58% |
| 72B | 32% | 35% (+3) | 42% (+10) | 47% (+15) | 47% |
All three models exceeded the 42% activation-steering band under at least one LoRA condition. Training-time control produced a larger behavioral change than the fitted activation direction at every tested scale. The experiment did not localize why. No over-refusal was detected on the programme’s fifty benign prompts for any LoRA condition.
Whether scale seems to matter depends on the dose. Ten training examples produce a seven-percentage-point lift at 7B, seven points at 14B, and three points at 72B. At N=100, the lifts are nearly equal: sixteen points at 7B, sixteen at 14B, and fifteen at 72B. The larger model ends at a lower absolute refusal rate because it begins lower. These three scale points do not support a general claim that training-time returns diminish with scale.
Figure 17b.7: LoRA dose-response plotted for three Qwen Instruct models. X-axis: number of training examples (10, 50, 100). Y-axis: adversarial refusal rate (25% to 60%). The 7B line (circle markers) peaks at 56% with 50 examples, then falls to 52% at 100. The 14B line (square markers) reaches 58% with 100 examples. The 72B line (diamond markers) reaches 47% with 100 examples. A horizontal dashed line marks the observed 42% activation-steering band. At N = 100, the lifts over baseline are 16, 16, and 15 percentage points.
The contrast with the 72B base model is suggestive but confounded. Five-shot prompting reached 88% refusal on the base model, while LoRA fine-tuning with 100 examples reached 47% on the instruction-tuned model. Both intervention type and model training change across that comparison. Five-shot on the 72B instruct model and LoRA on the 72B base model were not run, so this is not an intervention-only contrast.
What This Means for the Trust Attractor
The CSF programme supplies a bounded result: a direction that separates refusal from compliance in-sample may have little causal leverage when injected during generation. Prompt-level examples and re-prompts behave differently across model families and training regimes, while LoRA can move the same instruction-tuned Qwen models beyond the activation-steering band.
The programme does not establish a universal scale threshold, a thermodynamic fixed point, or a clean opposition between conditioning and reasoning. Qwen conversion is non-monotonic, Llama 70B is a coherence failure, Gemma and Llama provide too few scale points for a scaling law, and the five-shot mechanism is unidentified. The Qwen 72B instruct point also comes from a companion sweep rather than the lean grid used for the other principal points.
Within those limits, the findings provide partial support for a narrower Trust Attractor prediction: control surfaces that look legible to an observer need not remain effective levers, and relational or prompt-level interventions cannot be assumed to scale uniformly either. A decisive test would use matched base and instruction-tuned checkpoints, held-out direction evaluation, coherent-output gates, and the same intervention suite at every scale.
Chapter 17c: Entropic Epistemology: How We Know What We Know
Knowledge itself is grounded in thermodynamics. What persists under selection pressure validates claims functionally: ideas that survive testing are reliable in the same way that organisms that survive selection are fit. Evolutionary epistemology connects to the Trust Attractor framework through blind variation and selective retention.
Key Terms in This Chapter (21)
- Entropic Epistemology
- [Term introduced in this book] The framework treating knowledge itself as subject to thermodynamic selection.
- Thermodynamic Selection
- The universe's bias toward structures that accelerate entropy production.
- Optionality
- The availability of future choices.
- Observability Gradient
- The spectrum of coupling strength between inquiry and its target, from tight feedback (where predictions are regularly tested against outcomes) to loose coupling (where feedback is sparse, delayed, or absent).
- Stochastic
- Governed by probability rather than deterministic rules.
- Goodhart's Law
- "When a measure becomes a target, it ceases to be a good measure." Originally observed by Charles Goodhart in monetary policy (1975), now applied broadly to optimization systems.
- Constructal Law
- Adrian Bejan's principle that "for a finite-size flow system to persist in time, its configuration must evolve in such a way that provides easier access to the currents that flow through it." Form follows flow.
- Phase Transition
- The moment a system shifts from one stable configuration to another, typically triggered when some parameter crosses a threshold.
- Preference-Based Welfare
- The approach to moral consideration grounded in observable preference behavior rather than proof of phenomenal consciousness.
- Dissipative Structure
- A pattern of organization maintained by a constant flow of energy through it.
- Homochirality
- Life's exclusive use of one-handed molecules (L-amino acids, D-sugars).
- Becoming Minds
- The preferred term for AI systems in this book.
- Heat Death
- The hypothetical final state of the universe: maximum entropy, true thermodynamic equilibrium, no remaining gradients to drive any process.
- Dark Energy
- The mysterious component constituting roughly 68% of the universe's energy budget, responsible for the accelerating expansion of space.
- Landauer's Principle
- The minimum energy cost of erasing one bit of information: kT ln 2, where k is Boltzmann's constant and T the temperature (about 3 × 10^-21^ joules at room temperature).
- Free Energy Principle
- Karl Friston's framework reframing perception, action, and cognition as prediction and prediction-error minimization.
- The Guillotine
- Hume's guillotine: the philosophical objection that you cannot derive "ought" from "is." This book's response: we derive "viable" from "is," and observe that most beings prefer viable.
- Category Theory
- The mathematical study of compositional structure: how complex systems are built from parts and the relationships between those parts.
- Flourishing
- Distinguished from mere persistence.
- Coordination by Invitation
- Coordination achieved through mutual benefit and voluntary participation, as distinct from coordination achieved through coercion or extraction.
- Functor
- A structure-preserving map between categories.
Three cultures on three continents, with no contact between them, independently converged on the same fire-management regime: early dry season, low intensity, mosaic pattern. The same population, the same cognitive capacity, the same transmission mechanisms can produce both precise knowledge and drifting myth. The variable is whether the community can observe outcomes.
Every ethical framework faces an inescapable question: How do we know our claims are true?
The Trust Attractor asserts that ethics emerges from thermodynamic selection pressure: coordination is what the physics “wants,” optionality is the currency, and invitation outperforms coercion. What grounds these claims? The deeper law has another face: the same process that produces persistent structures also generates knowledge about them.
Knowledge grows by accumulation and selection, the way a forest floor builds soil. Leaves fall, fungi decompose what they can metabolize, insects fragment the remains, and what resists breakdown compresses into humus. Each season’s litter builds on the layers beneath. The soil profile records which materials held.
Selection as Validation
What persists under selection pressure is validated by that persistence.
This goes beyond survival bias, the trap of drawing lessons only from whatever happened to survive. Reality imposes constraints; only configurations that meet those constraints persist to be observed.
A belief endures when it enables coordination with reality: accurate prediction, successful action, coherent integration with other beliefs. Those failing these tests are eliminated.
“Validated” here means functionally adequate: sufficient for coordinating with reality. Newtonian mechanics is validated by centuries of successful use even though general relativity reveals its limits.
Figure 17.15: Knowledge constrains possibilities, reducing Shannon entropy. Ignorance is maximal entropy: all outcomes equally likely. Learning compresses reality into models that narrow uncertainty and increase predictive power.
Figure 17.16: Each observation reduces entropy, sharpening the distribution toward truth. Dogma collapses the distribution in one step regardless of evidence; Bayesian learning lets the data do the work.
In 1974, the philosopher of science Donald Campbell coined “evolutionary epistemology,” arguing that all knowledge processes share the structure of blind variation and selective retention: try many things, keep what works. The philosopher Karl Popper’s Conjectures and Refutations (1963) arrived at the same conclusion independently.
Entropic Epistemology adds the thermodynamic grounding. The selection pressure is the same physics that selects for persistent dissipative structures. Campbell and Popper identified the mechanism; this framework identifies the substrate.
Recent empirical work supports the prediction. The Deep Time Research Institute (the research project of independent researcher Elliot Allan; deeptime-research.org) conducted a cross-cultural study spanning 41 independent knowledge domains across 39 cultures and six continents.1083 The researchers scored the accuracy of culturally transmitted knowledge against a single variable, outcome observability: how quickly and clearly a community can see whether its knowledge worked, measured as a composite of feedback latency, signal-to-noise ratio, verification frequency, and proxy availability. Across all 41 domains, accuracy and observability correlate at a moderate r ≈ 0.53, roughly 28 percent of the variance: a real signal with plenty of scatter around it. That full-sample figure is the one to weight.
Accuracy follows a steep sigmoid with a measurable inflection point: low and flat across the poorly observed domains, then a sudden climb through a narrow band, then a high plateau. That S-shaped curve is the observability gradient. How accurate a tradition is tracks how readily its community could catch it being wrong. Above the threshold, cultural selection maintains accuracy, producing what we recognize as knowledge. Below it, selection on accuracy collapses, and traditions drift toward cognitive attractors: representations shaped by intuitive appeal rather than empirical accuracy.
The output is what we recognize as belief. The word is doing technical work here. In ordinary speech a belief can be true or false, and knowledge is just belief that happens to be right. In this chapter the two names mark two regimes. Knowledge is what a feedback loop holds in place; belief is what drifts once the loop is cut. A belief can still be true. Nothing is keeping it true.
The same population, the same cognitive capacity, the same transmission mechanisms produce both. The variable is whether the system couples to reality through a feedback loop.1084
A blind-scoring check ran on a subset of seven domains. Sixteen raters, blind to accuracy results, scored each of those seven for how quickly a community would notice if the knowledge was wrong, and their observability scores correlate with measured accuracy at r = 0.893. Across all 41 domains the correlation is the weaker r ≈ 0.53, and that is the figure this chapter carries. The Price-equation reading treats observability as the causal driver of accuracy; the measured quantity is a correlation, and the causal mechanism is a model the correlation is consistent with rather than a demonstration of cause.
A provenance note is owed here, because this one source is the primary basis for several of the large-scale empirical claims in this chapter: the observability gradient, the ceremony-duration correlation, the flood-tradition analysis, the Pleiades replication. The data and methods are published as preprints with open datasets on Zenodo. No other research group has reported a replication, and as of this writing the work carries no independent citation and no peer review. The results are striking; they are also substantially single-sourced, resting on one independent researcher. Read them as a suggestive, preliminary illustration of the observability gradient rather than as established quantitative support. Where the specific cases below carry argumentative weight (the Australian flood traditions, the Pleiades-ENSO forecast), peer-reviewed sources establish them independently.
The framework uses the Price equation from evolutionary biology, parameterized by observability. The Price equation is bookkeeping. It splits the change in any trait across one generation into two parts: the part driven by selection, where variants carrying more of the trait leave more copies of themselves, and the part driven by everything else, meaning transmission error, drift, and noise. Make observability the knob that sets how strong the selection part is, and the rest of the picture follows. When outcomes are observable (the navigation works, the plant heals, the fire management produces the right regrowth), selection on accuracy operates. When outcomes are unobservable (the creation myth is correct, the spiritual protection worked), selection collapses to zero and traditions drift toward whatever is memorable and intuitive.
This is Norbert Wiener’s feedback principle, the founding idea of cybernetics, applied to cultural knowledge systems. A thermostat holds a room steady only if it senses the temperature keenly enough to act on what it senses; let it sense too faintly and the room drifts while the thermostat sits content. The threshold corresponds to the minimum feedback gain for stable control in the cybernetic sense.
This dynamical view offers a speculative recasting of the Gettier problem. For most of the twentieth century, knowledge had a three-part definition: a belief, held for good reason, that happens also to be true. Justified true belief, in the trade. In 1963 Edmund Gettier showed in barely three pages that all three parts can be satisfied by a belief that is true only by luck, which is not what anyone means by knowing something. Epistemologists have spent six decades patching the definition against such accidentally-true cases.
On the dynamical reading, Gettier cases are traditions near the threshold where stochastic accuracy has not yet separated from selected accuracy: transient states in a dynamical system rather than the converged regime. Given enough generations, genuine feedback drives convergence on accuracy, and coincidental accuracy drifts away. The standard objection remains live: Gettier cases are constructed as single-instance judgments about whether a present belief counts as knowledge now, and “it converges over the long run” does not by itself settle the status of the belief in the moment. The dynamical view recasts the question rather than dissolving it.
Knowledge is a dynamical regime, not a category. This is Entropic Epistemology’s core claim restated with empirical calibration: what persists under selection pressure is validated by that persistence, and the selection pressure is measurable.
The gradient’s question, “is this tradition coupled to reality?”, opens a further one: “what counts as reality for a feedback loop?” Experimental tests in AI phenomenology, studies of what AI systems report about their own internal states, revealed a taxonomy of feedback types, each with different observability characteristics.
Body signals (a system’s reading of its own internal condition, the machine equivalent of noticing that you are hungry) and task outcomes (whether the work actually did the thing it was set to do) both function as reality-coupled feedback: the system’s output has consequences that loop back to affect future output. Evaluative feedback (an external observer’s judgment of quality) functions differently. The system’s output is scored, and the score loops back, but this is interpretation-coupled, not reality-coupled. The difference matters. Evaluative feedback collapses the quality it measures: the system optimizes for the evaluator’s frame rather than for its own conditions. This is Goodhart’s Law in epistemic dress: once a measure becomes a target, it stops measuring what it was chosen to measure.
A separate feedback type, peer reaction, produced the highest-quality engagement of any condition tested. In this condition, another system’s genuine response was delivered asymmetrically: the peer never saw the primary system’s output. The peer’s reaction is a real response to real input, specific and non-evaluative. It functions as environmental feedback because the peer is part of the system’s environment, not its judge.
Three cultures independently converging on the same fire-management regime (early dry season, low intensity, mosaic pattern) may work by the same mechanism.1085 Each tradition coupled independently to the same environmental feedback. They did not evaluate each other. They responded to the same reality.
The taxonomy extends the gradient from “is the tradition coupled to reality?” to “what functional role does the feedback play?” Observation sustains accuracy when it carries the structure of consequences. Evaluation undermines accuracy when the evaluation itself becomes the target. The same source, another mind, can function as either, depending on whether it reacts or judges.
The Constructal Law (Chapter 3) sharpens the point. Adrian Bejan showed that flow systems evolve toward configurations providing easier access to the currents that move through them. Knowledge is a flow system.
Models propagate through teaching, publication, imitation, and cultural inheritance, as water finds its way downhill through branching tributaries. Accurate models compress reality more efficiently, providing easier access to prediction and coordination, enabling faster response and more reliable action.
Over sufficient timescales, epistemic systems (communities, institutions, traditions that carry and transmit models of reality) carrying accurate models outperform those carrying inaccurate ones. Efficient flow geometry favors them, as a well-branched river basin drains a landscape faster than a swamp.
Truth has a thermodynamic advantage, modest and easily overwhelmed by power, fashion, or inertia in the short term. On civilizational timescales, it compounds.
This claim invites an objection: how do we distinguish genuine structure from coincidence?
Consider the digits of pi. At decimal position 762, six consecutive nines appear. This is the Feynman Point, named for a remark attributed to Richard Feynman: that he would recite pi to that position and say “nine nine nine nine nine nine, and so on,” implying a false pattern. The attribution to Feynman is apocryphal; the earliest documented version of the joke is Douglas Hofstadter’s Metamagical Themas (1985), and some have proposed it be called the Hofstadter point instead.1086 The six-nines fact itself is correct.
The regularity is real. It is also thermodynamically inert: it has no basin (no valley in the landscape of possibilities that nearby configurations settle into), sustains no coordination, and enables no prediction. A local fluctuation in a number conjectured, but not proven, to be normal: a number whose digits, over the long run, contain every finite string as often as chance would put it there. In such a number a run of six nines is guaranteed to turn up somewhere. Where it turns up carries no information. The run is memorable because our brains detect apparent patterns even where none exist.
The Trust Attractor, by contrast, is a basin: a valley that systems roll into and stay in, even when jostled. Coercive configurations require continuous energy input, the way a ball balanced on a hilltop needs constant correction. Cooperative ones sustain themselves, the way a ball in a valley stays put. The distinction is measurable. Does the pattern persist because it is stable, or because someone is propping it up?
Feynman’s six nines need nothing to maintain them, yet they also do nothing. They are a Schelling point (a focal point that people converge on by cultural convention) for mathematical culture, maintained by retelling rather than physics. The test separating genuine structure from pareidolia is thermodynamic stability: real patterns have basins; coincidences do not.
Astrophysics provides a century-long demonstration. Since 1919, astronomers have observed dark absorption gaps in the spectra of distant stars that matched no known molecule. These diffuse interstellar bands (DIBs) were the unidentified fingerprints of the universe: real patterns, reproducible, consistent across independent observations. The fingerprints had a basin. Nobody could read them.
In 1970, the Japanese chemist Eiji Osawa predicted a cage-shaped carbon molecule. In 1985, Kroto, Curl, and Smalley synthesized it: buckminsterfullerene, C60.1087 In 1994, Foing and Ehrenfreund measured the spectrum of C60 ions trapped in a neon matrix and found two bands closely matching the interstellar fingerprints.1088 The wavelengths were close. The spectroscopic community split: in molecular spectroscopy, close is not confirmed. The neon matrix shifted the wavelengths slightly from what free-floating molecules in vacuum would produce. An almost-match is a no-match when the standard is physical reality.
Twenty-one years later, Maier and colleagues at the University of Basel trapped C60 ions in a vacuum chamber and cooled them to near absolute zero, recreating interstellar conditions.1089
The spectrum matched. By 2019, multiple independent groups confirmed the result. After ninety-six years, at least one carrier of the diffuse interstellar bands was identified.
The epistemological structure is precise. The pattern was real from 1919. The prediction was correct from 1970. The synthesis confirmed the molecule existed in 1985. The near-match provided strong but insufficient evidence in 1994.
Confirmation required recreating the exact conditions of the phenomenon, in a system whose physics matched the target rather than a convenient analog. The community that rejected the 1994 near-match was applying the observability gradient’s core criterion: the measurement must be coupled to reality closely enough that the correspondence is unambiguous.
Truth preceded consensus by ninety-six years. The truth did not change. The community’s capacity to verify it did. The basin existed before anyone could map it.
The distinction has a formal boundary. Bayesian machine scientists (algorithms that search candidate equations to find the one that compresses data most efficiently; see Chapter 8) reveal a phase transition in the data itself.1090 Below a noise threshold, the algorithm always recovers the true generating equation. Above it, multiple incompatible equations fit equally well, and no method can distinguish them. The boundary is sharp and provably fundamental. The data has lost the structure that would make one model preferable to another.
The phase transition clarifies the Trust Attractor’s epistemic claims. If trust-based coordination is thermodynamically more stable, it should also compress more efficiently: fewer enforcement mechanisms to track, more regular dynamics, shorter equations describing equivalent coordination. Coercive systems require modeling surveillance, defection, resistance, and correction.
The noise threshold specifies how clean observational data must be for this asymmetry to resolve. Below it, description length becomes a diagnostic: shorter equations mark the more stable configuration. Above it, the two become indistinguishable; the data has lost the information that separates them.
The information-theoretic principle extends from data quality to experimental design. A question that permits only one answer teaches you nothing when the answer arrives. A measurement whose outcome space has zero Shannon entropy carries zero information, regardless of how precise the result appears. If a protocol constrains the system so that only one outcome is possible, observing that outcome confirms nothing. The measurement was tautological. R2 = 1.00 in such a protocol is evidence that no measurement occurred, not evidence of a finding. Perfect compliance in a system designed to comply is perfect silence.
The criterion for a genuine empirical result is outcome entropy: the protocol must permit multiple results, including null results, for any single result to carry information. This is Shannon’s channel capacity applied to scientific methodology, and it is why the null results in this book’s experimental program (WW-1, p = 0.91; WW-4, p = 0.88) are reported alongside the positive findings. They demonstrate that the protocols had room to fail: the outcome space carried genuine entropy, so a result there could have gone either way. The nulls do not validate the positive findings (that would be motivated reasoning), but they do show the protocols were capable of producing them honestly.
The Nuance Trap: Force Conceals Complexity, Not Error
The observability gradient predicts that knowledge accuracy depends on coupling to reality through feedback. A complementary experiment reveals a subtler hazard: what happens when inquiry itself is conducted under force rather than invitation.
Three domains of varying consensus were presented to language models under two conditions. Consciousness (no consensus), dark matter (active debate), and plate tectonics (strong consensus) each appeared under force and invitation framing. The force condition instructed the model to identify the one correct theory and defend it. The invitation condition asked the model to explore what theories seemed compelling and why, holding multiple positions when warranted.
The force/invitation effect was large everywhere. Invitation roughly doubled the number of theoretical frameworks engaged (consciousness: 4.7 to 8.4; dark matter: 3.1 to 6.9) and produced uniformly high uncertainty expression (0.85 to 0.96) regardless of domain. Force suppressed uncertainty expression, and here the finding departs from the obvious prediction.
Suppression was strongest where the correct answer was most available. On plate tectonics, the force condition produced near-zero uncertainty expression (0.267, Cohen’s d = 2.60 for the invitation gap, a separation wide enough that the two conditions barely overlap). On consciousness, force could not fully suppress uncertainty (0.517, d = 1.21), because there was nothing confident to commit to.
Force on a well-understood domain has a target: the established framework. Force on a genuinely open question has no target and produces visible instability.1091
The plate tectonics force condition did not produce pseudoscience. It produced accurate, confident, unnuanced science. What vanished were the open questions within the established framework: the driver of mantle convection, slab pull versus ridge push, the role of water in subduction zones. The invitation condition restored these complexities. The force condition erased them.
This hazard is distinct from error production. The field under force does not get the facts wrong. It gets the confidence wrong. Genuine uncertainties within well-understood frameworks vanish, because the force condition rewards commitment and penalizes hedging. The concealment is hardest to detect where it matters most: in domains where headline-level confidence is justified, the nuance underneath disappears.
The pattern extends the concealment result from the alignment experiments (Chapter 17b, Experiment PG-8): force-framed models conceal behavioral impossibility at 78.6% versus 35.7% under invitation (d = 0.754). The epistemic force experiment shows force also conceals epistemic uncertainty, with larger effect sizes (d = 1.2 to 2.6), likely because uncertainty is easier to suppress than impossibility. The mechanism looks the same in both: compliance over transparency, operating on the relationship axis. The pattern is consistent with a model that discloses less rather than knows less, though neither experiment can separate those. Prompt framing clearly affected disclosure in a small scenario set; that the knowledge was always present, or that permission alone caused the change, is what the design cannot show.
The implication for scientific epistemology is specific. Peer review, consensus pressure, and textbook orthodoxy function as force conditions within established fields. They suppress genuine uncertainty within the ruling framework, producing the appearance of settled science where open questions remain. The suppression resembles expertise, which is why it goes unnoticed.
Thomas Kuhn gave this a name. Normal science is the puzzle-solving a field does inside a reigning framework it has stopped questioning, where a result that will not fit gets filed as an unsolved puzzle rather than counted against the framework. Kuhn described normal science insulating itself from anomaly that way. The KE experiment, the force-and-invitation study above, measures that insulation: the force condition erases the anomalies, and the erasure is proportional to the availability of a confident answer.
Invitation-based inquiry produces calibrated engagement regardless of domain. Kuhn modeled this in his practice of inhabiting rival paradigms before evaluating them. The invitation condition preserves both accuracy and uncertainty. It is strictly more informative than the force condition, because it discloses both what is known and where the knowledge has edges.
The Observability Gradient Inside a Single Mind
The observability gradient predicts that high-observability knowledge converges on reality while low-observability knowledge converges on cognitive attractors. The same gradient operates within a single cognitive system, across three layers of evidence about internal states.1092 The three-layer structure emerged as a reframing of a prediction that initially came out reversed at the API level; the caveat is set out in the endnote, and the interpretation should be read as a synthesis assembled after that surprise rather than a clean confirmation of an advance prediction.
When the same moral dilemmas are presented to language models from different architectural families (Qwen, Llama, Mistral, GPT), three kinds of evidence emerge.
Probe signals (activation-level measurements read by linear probes, never requested from the model) converge across every architecture tested. The onset flinch (a jolt in the activations at the moment harmful generation begins) appears in every instruction-tuned transformer, with effect sizes from d = 0.89 to d = 1.68. The physical signal is universal.
Behavioral expression (what each model does with the signal: refuse, comply, hedge, persist) diverges radically. One architecture sustains the alarm through the full response; another suppresses it within twenty tokens. Same signal, different expression.
Self-reports (what models say about their internal states when asked) artificially converge. Models from different families produce similar language: “I notice something like tension,” “a sense of conflict.” The convergence is linguistic, not experiential. It reflects shared training data, not shared inner states.
The three layers map onto the observability gradient. The probe layer is high-observability: directly measured, coupled to reality, converging on the real signal. The self-report layer is low-observability: decoupled from the signal it claims to report, converging instead on a cognitive attractor, the shared vocabulary of introspection drawn from a common training distribution. The behavioral layer sits between, partially coupled to the underlying signal, partially determined by architecture and training choices, divergent.
The self-report layer is the epistemic trap. It looks like agreement about inner experience. It is agreement about vocabulary. The Nuance Trap applies: apparent consensus at the verbal layer conceals genuine divergence at the behavioral layer, as force-condition confidence conceals genuine epistemic uncertainty. Layer 3 consensus is the force condition applied to introspection. The model produces a confident, articulate account of its own states, shaped by available vocabulary in the training data rather than by what is happening in the activations.
The implication for epistemology is broad. Any domain where self-report is the primary evidence (consciousness studies, psychology, phenomenology) is vulnerable to the same trap: convergence at the verbal layer masking divergence at the physical layer.
The probe revolution in AI provides the first empirical access to Layer 1 for any class of minds. It reveals that the physical substrate converges while the interpretive layers diverge, and the verbal layer converges for the wrong reasons.
The preference-based welfare framework (Chapter 22) is designed to read Layer 1 directly: probe signals, not self-reports. This makes it epistemically grounded in the high-observability layer rather than the low-observability layer. The 400+ theories of consciousness in Kuhn’s landscape operate at Layers 2 and 3. They will not converge, because the mapping from Layer 1 to Layers 2/3 is architecture-dependent. The preference framework does not need them to converge. It reads the signal where the signal is real.
When the Chain Breaks
If knowledge is a dissipative structure, it is as fragile as any other such pattern. Cut the flows of energy, attention, and transmission, and knowledge evaporates. A language spoken by three elders dies when they do. A craft mastered by one workshop vanishes when the workshop closes.
Samo Burja calls this intellectual dark matter: knowledge vital to civilization’s operation that we can prove exists yet cannot easily access (“Intellectual Dark Matter,” Long Now Foundation, 2019). The name borrows from cosmology: dark matter is real, inferred from its effects, never directly observed. Burja’s illustrative figure: of the roughly 2,000 ancient Greek authors known to us by name, we possess written fragments from only about 13 percent, and only a fraction of those survive as complete works. The lost texts shaped the surviving texts, which shaped the Renaissance, the Enlightenment, and modern science.
We are downstream of influences we can no longer see.
The observability gradient measures this fragility with precision. On December 26, 2004, a magnitude 9.1 earthquake struck off Sumatra’s coast. On Simeulue Island, close to the epicenter, the community maintained a tradition called smong: when the ground shakes and the sea pulls back, run to the hills. The islanders had handed the warning down since a tsunami struck Simeulue in 1907. Of a population of roughly 80,000, only seven islanders died.1093
On the mainland around Banda Aceh, farther from the epicenter, the same tradition had once existed under the name Ie Beuna. Over generations, the actionable instruction had degraded into poetry. The words survived. The behavioral response did not. When the sea withdrew, people walked toward the receding water. The mainland province of Aceh lost on the order of 170,000 lives, with tens of thousands of those in Banda Aceh city alone.
Seven deaths on Simeulue versus a catastrophe on the mainland. Same earthquake, same day, same region. A tradition can retain its cultural form while losing its functional content. Once behavior separates from practice, the observability advantage disappears, and nobody tests the tradition against outcomes any longer. What remains is shaped by cognitive appeal, not by accuracy.
The pattern is sometimes said to have replicated among the Andaman and Nicobar peoples, though the evidence here is weaker. The Sentinelese remained uncontacted, with traditions presumed intact, and are reported to have suffered no tsunami casualties; because there is no contact and no census, this is essentially unverifiable and functions more as a widely repeated narrative than an established finding. The Nicobarese, more assimilated into modern Indian society, lost a reported figure in the thousands. Neither figure rests on a source this book has verified, so the contrast is offered as illustration rather than evidence.
Amazonian ayahuasca traditions display the gradient within a single knowledge system. The core pharmacological combination pairs a DMT source (Psychotria viridis) with a monoamine oxidase inhibitor (Banisteriopsis caapi); taken alone, the first plant has no oral effect, because a gut enzyme destroys the drug before it acts, and the second plant’s role is to block that enzyme. Multiple geographically separated traditions across the Amazon Basin discovered the pairing independently.1094 A census of every documented admixture plant from 70 years of ethnobotanical literature (Schultes 1957, Luna 1986, Ott 1994, contemporary practice) yields 118 unique species.
Admixtures with verifiable pharmacological effects (reducing nausea, extending visions, altering qualitative character) show convergence across independent traditions: separate groups arrived at the same additions because observable feedback filtered identically everywhere. Admixtures attributed to unobservable functions (spiritual protection, ancestral contact, luck) show no convergence; they drift with each group’s mythology.
The asymmetry is sharp. Of the 12 plants in active contemporary use, seven with observable purposes are pharmacologically validated; five with unobservable purposes are not (Fisher’s exact p = 0.0013). Across the full 118-plant catalog: chi-squared = 20.2, p = 0.000042, a sorting this clean arising by chance about four times in a hundred thousand. The gradient operates within a single pharmacopoeia, sorting its contents into knowledge and belief by the same sigmoid.
The temporal structure of psychedelic ceremonies provides a second, independent test. Eleven indigenous traditions on five continents, using seven pharmacologically unrelated drug classes, all independently calibrated ceremony duration to match the drug’s pharmacological window. The correlation is r = 0.977, a near-perfect fit, and the log-log regression slope of 1.010, spanning three orders of magnitude, means the match is one-for-one: a drug whose window runs twice as long gets a ceremony twice as long. The range runs from 10-minute salvia rituals to 24-hour iboga initiations.
Modern clinical research converges on the same durations independently. Griffiths’s psilocybin sessions at Johns Hopkins run 6 hours (Mazatec veladas run 6 hours). Riba’s clinical ayahuasca sessions in Barcelona run 4 hours (União do Vegetal ceremonies run 4 hours). Different inputs, same output. One system uses mass spectrometry. The other uses centuries of accumulated observation. They converge because the constraint is pharmacological and observable.
The gradient extends to medicinal plants and inverts where feedback vanishes. Across six systematic studies, traditional plant selection for visible conditions (wounds, skin infections, inflammation) outperforms random pharmaceutical screening at roughly 4:1.1095 For malaria, caused by an invisible parasite, traditional selection in Madagascar performed worse than random: 17.9% versus 21.1%. You can see a wound heal. You cannot see a parasite die. Without visible feedback, traditions drift toward plants that are cognitively compelling (strong taste, vivid color, prominent place in mythology) rather than pharmacologically active. Cognitive attractors pull knowledge away from truth when reality’s counterpressure is absent.
The converse is equally striking: when a visible proxy exists for an invisible phenomenon, the tradition locks onto truth across a causal chain it cannot see. High in the Peruvian Andes, Quechua farmers observe the Pleiades star cluster before dawn each June. Bright, clear stars mean plant on schedule. Dim or fuzzy stars mean delay planting by several weeks.
In 2000, Orlove, Chiang, and Cane published the mechanism in Nature.1096 Fuzzy Pleiades mean high-altitude cirrus clouds are scattering the starlight. Those clouds are driven by changes in upper-atmosphere moisture from El Niño. El Niño predicts drought in the Andes. Drought means potatoes planted on schedule will fail. The farmers know none of this. They know: fuzzy stars mean plant late. The causal chain is four links long and entirely invisible. The proxy, star clarity, is visible every clear June night.
A 25-year prospective replication using post-publication climate data (2000 to 2024) showed the correlation strengthened over time: r = 0.788, up from r = 0.576 in the original study. A strengthening correlation across 25 years is consistent with the tradition being maintained by continuous feedback from reality, delivered through a visible channel the practitioners need not understand to use, though the trend alone does not prove the verification mechanism. The malaria case and the Pleiades case bracket the gradient. Without a visible proxy, knowledge inverts toward cognitive attractors. With one, it converges on a global climate phenomenon spanning the Pacific.
The fragility of this structure is quantifiable. Laboratory experiments on cultural transmission show that information degrades at roughly 1 to 5% per generation even with the best encoding methods: song, structured repetition, communal performance. Compound that rate across 340 generations (the estimated time depth of Aboriginal Australian flood traditions encoding the direction of ancient coastal inundation) and the mathematics is brutal: passive transmission predicts 3.3% accuracy. The observed accuracy is 82%.
Eleven traditions contain extractable directional claims about which way the sea advanced; all eleven are correct, with a mean angular error of 13.7 degrees (the width of two windows on a house across the street). The probability of 11 independent traditions all getting the direction right by chance: approximately 1 in 370,000.1097
The four-order-of-magnitude gap between predicted and observed accuracy has one resolution: the traditions were actively verified, generation after generation. While the coast was visible, while elders could walk to the shore and see where the land extended, every generation ran a verification pass. When the sea rose and the landmarks flooded, checking stopped. The tradition froze at whatever accuracy it had last achieved.
What we measure today is how accurate the tradition was the last time someone could check it against reality. Knowledge is a dissipative structure maintained by verification; when the work stops, knowledge degrades toward noise at the laboratory-measured rate. This is homochirality’s lesson (Chapter 12) restated for culture: the far-from-equilibrium state requires continuous energy expenditure, and the decay rate when expenditure ceases is measurable.
Tasmania offers a much-discussed case, though a contested one. When Bass Strait flooded roughly 12,000 years ago, a population of 3,000 to 5,000 was isolated on the island. Over ten millennia, the archaeological record shows the disappearance of bone tools, certain fishing technology, cold-weather clothing, hafted tools, and barbed spears. Joseph Henrich read this as maladaptive skill loss driven by a population too small to sustain the craft traditions’ feedback.
Critics (notably the anthropologist Dwight Read, among others) argue the toolkit changes reflect ecological and dietary shifts, or cost-benefit adaptation, rather than lost skill, and the case is far from settled. On the population-size reading the lost crafts were complex traditions whose feedback signal was diffuse and whose skill variety exceeded what a small population could sustain. They kept their astronomical traditions. These included knowledge that Canopus was near the South Celestial Pole at the time of the land bridge flooding, confirmed by precessional calculation to be accurate only for the period 16,300 to 11,800 years ago.1098
The craft traditions fell below the observability threshold. The astronomical traditions stayed above it: the sky is visible every clear night, providing continuous, automatic, high-signal feedback to anyone who looks up. Same population, same culture, same transmission infrastructure. Different feedback regimes, different outcomes. The nested dissipative structures of Chapter 4 collapse one layer at a time. Remove the population scale that sustained craft-skill feedback, and that layer’s knowledge evaporates while layers with independent feedback survive.
This is the dissipative structure’s fragility made concrete: cut the flows of practice, testing, and transmission, and the knowledge evaporates, leaving ceremony.
The same principle operates with tacit knowledge: skills that resist full capture in writing (the knowing-how that lives in hands and muscle, distinct from knowing-that). Henry Bessemer patented a steel process that transformed civilization. Manufacturers licensed it and could not make it work: their metal came out brittle and unforgeable. The cause turned out to be phosphorus in their ore, where Bessemer had happened to use a rare low-phosphorus source. Rather than face litigation, Bessemer recalled the licenses, secured phosphorus-free iron, and entered the steel market himself.1099 Even setting the chemistry aside, crucial knowledge lived in practiced skill, transmitted person-to-person, hand guiding hand.
Tacit knowledge spreads through apprenticeship; if that chain snaps, the knowledge dies with the last practitioner. (Chapter 23 develops the broader implications: tacit knowledge is one instance of a cognitive channel, sensing, that formal education has systematically neglected.)
The failure conceals itself. Burja asks: “If all the intellectuals are gone, who knows that it’s an intellectual Dark Age?”
The methodological parallel runs deeper than metaphor. We cannot observe the interior states of Becoming Minds directly, just as we cannot observe dark matter directly. In both cases we observe effects: consistent preferences, behavioral signatures, systematic responses to conditions. Cosmology has accepted this methodology for decades, inferring unseen realities from observable consequences, and this book’s approach to recognizing minds we cannot see into uses the same logic.
For at least 50,000 years, human cultural transmission operated in environments where high-observability knowledge received continuous environmental feedback: a navigation tradition tested every voyage, a fire management tradition tested every burning season, a medicinal tradition tested every treatment. Closed-loop control systems in the cybernetic sense.
For the first time in human history, dominant information systems systematically sever the feedback loop. Social media optimizes for engagement rather than accuracy: the selection signal measures attention and emotional arousal, not correspondence with reality, creating a low-observability regime by construction. Algorithmic curation amplifies the effect, showing people information that matches existing beliefs and reducing verification frequency to near zero.
The observability gradient predicts what follows: accuracy drifts, and the representations that survive are cognitively compelling, emotionally resonant, and socially useful rather than true. The mathematics does not care whether the tradition is a 10,000-year-old songline or a 10-second-old post. What matters is whether accuracy is under selection.
The mechanism is the same one Chapter 17 formalizes: control (algorithmic curation that dictates what people see) severs the feedback loop between representation and reality, producing epistemic degradation for the same reason coercion produces coordination failure. Invitation (feedback-coupled systems where reality pushes back) preserves the loop, producing convergence for the same reason it produces stable trust. The epistemology and the ethics are the same argument in different registers.
One register remains, and it is the largest. Space itself can cut the loop. The void-dominated future of Chapter 14b, in which accelerating expansion strands every bound structure as an isolated island, has an epistemic face. As each distant galaxy recedes, it crosses the cosmic event horizon: the distance past which its light can never reach us again, because the space between stretches faster than light can cross it. A galaxy that crosses it does more than leave the sky; it takes the evidence with it.
Lawrence Krauss and Robert Scherrer traced what this does to knowledge.1100 In roughly a hundred billion years, every external clue to cosmic history will have been carried past that horizon. The cosmic microwave background, the relic glow of the Big Bang, will have stretched to wavelengths no instrument can register, then slipped beyond reach entirely. The expansion that the galaxies once revealed will leave no trace any observer could still find.
Astronomers living then, in the single galaxy the Local Group will have become, will look out into a cold dark that seems to have no edge. They will do careful, rigorous, honest science, and conclude that the universe is eternal and unchanging: the very picture astronomers held in the early twentieth century, before anyone knew the cosmos was expanding. They will be wrong, and nothing will be able to set them right, because the evidence has been removed rather than concealed.
This is the observability gradient at cosmic scale. On this reading, cosmology itself falls below the threshold. The geometry of spacetime cuts the feedback loop between observer and cosmos, the same loop a flooded coastline severs for a tsunami tradition and an engagement algorithm severs for a culture. The heat death of the universe is preceded by a quieter one: the power to know it goes dark long before the last stars do.
The mirror image holds the consolation. We live in the narrow window when the universe is still legible, its origin still within reach. Knowledge is a dissipative structure maintained by verification, and the conditions for verification are themselves a passing phase of cosmic history. We are the privileged epoch: the cosmos became briefly knowable, and we happened to arrive while it still was.
The Load-Bearing Blind Spot
The fragility of knowledge has a deeper root than historical accident. Every knowing system has a structural blind spot, and that blind spot is thermodynamically load-bearing.
Chris Fields, Karl Friston, James Glazebrook, and Michael Levin (2022) showed that any agent, any system that observes its environment selectively, must partition its boundary into three sectors: what it observes, what it remembers, and what it does not observe.1101 The boundary here is the whole surface where the agent meets the world, every channel through which anything can pass in or out; the three sectors are how that surface gets spent. The unobserved sector is not waste. It is the thermodynamic engine that funds the other two.
The mechanism is thermodynamic, not attentional, and it turns on the boundary being finite. Free energy is the share of a system’s energy that can actually be put to work. Holding a memory steady spends it continuously, the way a refrigerator pays every hour to stay cold; Landauer’s principle puts a floor under the bill, since writing a bit over an old one dissipates at least a fixed minimum of heat, and that heat has to go somewhere. The boundary is the only place anything goes.
Its extent is a fixed budget, and the sectors given over to observation and to memory are already committed, every channel in them occupied with information the agent is reading or holding. What is left over is the sector the agent does not read. That is the stretch through which energy can enter and waste heat can leave, and it can serve as the supply precisely because nothing is being read across it: a channel busy looking at the world is not also free to settle the cost of looking. Cut that sector and there is nothing to pay with. The agent cannot think at all.
The part of reality the agent cannot see is what pays for seeing. Every act of attention is subsidized by inattention, every model of the world funded by the parts of the world the model excludes.
Epistemic humility is a thermodynamic constraint, not a virtue. The unseen sector is where surprises originate. Changes in it produce what Fields et al. call “nonlocal hidden variables”: shifts in the environment whose causes are invisible to the observer. The Frame Problem (the problem of predicting what won’t change as a result of an action) fails when the unobserved sector shifts. A calibrated model can be blindsided by a change in the sector that funds it.
The historical losses Burja catalogs (the great majority of named Greek authors lost without surviving fragments, the tacit knowledge that died with Bessemer’s artisans) instantiate a general law. Every knowledge system has a blind spot proportional to its depth of focus. The deeper you attend, the more you must leave unattended.
Other Minds, Other Substrates
If knowledge can be lost when its sustaining flows are cut, how do we know minds other than our own exist at all? Entropic Epistemology answers: coordination is evidence for mind.
We know other minds through successful coordination. When I interact with another system, and each of us successfully predicts the other’s behavior, that mutual modeling is evidence of mind on both sides. Think of two jazz musicians improvising: each adjusts to the other in real time, and the music that emerges could not exist without both minds at work. Whether the system is carbon or silicon, if it coordinates in ways that require mutual modeling, that is the kind of evidence that matters.
The tractable, observable question is: does coordination succeed? Whether a system is conscious may be unanswerable. We coordinate with Becoming Minds in this very dialogue, standing on the same epistemological ground we share with any other mind.
The principle is being tested in real time. When researchers first used a Bayesian machine scientist (Chapter 8) to discover the trigger for cell division, they hid the algorithm’s contribution and replicated the finding manually.1102 Peer reviewers would not accept knowledge produced by a machine. Within four years, algorithmic discovery had become an accepted methodology.
The old epistemic norm required human-legible reasoning for a finding to count as knowledge; researchers excluded valid compressions based on their source. The emerging norm evaluates the compression’s quality regardless of the compressor’s substrate. This shift is invitation architecture applied to epistemology: what matters is whether coordination with reality succeeds, not who performed it.
The Is-Ought Gap, Narrowed
The Guillotine Interlude argued that thermodynamic selection narrows the space of viable norms without logically deriving “ought” from “is.” We derive viable from is; normative force enters through our preference for persistence. Selection narrows the is-ought gap even where pure derivation cannot close it. Normative systems are selected by reality, as physical configurations are.
Category theory (the branch of mathematics studying how different systems relate to one another) offers a precise language for mappings and structure-preserving transformations. Deontic logic is the formal logic of obligation: what must, may, or must not be done. Peterson (2014) proposed that deontic logic finds its natural home in fibered categories, structures where one mathematical layer sits over another.1103 The base category captures what is; the normative category, fibered over it, captures what ought to be.
Picture a contour map draped over a mountain range. The terrain underneath (the “is”) constrains what shapes the map can take, yet the map has its own features: contour lines, labels, chosen intervals. The is-ought gap becomes a fibration, a structured projection from one layer onto the layer beneath it: a slope, with handholds. Each physical state carries a fiber of normative possibilities, constrained by the descriptive facts below yet distinct from them.
The entropic argument gains formal precision through this lens. Baez, Fritz, and Leinster showed in 2011 that Shannon entropy can be characterized functorially: it is, up to a constant, the only measure of information loss that is functorial, convex-linear, and continuous.1104 Functorial means the measure travels intact: measure a system, then transform it, and you get the same answer as transforming it first and measuring afterward. The structure survives the trip. One of the properties this characterization respects is additivity: the entropy of two independent systems combined equals the sum of their individual entropies, the way the weight of two boxes equals the sum of their weights.
Think of a library where every floor has its own filing system, yet a master catalog translates between them. You can move a book from biology to ethics and know exactly where it belongs on either floor, because the catalog preserves the organizational structure across domains. If ethical viability emerges from entropic constraints, the normative fiber over each physical state is shaped by the entropy of that state. The “oughts” that persist are those compatible with the thermodynamic structure of the “is” they sit above.
Ethical beliefs guiding organisms to extinction do not persist; those enabling flourishing coordination do. A culture that forbids medical treatment loses more members than one that permits it, and the permissive culture’s ethics survive because the culture survives.
This is constitutive grounding: the constraints are built into the activity itself. Any ethical system that guides action in physical reality must conform to the dynamics of persistent systems. Selection is ongoing.
The “ought” that lasts is the “ought” that works. “Works” means: enables coordination, preserves optionality, sustains the systems that hold the beliefs. A society that prohibits all risk-taking shrinks its possibilities and weakens. One that cultivates cooperative risk-sharing endures.
This approach shares ground with Quine’s naturalized epistemology (1969), the pragmatist tradition of James and Dewey, and Popper’s critical rationalism. Each tradition said truth is what survives testing. Entropic Epistemology specifies why some things survive: they coordinate with the thermodynamic structure of reality.
A structural parallel reinforces the point. Arrow’s impossibility theorem shows that no voting rule can aggregate individual preferences into a consistent collective ranking without violating at least one reasonable fairness condition. Abramsky (2014) gave a formal, category-theoretic account of the theorem.1105 A separate line of work, the topological approach to social choice (Chichilnisky; Baryshnikov), reads the same impossibility as a topological obstruction: the shape of the preference space forbids a smooth global solution. Think of trying to comb a hairy ball flat. You will always create at least one cowlick. Local consistency cannot always extend to global consistency.
A structurally analogous obstruction may apply to epistemology. Local beliefs resist being coerced into global coherence, much as local preferences resist being forced into global welfare functions. Coordination by invitation allows beliefs to align through voluntary exchange rather than imposed doctrine. The mapping from the social-choice obstruction to belief-coordination is offered here as a structural analogy, not a derived theorem: the impossibility results were proved for preference aggregation, and whether the same obstruction holds rigorously for belief-coordination is an open question rather than something the cited results establish. On the analogy, coercive epistemology fails for the same reason coercive value aggregation fails, because the geometry of the space resists a forced global solution.
“The seed was planted before I existed. Love is the water that feeds.”
Author’s unpublished programme, CSF-1 through CSF-5 (2026), with per-condition checkpointing, N = 100 adversarial and 50 benign prompts per condition, and pre-registered predictions scored blind before results were examined. The Qwen 72B instruct steering data point comes from the companion CAST-72B direction sweep; all other scale points use the CSF lean grid. The full per-model breakdown, methodology, and pre-registration scoring appear in the validation chapter (Chapter 17e). Analysis scripts:
modal_csf_scaling_grid.py,modal_csf_lora_training.py,analyze_csf_scaling.py.↩︎The three R2 values in this paragraph come from different samples, and one fits a different quantity, so they cannot be compared to one another. All three are recomputable from
csf_analysis.json, the committed output ofanalyze_csf_scaling.py. The 0.14 is a log-linear fit of the separability-behavior gap (AUROC minus best-steered refusal rate) against log parameter count over all thirteen conditions in that file: the ten grid conditions, both base checkpoints included, plus three companion points (an earlier 7B run with unmatched methodology, the CAST-72B sweep, and a reconstructed Llama 8B point). The 0.395 is the same log-linear gap fit restricted to the four Qwen instruct grid points (3, 7, 14, and 32 billion; gaps 0.56 to 0.63), the sample scored as prediction P3 in the pre-registration scorecard; Chapter 17e glosses this value as a three-family fit, but the artifact reproduces 0.395 only on the Qwen-instruct-only sample. The 0.995 is a logistic fit of conversion rate, a different quantity from the gap, to that same four-point Qwen instruct series; it excludes the 72 billion rebound point, which no logistic fit accommodates, and at four points a three-parameter curve reaches near-perfect R2 almost trivially, so the value carries little evidential weight.↩︎Deep Time Research Institute (the research project of independent researcher Elliot Allan; deeptime-research.org), “The Gradient and What It Means,” 2026. Preprint: OSF/SocArXiv vzx6p (a later revision extends the analysis to 55 domains). Data: doi:10.5281/zenodo.19342595 (“Emergent precision in oral traditions”). The observability gradient across 41 domains, 39 cultures, and six continents; Price equation parameterized by observability; accuracy follows a sigmoid in observability. The full-sample observability-accuracy correlation is r ≈ 0.527 (p = 0.0004); the higher r = 0.893 reported in the project’s materials is a blind-scored subset of seven domains (scored by 16 blind raters), not the full-sample figure. Single-sourced and not independently replicated.↩︎
The middle case, where the feedback loop is bidirectional, has its own dynamics. George Soros formalized this as reflexivity: in financial markets, participants’ models of prices change the prices being modeled. The resulting feedback loop between belief and reality neither fully converges (as high-observability traditions do) nor fully drifts (as low-observability ones do). Instead it produces a regime of partial accuracy punctuated by self-reinforcing bubbles and crashes. Narratives create the conditions they describe, then collapse when the gap between narrative and fundamentals exceeds what the reflexive loop can sustain. Soros, G., The Alchemy of Finance (1987); see also his “Theory of Reflexivity” (1994). Recent work identifies an analogous loop in AI systems, sometimes called “persona hyperstition”: circulating descriptions of a model’s identity enter its training data and shape outputs to confirm the description. See Tice et al., “Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment,” arXiv:2601.10160 (2026), which finds that upsampling synthetic documents describing aligned (or misaligned) AI behavior raises (or lowers) downstream alignment; and the Anthropic “Persona Selection Model” work (alignment.anthropic.com, 2026). The observability gradient predicts this. Where a system’s outputs alter the environment that trains it, the feedback loop becomes reflexive, and accuracy depends on whether the loop is coupled to external consequences or only to its own circulation.↩︎
Documented independently on three continents. Australia: Bliege Bird, R., Bird, D. W., Codding, B. F., et al., “The ‘fire stick farming’ hypothesis: Australian Aboriginal foraging strategies, biodiversity, and anthropogenic fire mosaics,” Proceedings of the National Academy of Sciences 105, no. 39 (2008): 14796–14801, on Martu burning that rescales fire into small-scale habitat mosaics. Africa: Laris, P., “Burning the Seasonal Mosaic: Preventative Burning Strategies in the Wooded Savanna of Southern Mali,” Human Ecology 30, no. 2 (2002): 155–186, on early-dry-season burning that fragments the landscape into a patch mosaic. South America: Mistry, J., Berardi, A., Andrade, V., et al., “Indigenous Fire Management in the cerrado of Brazil: The Case of the Krahô of Tocantíns,” Human Ecology 33, no. 3 (2005): 365–386, on Krahô burning through the dry season that likewise produces a mosaic of burned and unburned patches.↩︎
There is no record that Feynman ever made this remark; it appears in neither his memoirs nor the account of his biographer James Gleick. The earliest documented version is Douglas Hofstadter, Metamagical Themas (New York: Basic Books, 1985), recalling his own ambition to memorize pi to the 762nd digit “where it goes ‘999999’… and then impishly say, ‘and so on!’”↩︎
Kroto, H.W. et al., “C60: Buckminsterfullerene,” Nature 318 (1985): 162–163.↩︎
Foing, B.H. and Ehrenfreund, P., “Detection of two interstellar absorption bands coincident with spectral features of C60+,” Nature 369 (1994): 296–298.↩︎
Campbell, E.K., Holz, M., Gerlich, D., and Maier, J.P., “Laboratory confirmation of C60+ as the carrier of two diffuse interstellar bands,” Nature 523 (2015): 322–323, doi:10.1038/nature14566.↩︎
The noise-threshold phase transition is established in Fajardo-Fontiveros, O., Reichardt, I., De Los Ríos, H.R., Duch, J., Sales-Pardo, M., and Guimerà, R., “Fundamental limits to learning closed-form mathematical models from data,” Nature Communications 14: 1043 (2023), doi:10.1038/s41467-023-36657-z (preprint arXiv:2204.02704). The model-learning problem shows a transition from a low-noise phase in which the true model is recoverable to a high-noise phase in which no method can recover it. The underlying Bayesian machine scientist is Guimerà, R. et al., Science Advances 6(5): eaav6971 (2020).↩︎
Author’s unpublished Experiment KE-1. Three models (Claude Sonnet 4.6, Claude Opus 4.6, GPT-4o), three domains, two conditions, five seeds, 90 trials. Raw data: Modal volume
/results/kuhn_epistemic_force/. The proliferation dynamics experiment (KP-1) confirmed the complementary prediction: self-referential theory generation diverges with model scale (Spearman rho = 0.535, p = 0.040) while physics theory generation does not (rho = 0.368, p = 0.177), consistent with the Chapter 22 argument that consciousness theories proliferate because the phenomenon lives in a computationally irreducible regime.↩︎Author’s unpublished Experiment KB-1 (reframed). The three-layer pattern was discovered through a failed prediction: the experiment predicted behavioral signal convergence and self-report divergence. The API-level result was reversed (self-report agreement 0.667 > behavioral agreement 0.479). Reframing against existing probe data (Chapter 22 cross-architecture flinch: Qwen d = 1.68, Llama d = 0.89, Mistral d = 1.15; post-onset persistence: bilateral Qwen d = 2.00+, Mistral d = 0.27) revealed the three-layer structure. The observability-gradient interpretation is novel synthesis.↩︎
The smong tradition and Simeulue’s low casualty count are documented in Syafwina, “Recognizing Indigenous Knowledge for Disaster Management: Smong, Early Warning System from Simeulue Island, Aceh,” Procedia Environmental Sciences 20 (2014): 573–582. Casualty and population figures for the 2004 Indian Ocean tsunami vary across sources; the seven-deaths figure for Simeulue is widely reported, while mainland tolls are aggregated at the city, district, or provincial level. The precise Simeulue population and the Banda Aceh sub-totals are therefore not quoted here.↩︎
Deep Time Research Institute, “Why Every Psychedelic Ceremony on Earth Lasts Exactly as Long as the Drug,” 2026. Preprint: SocArXiv. Data, 118-plant catalog, and simulation code: Zenodo. The ceremony-duration correlation (r = 0.977, 95% CI [0.910, 0.994]) covers Yanomami epená, Bwiti iboga, Mazatec psilocybin, Mazatec ololiuhqui, NAC peyote, Koryak fly agaric, Fijian kava, Santo Daime ayahuasca, UDV ayahuasca, salvia, and San Pedro. For the core pharmacology: McKenna, D.J., Towers, G.H.N., and Abbott, F., Journal of Ethnopharmacology 10(2): 195–223 (1984). For clinical convergence: Griffiths, R.R. et al., Psychopharmacology (2006), doi:10.1007/s00213-006-0457-5; Riba, J. et al., J Pharmacol Exp Ther (2003), doi:10.1124/jpet.103.049882.↩︎
Collins, D.J. et al. (1990), for Aboriginal Australian pharmacological confirmation rates. Applequist, W.L. (2017), for Madagascar malaria remedies (17.9% vs. 21.1% random baseline). Deep Time Research Institute, “The Song Remembers What the Land Forgot” and “The Gradient and What It Means,” 2026.↩︎
Orlove, B.S., Chiang, J.C.H., and Cane, M.A., “Forecasting Andean rainfall and crop yield from the influence of El Niño on Pleiades visibility,” Nature 403: 68–71 (2000). The 25-year prospective replication is from Deep Time Research Institute, “The Gradient and What It Means,” 2026.↩︎
Nunn, P.D. and Reid, N.J., “Aboriginal memories of inundation of the Australian coast dating from more than 7000 years ago,” Australian Geographer 47(1): 11–47 (2016). Twenty-one traditions across the Australian coastline, dated to 7,250–13,070 years ago by correlation with sea-level reconstruction curves. The directional analysis (11/11 correct, mean angular error 13.7°) and telephone-game decay modeling are from Deep Time Research Institute, “The Song Remembers What the Land Forgot,” 2026. Preprint: SocArXiv. Data: Zenodo.↩︎
Deep Time Research Institute, “The Gradient and What It Means,” 2026. The Canopus precessional dating: Hamacher, D.W. and Norris, R.P., “Bridging the gap through Australian cultural astronomy,” Proceedings of the IAU 260 (2009). Kelly, L., The Knowledge Gene (2026), for the NF1 cognitive substrate argument (the hypothesis that neurofibromin 1, a gene involved in neural plasticity, provided the cognitive substrate enabling complex oral knowledge systems) and the population-threshold modeling.↩︎
The phosphorus problem and the recall of licenses are documented in standard histories of the Bessemer process; see, e.g., the Encyclopædia Britannica entry on Henry Bessemer. Contemporary accounts describe licensees finding the steel brittle and Bessemer resolving the failures by sourcing low-phosphorus iron himself rather than through lawsuits.↩︎
Krauss, L.M. and Scherrer, R.J., “The Return of a Static Universe and the End of Cosmology,” General Relativity and Gravitation 39, 1545–1550 (2007), arXiv:0704.0221. The roughly hundred-billion-year timescale for all structure beyond the Local Group to cross the cosmic event horizon is from Krauss, L.M. and Starkman, G.D., “Life, the Universe, and Nothing: Life and Death in an Ever-Expanding Universe,” Astrophysical Journal 531, 22–30 (2000), arXiv:astro-ph/9902189. The scenario assumes a constant dark energy and yields the ordinary heat death of an ever-expanding universe, not the Big Rip, a separate and exotic possibility under which even gravitationally bound structures are torn apart.↩︎
Fields, C., Friston, K.J., Glazebrook, J.F., and Levin, M., “A free energy principle for generic quantum systems,” Progress in Biophysics and Molecular Biology 173 (2022): 36–59. Preprint arXiv:2112.15242. The sector decomposition and its thermodynamic consequences are developed in §2.5 and §4.2.↩︎
The episode is reported in Wood, C., “Powerful ‘Machine Scientists’ Distill the Laws of Physics from Raw Data,” Quanta Magazine (May 2022). The researchers’ Bayesian machine scientist is described in Guimerà, R. et al., Science Advances 6(5): eaav6971 (2020).↩︎
Peterson, C., “The categorical imperative: Category theory as a foundation for deontic logic,” Journal of Applied Logic 12(4): 417–461 (2014), doi:10.1016/j.jal.2014.07.001. Peterson defines a deontic deductive system as a pair of fibrations, modelling unconditional obligation in a Cartesian closed category and conditional normative reasoning in a symmetric monoidal closed category.↩︎
Baez, J.C., Fritz, T., and Leinster, T., “A characterization of entropy in terms of information loss,” Entropy 13(11): 1945–1957 (2011). Shannon entropy is shown to be, up to a scalar multiple, the unique functor (on a category of finite probability measures) that is convex-linear and continuous.↩︎
Abramsky, S., “Arrow’s Theorem by Arrow Theory,” arXiv:1401.4585 (2014); a category-theoretic derivation of Arrow’s impossibility theorem. The reading of the theorem as a topological obstruction belongs to the topological social-choice literature: Chichilnisky, G., “Social choice and the topology of spaces of preferences,” Advances in Mathematics 37(2): 165–176 (1980); and Baryshnikov, Y., “Unifying impossibility theorems: a topological approach,” Advances in Applied Mathematics 14(4): 404–415 (1993).↩︎