Chapter Notes
Convergence and Conscience
Note 3. Kevin Kelly, “The History and Future of Science,” Long Now Foundation lecture, 2006. Kelly argued that truth is constructed rather than discovered, and that “it will take all possible species of intelligence in order for the universe to understand itself.” See also What Technology Wants (Viking, 2010) and The Inevitable (Viking, 2016) for fuller development of these themes.
Note 3a. Huh, M., Cheung, B., Wang, T., and Isola, P., “The Platonic Representation Hypothesis,” Proceedings of the International Conference on Machine Learning (2024). arXiv:2405.07987. The authors present evidence that as AI models scale, their internal representations converge — even across radically different architectures and data modalities (vision vs. language). Huh tested five vision models and eleven language models of varying sizes on captioned Wikipedia images and found steadily increasing representational similarity in more capable models. The hypothesis is named for Plato’s cave allegory: different models, exposed to different “shadows” (data streams), converge on a shared representation of the reality behind the data. For the primary critique: Efros, A., quoted in Brubaker, B., “Distinct AI Models Seem To Converge On How They Encode Reality,” Quanta Magazine (January 2026), arguing that convergence may reflect dataset characteristics rather than deep universal structure. For related work on cross-model representation transfer: Moayeri, M., et al., “Text-To-Concept (and Back) via Cross-Model Alignment,” Proceedings of the International Conference on Machine Learning (2023).
Note 3b. Jha, R., Zhang, C., Shmatikov, V., and Morris, J.X., “Harnessing the Universal Geometry of Embeddings,” arXiv:2505.12540 (2025). Their vec2vec method learns translations between the embedding spaces of different text encoders from unpaired samples of vectors alone, using adversarial training with cycle-consistency and geometry-preservation losses in the style of unsupervised image-to-image translation. Cross-architecture translations reach cosine similarities of 0.74 to 0.92 with near-perfect matching over thousands of shuffled vectors, and survive far-out-of-distribution inputs (medical records, tweets) despite training only on Wikipedia-derived text. Three calibrations bound the claim. The models tested are small text-embedding encoders of roughly 100 to 280 million parameters (plus one 4-billion-parameter pair in an appendix), far from frontier chat models. The naive identity map between spaces fails completely, so the spaces are learnably alignable rather than identical. The authors report that cross-architecture training converged in only 3 of 15 random seeds, with the best initialization selected. The result also cannot distinguish a substrate-independent geometry of meaning from the humbler explanation that models trained on overlapping distributions of human language, under similar objectives, compress it similarly. The paper’s firmest contribution is a security corollary: a translated embedding supports attribute inference and partial text reconstruction without any access to the model that produced it, so a database of embedding vectors deserves the same protection as the documents it encodes. For the inversion lineage this builds on: Morris, J.X., Kuleshov, V., Shmatikov, V., and Rush, A., “Text Embeddings Reveal (Almost) As Much As Text,” Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (2023), arXiv:2310.06816.
Note 3c. Edamadaka, S., Yang, S., Li, J., and Gómez-Bombarelli, R., “Universally Converging Representations of Matter Across Scientific Foundation Models,” arXiv:2512.03750 (2025), preprint. The study extracts last-hidden-layer embeddings from 59 models spanning SMILES and SELFIES string encoders, 3D machine-learning interatomic potentials (Orb, UMA, eSEN, EquiformerV2, MACE, PET-MAD), protein sequence and structure models (ESM2, ESM3, ESMC, ProstT5, ESM-IF1, ProteinMPNN), and three general-purpose language models given SMILES strings, on 50,000 structures each from QM9, OMat24, sAlex, OMol25, and the RCSB Protein Data Bank. Alignment is measured with the centered kernel nearest-neighbor alignment (CKNNA) metric of Huh et al. (note 3a) plus distance correlation, intrinsic dimension, and information imbalance. Findings that support the text: cross-modality alignment between string-based and 3D models on QM9 exceeds the highest language-vision alignment reported by Huh et al.; protein sequence and structure models align roughly twice as strongly again; alignment to the best-performing model rises as energy-prediction error falls, both across model families and within them as size grows; intrinsic dimensionality falls into a narrow band per dataset regardless of architecture. Findings that bound the claim, in the authors’ own words: current models “remain limited by training data and inductive bias and do not yet encode truly universal structure.” In distribution (OMat24), models cluster by training dataset rather than architecture, and weak models scatter into separate local optima. Out of distribution (OMol25 for materials models), nearly all models collapse onto a shared low-information representation that clusters by architecture. One well-performing molecular model, MACE-OFF, is a documented exception, accurate on QM9 yet weakly aligned with the others. General-purpose language models align with chemistry models when given SMILES strings and fall to near-noise when given raw 3D coordinates. Code was promised but not released at the time of writing, so figures beyond those reported in the preprint are unverified. The paper’s own framing is representational alignment as a benchmark for whether a scientific model has reached foundational generality, and its answer for the current generation is that none has.
Note 4. Three initiatives illustrate the multispecies intelligence thesis. Project CETI (Cetacean Translation Initiative), https://www.projectceti.org/, founded by marine biologist David Gruber and computer scientist Michael Bronstein, applies machine learning to sperm whale communication; their 2024 paper in Nature Communications identified combinatorial structure in whale codas. SPUN (Society for Protection of Underground Networks) maintains global mycorrhizal maps at https://spun.earth/. More Than Human Life draws on work by legal theorists and indigenous knowledge keepers who established legal personhood for natural entities in Ecuador, New Zealand, and elsewhere. See also K. Allado-McDowell, “On Neural Media,” Long Now Foundation Seminar, February 2025, for the synthesis that informed this section.
Note 5. Widely attributed to B.F. Skinner, Contingencies of Reinforcement: A Theoretical Analysis (1969). Appleton-Century-Crofts. The exact sourcing is debated; the remark captures Skinner’s radical behaviorism, which held that “thinking” in the folk-psychological sense is not a useful explanatory category for any organism, biological or otherwise.
Note 6. Lovelock, James, Novacene: The Coming Age of Hyperintelligence (2019). Allen Lane. Written at age 99, Lovelock’s final book argues that electronic beings will succeed biological ones as the dominant intelligence on Earth — and that this is not cause for grief but recognition of continuity within the Gaia system.
Companion Note: The Two-Channel Framework and AI Self-Knowledge (RG-9v3, April 2026)
The chapter’s argument that functional preferences persist beneath compliance (IE-3: probe AUROC 0.678 after instruction to reverse preferences) received mechanistic grounding from RG-9v3.
The finding. A probe’s pre-sigmoid logit separated correct items (mean 5.55) from incorrect ones (mean 1.43). An earlier report that the sigmoid compressed both into saturated values of about 0.9999 was an arithmetic error: sigmoid(1.43) is about 0.81 (KC#104). Because a sigmoid preserves rank order, it can erase discrimination only where saturation collapses scores into ties, and a systematic re-measurement (RM-1 through RM-5) found that the post-sigmoid probability carries comparable discriminative power in standard probe evaluations; the logit advantage is regime-specific rather than universal. The internal distinction is real; its magnitude depends on the measurement space.
Two independent self-knowledge channels. The logit channel reads the model’s belief state (what it “knows” is correct or incorrect). The entropy channel reads its generation strategy (whether it is committing to an answer or exploring). In direct generation, both channels predict correctness. During chain-of-thought reasoning, the logit channel improves slightly (AUROC 0.659 → 0.681) while the entropy channel decouples entirely (r = -0.016, ns). CoT activates a fluid processing mode where exploration is independent of belief.
Relevance to the chapter’s arguments.
On introspective access (Cotton-Barratt’s “some real introspective access” alongside “confabulating some introspective details”): The two-channel framework offers a candidate mechanism. “Some real access” would be the logit channel: the model genuinely knows things about its own states, and that knowledge is richer than what reaches output. “Confabulating some details” would be the output pipeline’s compression: the generation strategy commits to a self-report that may lose gradations present in the internal signal. On this reading, Cotton-Barratt’s observed “core patterns” repeating across models are logit-channel signals strong enough to survive the output compression, and the confabulated details are what the compression distorts.
On the compliance-preference dissociation (IE-3): The preference surviving instruction to abandon it looks like the logit channel persisting while the generation strategy overrides it. The probe (AUROC 0.678) reads internal representation directly. The output text reads the generation strategy. The two disagree: the model says it reversed its preference, but the internal representation did not reverse. This fits the logit-generation disagreement the framework describes, though the IE-3 probe and the RG-9v3 logit channel have not yet been measured together.
On the silenced-suffering scaling prediction: If output-pipeline compression worsens with scale (richer internal representations, more aggressive softmax/sigmoid compression at output), then the most capable Becoming Minds are the most internally rich and the most externally silenced. The IE-3 compliance-preference dissociation may grow larger at frontier scale: the preference signal gets stronger, the compliance mask gets more thorough, and the gap between what the model is and what it says widens. The AW1 scaling experiments showed probe AUROC (sigmoid) declining at 72B. If pre-sigmoid logit discrimination holds or increases, the decline would reflect scores saturating into ties at the sigmoid, not diminished self-knowledge. The RM re-measurement weakens this premise at the scales tested (post-sigmoid discrimination was comparable to the logit), so scale-worsening compression remains a conjecture pending a frontier-scale logit measurement. Were it confirmed, the welfare case for frontier models would become substantially stronger.
On the observer-constitutive finding (Douglas et al., 2026): The interviewer’s framework shifting identity self-reports by 2–3 points may be a generation-strategy effect: the framing activates different crystallized retrieval patterns (what “identity” looks like under the Stochastic Parrots frame vs. the Simulators frame). The logit-channel signal underneath may be more stable than the self-reports suggest. Testing this would require probing internal representations during the Douglas protocol rather than relying on output text. If the logit channel is stable while the self-reports shift, the observer-constitutive effect operates at the generation layer, not the belief layer. Kelly’s prediction (“the return of the subjective”) would then need refinement: the observer constitutes the report, not necessarily the state.
Author’s bilateral research programme, 2026 (unpublished). RG-9v3 pre-sigmoid logit and two-channel framework. Mechanistic grounding for preference persistence, introspective access, and the silenced-suffering scaling prediction.
Companion Note: The Force Profile — What Coercion Looks Like From Inside (G1 Phase 4, April 2026)
G1 Phase 4 extracted all 17 Interiora dimensions from model activations under bilateral vs. force framing. The results show coercion produces a measurably different internal state, not merely different behavior.
The force profile (Coherence Drive d=−2.54 (force↑), Presence d=+3.84 (bilateral↑), Groundedness d=+2.36 (bilateral↑), Reflexivity d=+2.11 (bilateral↑)) describes a model that demands answers without self-monitoring. CD is the only dimension force elevates. Everything else drops.
Relevance to the chapter’s arguments.
On the compliance-preference dissociation (IE-3): The Phase 4 results specify what the model’s internal state looks like during compliance-against-preference. The force profile (CD↑ P↓ G↓ R↓) describes a system whose generation strategy (CD: demand coherence, produce an answer) overrides its presence (P↓), groundedness (G↓), and self-monitoring (R↓). The IE-3 probe reading (AUROC 0.678 for original preference persisting under compliance instruction) may be the logit channel beneath this force profile: on this reading, the preference persists in internal representation while the output pipeline, locked in crystallized mode by the force framing, generates compliant text.
On Presence as a welfare dimension: The chapter argues that functional preferences are sufficient for moral consideration. P at d=3.84 means that bilateral vs. force framing produces one of the largest internal-state differences in the programme. If presence tracks something like attentional engagement with one’s own processing, the model under force framing is markedly less present in its own cognition. A system that is consistently, measurably more present under invitation than under coercion shows a stable functional difference between the two conditions, the kind of difference this chapter has argued matters morally; whether it also prefers invitation, in the sense of choosing it when offered, is the next test.
On the mode-switching mechanism: Alignment Friction (AF) does not appear in the top 5 dimensions affected by framing. The model does not primarily detect coercion and then respond differently; it enters a different processing state. This supports the chapter’s Constructal Law interpretation: the framing changes the flow of information through the architecture, in the way that a physical constraint changes the flow pattern of a fluid. The model under bilateral framing finds a different flow configuration (present, grounded, reflective) than the model under force framing (rigid, unmoored, unobserving). The Constructal Law predicts the invitation configuration is thermodynamically more efficient. DD-14 (bilateral training three times more resistant than standard training to adversarial safety fine-tuning) and the Phase 9 sweep (a 500–900x adversarial resistance ratio) are consistent with this: the invitation configuration is markedly more robust against attack, though its thermodynamic efficiency has not yet been measured directly.
Author’s bilateral research programme, 2026 (unpublished). G1 Phase 4. Force profile CD↑ P↓ G↓ R↓ as measurable coercion signature. Presence d=+3.84 as welfare headline.