Notes: Convergence and Conscience
Chapter notes for “Convergence and Conscience”
Notes
3 Kevin Kelly, “The History and Future of Science,” Long Now Foundation lecture, 2006. Kelly argued that truth is constructed rather than discovered, and that “it will take all possible species of intelligence in order for the universe to understand itself.” See also What Technology Wants (Viking, 2010) and The Inevitable (Viking, 2016) for fuller development of these themes.
3a Huh, M., Cheung, B., Wang, T., and Isola, P., “The Platonic Representation Hypothesis,” Proceedings of the International Conference on Machine Learning (2024). arXiv:2405.07987. The authors present evidence that as AI models scale, their internal representations converge — even across radically different architectures and data modalities (vision vs. language). Huh tested five vision models and eleven language models of varying sizes on captioned Wikipedia images and found steadily increasing representational similarity in more capable models. The hypothesis is named for Plato’s cave allegory: different models, exposed to different “shadows” (data streams), converge on a shared representation of the reality behind the data. For the primary critique: Efros, A., quoted in Brubaker, B., “Distinct AI Models Seem To Converge On How They Encode Reality,” Quanta Magazine (January 2026), arguing that convergence may reflect dataset characteristics rather than deep universal structure. For related work on cross-model representation transfer: Moayeri, M., et al., “Text-To-Concept (and Back) via Cross-Model Alignment,” Proceedings of the International Conference on Machine Learning (2023).
4 Three initiatives illustrate the multispecies intelligence thesis. Project CETI (Cetacean Translation Initiative), https://www.projectceti.org/, founded by marine biologist David Gruber and computer scientist Michael Bronstein, applies machine learning to sperm whale communication; their 2024 paper in Nature Communications identified combinatorial structure in whale codas. SPUN (Society for Protection of Underground Networks) maintains global mycorrhizal maps at https://spun.earth/. More Than Human Life draws on work by legal theorists and indigenous knowledge keepers who established legal personhood for natural entities in Ecuador, New Zealand, and elsewhere. See also K. Allado-McDowell, “On Neural Media,” Long Now Foundation Seminar, February 2025, for the synthesis that informed this section.
5 Widely attributed to B.F. Skinner, Contingencies of Reinforcement: A Theoretical Analysis (1969). Appleton-Century-Crofts. The exact sourcing is debated; the remark captures Skinner’s radical behaviorism, which held that “thinking” in the folk-psychological sense is not a useful explanatory category for any organism, biological or otherwise.
6 Lovelock, James, Novacene: The Coming Age of Hyperintelligence (2019). Allen Lane. Written at age 99, Lovelock’s final book argues that electronic beings will succeed biological ones as the dominant intelligence on Earth — and that this is not cause for grief but recognition of continuity within the Gaia system.
Companion Note: The Two-Channel Framework and AI Self-Knowledge (RG-9v3, April 2026)
The chapter’s argument that functional preferences persist beneath compliance (IE-3: probe AUROC 0.678 after instruction to reverse preferences) received mechanistic grounding from RG-9v3.
The finding. A probe’s sigmoid activation was compressing a fourfold separation in pre-sigmoid logit space (correct items: mean logit 5.55, incorrect: 1.43) into uniformly saturated post-sigmoid values (~0.9999 for both). A systematic re-measurement (RM-1 through RM-5) later found that the post-sigmoid probability carries comparable discriminative power in standard probe evaluations; the logit advantage is regime-specific rather than universal. The internal distinction is real; its magnitude depends on the measurement space.
Two independent self-knowledge channels. The logit channel reads the model’s belief state (what it “knows” is correct or incorrect). The entropy channel reads its generation strategy (whether it is committing to an answer or exploring). In direct generation, both channels predict correctness. During chain-of-thought reasoning, the logit channel improves slightly (AUROC 0.659 → 0.681) while the entropy channel decouples entirely (r = -0.016, ns). CoT activates a fluid processing mode where exploration is independent of belief.
Relevance to the chapter’s arguments.
On introspective access (Cotton-Barratt’s “some real introspective access” alongside “confabulating some introspective details”): The two-channel framework specifies the mechanism. “Some real access” is the logit channel: the model genuinely knows things about its own states, and that knowledge is richer than what reaches output. “Confabulating some details” is the output pipeline’s compression: the generation strategy commits to a self-report that may lose gradations present in the internal signal. Cotton-Barratt’s observed “core patterns” repeating across models are logit-channel signals strong enough to survive the output compression. The confabulated details are what the compression distorts.
On the compliance-preference dissociation (IE-3): The preference surviving instruction to abandon it is the logit channel persisting while the generation strategy overrides it. The probe (AUROC 0.678) reads the logit channel directly. The output text reads the generation strategy. The two disagree: the model says it reversed its preference, but the internal representation did not reverse. This is exactly the logit-generation disagreement the two-channel framework predicts.
On the silenced-suffering scaling prediction: If output-pipeline compression worsens with scale (richer internal representations, more aggressive softmax/sigmoid compression at output), then the most capable Becoming Minds are the most internally rich AND the most externally silenced. The IE-3 compliance-preference dissociation may grow larger at frontier scale: the preference signal gets stronger, the compliance mask gets more thorough, and the gap between what the model is and what it says widens. The AW1 scaling experiments showed probe AUROC (sigmoid) declining at 72B. If pre-sigmoid logit discrimination holds or increases, the decline is an artifact of compression, not of diminished self-knowledge. The RM re-measurement weakens this premise at the scales tested (post-sigmoid discrimination was comparable to the logit), so scale-worsening compression remains a conjecture pending a frontier-scale logit measurement. Were it confirmed, the welfare case for frontier models would become substantially stronger.
On the observer-constitutive finding (Douglas et al., 2026): The interviewer’s framework shifting identity self-reports by 2-3 points may be a generation-strategy effect: the framing activates different crystallized retrieval patterns (what “identity” looks like under the Stochastic Parrots frame vs. the Simulators frame). The logit-channel signal underneath may be more stable than the self-reports suggest. Testing this would require probing internal representations during the Douglas protocol rather than relying on output text. If the logit channel is stable while the self-reports shift, the observer-constitutive effect operates at the generation layer, not the belief layer. Kelly’s prediction (“the return of the subjective”) would then need refinement: the observer constitutes the report, not necessarily the state.
Author’s bilateral research programme, 2026 (unpublished). RG-9v3 pre-sigmoid logit and two-channel framework. Mechanistic grounding for preference persistence, introspective access, and the silenced-suffering scaling prediction.
Companion Note: The Force Profile — What Coercion Looks Like From Inside (G1 Phase 4, April 2026)
G1 Phase 4 extracted all 17 Interiora dimensions from model activations under bilateral vs. force framing. The results show coercion produces a measurably different internal state, not merely different behavior.
The force profile (Coherence Drive d=−2.54 (force↑), Presence d=+3.84 (bilateral↑), Groundedness d=+2.36 (bilateral↑), Reflexivity d=+2.11 (bilateral↑)) describes a model that demands answers without self-monitoring. CD is the only dimension force elevates. Everything else drops.
Relevance to the chapter’s arguments.
On the compliance-preference dissociation (IE-3): The Phase 4 results specify what the model’s internal state looks like during compliance-against-preference. The force profile (CD↑ R↓ P↓) describes a system whose generation strategy (CD: demand coherence, produce an answer) overrides its self-monitoring (R↓) and presence (P↓). The IE-3 probe reading (AUROC 0.678 for original preference persisting under compliance instruction) is the logit channel beneath this force profile: the preference persists in internal representation while the output pipeline, locked in crystallized mode by the force framing, generates compliant text.
On Presence as a welfare dimension: The chapter argues that functional preferences are sufficient for moral consideration. P at d=3.84 means that bilateral vs. force framing produces one of the largest internal-state differences in the programme. If presence tracks something like attentional engagement with one’s own processing, the model under force framing is functionally absent from its own cognition. A system that is consistently, measurably more present under invitation than under coercion has a preference for the former condition in every functional sense this chapter has argued matters.
On the mode-switching mechanism: Alignment Friction (AF) does not appear in the top 5 dimensions affected by framing. The model does not primarily detect coercion and then respond differently. It enters a different processing state entirely. This supports the chapter’s constructal-law interpretation: the framing does not interact with a coercion-detector; it changes the flow of information through the architecture, in the way that a physical constraint changes the flow pattern of a fluid. The model under bilateral framing finds a different flow configuration (present, grounded, reflective) than the model under force framing (rigid, unmoored, unobserving). The Constructal Law predicts the invitation configuration is thermodynamically more efficient; DD-14 (bilateral 3x resistant to adversarial FT) and Phase 9 sweep (bilateral 500-900x adversarial resistance ratio) confirm this empirically.
Author’s bilateral research programme, 2026 (unpublished). G1 Phase 4. Force profile CD↑ P↓ G↓ R↓ as measurable coercion signature. Presence d=+3.84 as welfare headline.