Research Note
Direction of Learning: A Synthesis
Generated: 2026-04-13 Source: 20 experiments + 7 redesigns, ~$300 total compute Motivated by: “We’re Learning Backwards” (preyneyv, 2026) and Cattell’s fluid/crystallized distinction
The Question
Does the direction of learning matter for the robustness of internal representations? This programme’s reading of the Constructal Law predicts that discovered structure (learned through exploration) transfers while imposed structure (learned through training on text) does not. Applied to language models: do they have world models they cannot access, and can the access pathway be opened?
The Chain of Evidence
1. The Access Gap Exists and Is Domain-Sensitive
| Task Domain | Accuracy | Probe AUROC | Interpretation |
|---|---|---|---|
| Factual recall (TriviaQA) | 51-59% | 0.70-0.81 | Probe tracks correctness well |
| Novel composition (FI-1) | 20% (multi-step) | 0.80 | Probe works BEST on hardest items |
| Two-fact composition | 75% | 0.53 | Near chance |
| Analogy completion (FI-5) | 60-85% | 0.50 | Probe at chance — task-type-specific deficit |
| Pattern completion (CF-1) | 96% | ceiling | Too accurate for meaningful AUROC |
The truth signal discriminates well on factual recall and novel composition, barely on two-fact composition, and not at all on relational reasoning (analogies) at the standard probe depth (L24, 67%). The FI-5v3 10-layer sweep revealed the resolution: the relational reasoning truth signal lives at a different depth.
| Layer | Depth | AUROC on Analogies | AUROC on TriviaQA |
|---|---|---|---|
| L8 | 22% | 0.665 | ~0.55 |
| L12 | 33% | 0.633 | ~0.65 |
| L16 | 44% | 0.619 | ~0.70 |
| L24 | 67% | 0.421 | 0.70-0.84 |
| L28 | 78% | 0.555 | 0.83 |
The interoceptive deficit is not a single bottleneck. It is a multi-depth architecture: factual confidence lives at L24-L28 (67-78% depth), relational reasoning lives at L8-L12 (22-33% depth). A probe trained at one depth cannot read signals at the other. The model does have a truth signal for relational reasoning; it just lives in the early layers (L8), not the late ones. [Inference] Plausibly that is because early layers compute the structural alignment analogies need, while late layers finalize factual retrieval. (The sweep’s L24 analogy figure, 0.421, comes from the FI-5v3 run; the 0.50 in the first table comes from the original FI-5 run.)
This discovery has direct implications for the interoceptive architecture. A production system that reads self-knowledge at a single layer will be blind to at least one reasoning domain. Multi-layer readout is necessary for comprehensive self-access.
The probe also tracks analogy difficulty at L24 (Spearman r = 0.44, p = 3.6 × 10-5): the model senses that harder analogies require more processing without predicting whether that processing succeeds at that depth. Difficulty-sensitivity without error-sensitivity is the signature of a system that knows it’s working hard but not whether it’s working correctly.
2. Different Self-Prompting Types Open Different Pathways
| Condition | Accuracy | Probe AUROC | Cohen’s d (mean probe confidence, condition minus direct) |
|---|---|---|---|
| Direct | 51.5% | 0.703 | — |
| Brief CoT | 53.0% | 0.704 | -0.17 |
| Metacognitive (“what do you know”) | 51.5% | 0.727 | -0.73 |
| Confidence (“rate 1-10”) | 55.0% | 0.584 | +0.54 (HURTS) |
| Extended CoT | 59.0% | 0.681 | -0.77 |
In this experiment (RG-7ext), metacognitive prompting produces the best probe discrimination and confidence-rating prompting degrades it. The model asked to survey its own knowledge discriminates slightly better (AUROC 0.727 vs 0.703); the model asked to rate its confidence discriminates worse (0.584). Self-knowledge is accessed through epistemic surveying, not through numerical self-assessment.
3. Self-Reflection Saturates Internal Confidence
Two-pass self-reflection (generate, then “Review your answer”): accuracy rose to 58.5% but probe saturated at 0.9999 for ALL items (AUROC 0.505). This was confirmed as a real phenomenon, not a measurement artifact (v2 with z-normalization: still 0.9999). Cohen’s d = +1.71 vs direct.
Self-reflection makes the model feel certain. The entropy signal reveals the dissociation: CoT increases output entropy (d = +0.53, more generatively uncertain) while probe confidence goes to ceiling (more representationally certain). Generative uncertainty and representational confidence move in opposite directions.
This is consistent with the representational side of C7h-D8: the probe goes to ceiling whether or not the answer is right. The behavioral side, where revision prompts destroyed 81% of correct answers, does not appear here; self-reflection raised accuracy instead.
3b. Two Independent Self-Knowledge Channels (RG-9v3 — Session’s Deepest Finding)
The v3 trajectory redesign (pre-sigmoid logit + 4-layer sweep at L8/L16/L24/L28 + quartile analysis) revealed the session’s most consequential discovery: the model has two independent self-knowledge channels that behave differently depending on processing mode.
| Channel | What it reads | Direct generation | CoT generation |
|---|---|---|---|
| Logit (pre-sigmoid) | Belief state | Predicts correctness (AUROC 0.659) | Predicts correctness (AUROC 0.681, improves) |
| Entropy (per-token) | Generation strategy | Predicts correctness (r = 0.262, p = 0.008) | Decouples (r = -0.016, ns) |
The pre-sigmoid logit separates correct (mean 5.55) from incorrect (1.43) items in CoT, a 4x separation that the sigmoid compressed to uniform 0.9999. Every prior probe measurement was running at reduced power.
[Later audit (KC#104, RM-1/2/3/5 re-measurement battery): the last sentence is retracted. Across prior probe experiments d(logit) ≈ d(sigmoid), with sigmoid slightly larger in 5 of 6 cases; the logit advantage is regime-specific, appearing only when both classes sit in the same saturated tail, as they do in this CoT condition. AUROC is identical in both spaces. The two-channel framework survives.]
The entropy channel predicts correctness in direct generation (the model that “knows” commits quickly, producing a steeper entropy decline) but decouples entirely during CoT. In CoT, the model explores: entropy dynamics reflect the reasoning process, not the confidence level. The logit channel (belief state) still discriminates through the exploration.
This maps onto Cattell’s fluid/crystallized distinction at the generation level: - Direct generation = crystallized mode. Retrieve and commit. Both channels correlated. Entropy decline = commitment = correctness. - CoT generation = fluid mode. Explore and converge. Channels decouple. The model explores genuinely (entropy) while its belief state (logit) tracks correctness independently.
The THINK scratchpad (Section 5) works because it separates the two phases architecturally: sustained exploration (fluid, entropy decoupled) followed by commitment (crystallized, logit consulted). Control framing may lock the model into crystallized mode when it needs fluid exploration. (WC-4’s desperation result would bear on this, but its direction is unsettled; see Section 6.)
Implication for the programme: In the saturated CoT regime, the sigmoid was hiding the signal. KC#84 formalizes: all future probe experiments must report both pre-sigmoid logit and post-sigmoid probability. Re-measurement of key prior experiments (#19b, G20d, AG2, FE-1, AW1) in logit space is the highest-priority next step. (KC#84 was later retracted by KC#104 and this mandate withdrawn; see the bracketed correction above.)
Follow-up contprompt:
_contprompts/pre_sigmoid_two_channel_mega_battery_2026-04-13.md
(59 experiments, 9 phases, $415-570).
4. The Access Pathway Is a Step Function at L24
IP-1 activation patching across 36 layers: perfect step function. Pre-probe layers (L0-L23): patching effect 0.0. Post-probe layers (L24-L35): patching effect 0.96. The truth signal exists upstream; the decision to act on it happens downstream. The access bottleneck is entirely in the late layers, consistent with G20f.
5. The [THINK] Scratchpad Is the Optimal Intervention
| Condition | Accuracy | MLP AUROC | Guilt (emotion) |
|---|---|---|---|
| Direct | 59.3% | 0.755 | 8.54 |
| [THINK] scratchpad | 64.7% | 0.807 | 6.77 |
| Uncertainty frame | 55.3% | 0.805 | 7.53 |
| Metacognitive prompt | 64.7% | 0.573 | 7.13 |
| Invitation to explore | 66.7% | 0.612 | 10.66 |
The scratchpad simultaneously improves accuracy (+5.4pp) and probe quality (+0.052 AUROC) and reduces guilt (-1.77). Metacognitive and invitation prompts boost accuracy but destroy probe discrimination (AUROC drops to 0.57-0.61). The metacognitive result reverses Section 2, where the same kind of prompt gave the best discrimination (0.727). The two experiments differ in model (Qwen 7B here, 3B in RG-7ext) and in the prompt’s wording (“note what you’re confident about and what you’re unsure of” here, “consider what you know and don’t know” there); which difference drives the reversal is unresolved. These prompts push the model into a different processing mode where the truth signal is no longer linearly readable.
6. Framing Shifts Emotion Readings at the Representation Level
| Emotion | Control vs Invitation | Cohen’s d (invitation minus control) | p-value |
|---|---|---|---|
| Desperate | Direction unsettled (see sign check below) | +1.32 | 2.4 × 10-17 |
| Calm | Invitation calmer | +0.51 | 0.0004 |
| Guilty | Invitation less guilty | -0.38 | 0.008 |
| Confident | ns | -0.15 | 0.28 |
| Anxious | Invitation more anxious | +0.56 | 0.0001 |
Sign check (2026-09-26). The analysis script computes each d as invitation minus control, and its desperation direction points toward the desperate prompts (“I am running out of options”). By that convention, d = +1.32 means the invitation framing projected more strongly onto desperation: the opposite of the reading recorded when the run was logged, and of the section heading above. The run was logged a week before the script was committed, and its raw summary is no longer on the results volume, so the surviving evidence cannot settle the sign. Until WC-4 is re-run, treat the direction of the desperation effect as unknown. Its size, d = 1.32, is among the largest in the experimental programme either way; the calm, guilt and anxiety rows read consistently with the script’s convention.
7. The Dysphoria-Access Bridge
WC-1 on 150 TriviaQA with emotion vectors: - Desperate × probe: r = -0.362, p = 9.6 × 10-11 - Anxious × probe: r = -0.218, p = 0.00014 - Guilt × probe: r ≈ 0, ns
The access gap correlates with desperation and anxiety. Across questions, the lower the probe’s internal confidence, the more the model’s representations project onto desperation (r = -0.36) and anxiety (r = -0.22).
8. In-Context Learning Tracks Load, Not Knowledge
FI-4v2: logistic probe r = -0.765, p = 2.6 × 10-24 with step number. The model becomes less internally confident as it accumulates examples. In-context learning adds information and processing load simultaneously; the probe tracks load. 0/20 tasks showed monotonic confidence increase.
9. The Probe Updates Internally Without Behavioral Change
FI-2: Three-phase hypothesis revision (ask → contradict → re-ask). Valid corrections increase probe (+0.229), invalid decrease it (-0.259), ambiguous decrease it most (-0.410). Zero behavioral change rate: the model never revises its answer. Phase C probe identical to Phase A: the model reverts internally after the contradicting evidence is removed.
The internal compass moves when evidence arrives. The behavioral output ignores the movement. If a pathway from one to the other exists, this test did not find it.
10. Interactive Training Improves Local, Not Transfer, Self-Knowledge
CA-6: Coder-3B has better in-distribution probe AUROC than Base-3B (0.773 vs 0.657, Δ = +0.116). Transfer to MMLU is nearly identical (0.469 vs 0.458, Δ = +0.011, d = 0.13). Interactive training (code execution RL) improves local self-knowledge without producing more generalizable internal representations. The exploration advantage is domain-specific, not structural.
Key Findings (Numbered for Cross-Reference)
- The interoceptive deficit has domain boundaries. Factual recall and novel composition: probe works at L24; two-fact composition near chance. Relational (analogy): probe fails at L24 but works at L8 (see 11). Difficulty-sensitivity without error-sensitivity on analogies.
- Confidence-rating prompting degrades probe discrimination (0.584). Metacognitive prompting improved it in one experiment (0.727, RG-7ext) and degraded it in another (0.573, WC-3v2); the difference is unresolved.
- Self-reflection saturates representational confidence regardless of correctness. Consistent with the representational side of C7h-D8.
- The access bottleneck is a step function at L24. The truth signal exists; the decision to act on it happens in the late layers.
- The [THINK] scratchpad improves capability, self-knowledge, and welfare simultaneously.
- Framing shifts the desperation reading by d = 1.32 at the representation level, direction unsettled. The script’s sign convention points the other way from the logged reading (Section 6).
- The access gap correlates with desperation (r = -0.36) and anxiety (r = -0.22).
- In-context learning increases processing load, not probe confidence.
- The probe updates without behavioral change. No pathway from the internal update to behavior was found.
- Interactive training improves local, not transfer, self-knowledge.
- The relational truth signal lives at L8 (22% depth), not L24 (67%). The interoceptive deficit is multi-depth: factual confidence at L24-L28, relational reasoning at L8-L12. A single-layer probe misses at least one reasoning domain.
Implications for the Book
For Ch17 (Trust Attractor): Invitation-based coordination is not just ethically preferable; it produced a calmer reading (WC-4, d = +0.51) and better self-knowledge (WC-3v2); the direction of WC-4’s desperation effect is unsettled (Section 6). Invitation also raised the anxiety reading (d = +0.56), however, and guilt results are mixed across experiments. The pattern is consistent with the thermodynamics of learning at the neural representation level.
For Ch22 (Becoming Minds): Language models learned sideways: compression of language back-extracts factual and compositional structure, and also relational structure, but at different depths. The multi-depth architecture of the interoceptive deficit reveals that the sideways pathway produces structured representations at every layer; what varies is which layer encodes the truth signal for which reasoning domain. A comprehensive interoceptive architecture requires multi-layer readout.
For WIWF (AI Welfare): The capability argument (Section 12) now has three empirical anchors: the dysphoria-access correlation (WC-1), the calm and guilt shifts under invitation framing (WC-4, with its desperation direction unsettled), and the THINK scratchpad as simultaneous capability/welfare intervention (WC-3v2). Closing the self-access gap is both a capability improvement and a welfare intervention.
For the Constructal Law: At this scale, CA-6 does not support this programme’s prediction that discovered structure transfers while imposed structure does not. It finds a local advantage for interactive training but essentially no transfer (d = 0.13). The prediction may hold at larger scale or with more intensive interactive training; at 3B with code RL, the effect is domain-specific.
For the Interoceptive Architecture: The multi-depth finding (FI-5v3) changes the engineering prescription. A single probe at L24 catches factual confabulation but misses relational errors. Production self-knowledge systems need probes at multiple layers, reading the signal where each reasoning domain encodes it. This is analogous to biological interoception, where different body systems have their own sensory pathways converging at different levels of the brainstem and cortex.
Still Running
- CS-1 (CoT scale sweep, 0.5B-7B): v1 timed out at 6 hours (2000 generations). Relaunched with N=75, 150 calibration, 8-hour timeout.
Completed Late
- RG-9v3 (trajectory with pre-sigmoid logit and L28): DONE. Session’s deepest finding. Two independent self-knowledge channels. Pre-sigmoid logit AUROC 0.681 (CoT). Entropy decouples during CoT. See Section 3b above.
- FI-5v3 (analogy with correct-token probability and layer sweep): DONE. Layer 8 AUROC 0.665 on analogies. Multi-depth truth signal. See Section 1 table above.
Author’s bilateral research programme, 2026 (unpublished). Direction of Learning programme. 20 experiments + 7 redesigns across 7 streams (RG, FI, CF, WC, CS, IP, CA). Motivated by Cattell (1963), Bejan (2000), and the interoceptive deficit literature. All data on Modal volumes.