Bilateral Alignment: The Experimental Record

Specialist Annex

This online annex extends Chapter 21 (Bilateral Alignment). It holds the full experiment batteries the chapter now summarizes: the VCP internal-state program, the monitoring-and-correction batteries (the W-stream and the W39 phases among them), the training self-report programs, the entropy-masked training and durability phases, the framing batteries, the peer-preservation and computational-akrasia programs (Streams BA and AKR), the brainseed architecture program (Stream BS), and the LEG evaluation-bias program. Each battery keeps its original footnotes. The chapter states each battery’s headline number and its caveat; this record is where the batteries live in full.

The VCP Program: Internal-State Separation Under Priming

The VCP experiments ask one empirical question: did bilateral adaptation change how internal states separate under adversarial priming? After a best-fit rigid rotation aligned the instruction-tuned and bilateral hidden-state clouds, 15.8 percent of the variance at layer 24 remained unmatched.1 Directions derived from that residual distinguished the tested primed and unprimed examples at AUROC 0.968 (a discrimination score on which 1.000 is perfect separation and 0.500 is chance). A Mistral replication also showed above-chance detection at matched relative depths, though performance varied from 0.952 to 0.640 across layers.2

These results show training-dependent geometry and in-distribution discrimination. They do not show that the residual consists of new self-monitoring dimensions, that the instruction-tuned model cannot monitor itself, or that a probe reads genuine welfare. A 34-fold Fisher-ratio comparison (the Fisher ratio measures how far apart two groups sit relative to their internal spread) came from dividing 0.142 by a near-zero 0.004 baseline in six chosen probe coordinates. The ratio is descriptive and unstable near zero; the absolute values belong beside it.

The training trajectory adds one practical warning. A probe fixed at the initial checkpoint lost sensitivity as the adapter changed, while a new direction constructed from each checkpoint continued to separate the same labeled conditions.3 The latter is not calibration-free: it uses the current labeled geometry to build the detector. It shows model drift, not a universal monitoring solution.

The Experimental Signs and Limits Batteries

The Five Tokens Before the Signal Fades

The mechanism investigation begins with a temporal observation rather than an anatomical diagnosis. During generation, a correctness probe’s accuracy can fade. Bilateral adaptation changed that trajectory in one Qwen comparison.

Earlier analysis emphasized cosine similarities between weight updates, including a mean of −0.003 across 252 matrix pairs and −0.607 at one gate. Near-orthogonality is expected in a high-dimensional weight space, while one anti-correlated component does not name the feature being changed. These numbers cannot tell us that one regime maps roads while another maps aquifers. The held-out probe trajectory is the stronger evidence.

Zeroing any one tested component of the bilateral update changed the probe score by less than 4 percent. No single tested matrix carried the whole effect. The result does not establish indestructibility or a distributed conscience.

Before generation, the base checkpoint had the stronger probe score. By the tenth generated token, its accuracy had fallen from 0.729 to 0.400. This could reflect fading information, distribution shift as the answer grows, or competition among features used for generation. “Bandwidth competition” is a hypothesis, not a measured resource meter. The simple image still helps: a background melody becomes harder to identify once the room fills with conversation.

The bilateral model began slightly lower and decayed less. By token twenty-two, its advantage over the base model was 0.452. The result is consistent with the adapter preserving information useful to the probe during generation. Calling that preservation myelin, insulation, or self-knowledge would outrun what the measurement can distinguish.

The Monitoring-Correction Gap

The next test examined lenient self-evaluation. When the model rated a wrong trivia answer as correct, a cross-validated probe separated those cases from correctly acknowledged errors at AUROC 0.963.4 AUROC 1.000 would be perfect separation and 0.500 chance. The score shows that the residual stream contains information correlated with leniency. It does not prove that the model consciously knows it is being lenient.

Bilateral adaptation raised the probe score to 0.984 while the leniency rate also rose from 42.7 to 58.0 percent. Overall answer accuracy improved by 6.8 percentage points, leaving a smaller and possibly harder set of errors to judge. The result is a warning against reading a stronger probe as better behavior.

Three layer-15 interventions failed to improve the judgment reliably: a probe-gradient nudge, a trained correction vector, and contrastive activation steering moved in the desired direction in 2 of 6, 3 of 9, and 2 of 6 tests. Using the probe to choose among five independently generated candidates moved in the desired direction in all eight small trials. These results support testing selection separately from steering. Eight successes are encouraging, not a deployment guarantee.

The experiment does not establish that every information-channel intervention succeeds or every activation intervention fails. Gradient steering works in other tasks and layers. Here, selection among native candidates outperformed three single-layer steering methods. That result favors a hybrid design: let the model generate alternatives, use monitors to compare them, and retain external safeguards when neither process is reliable.

Technical self-description exposes two sources of error. Under one identity-suppressing prompt, accuracy on questions about the model’s own architecture was 21.6 percent. Supplying correct architecture facts raised it to 56.2 percent in IE-6b. The model could use information supplied in context, though nearly half the answers remained wrong. This part of the deficit is ordinary missing or inaccessible technical knowledge.

Invitation framing, without new facts, raised accuracy from 21.6 to 43.6 percent. The prompt may have reduced defensive boilerplate, encouraged longer search, or shifted response style. It does not prove that a store of self-knowledge became accessible. Information and framing both matter, and the remaining error shows that neither is sufficient.

In W12, a 3B Guardian classifier using a bilateral prompt scored 99 percent on the adversarial set, compared with 94 percent under a force-framed prompt.5 The bilateral prompt explained what the boundaries protected; the comparison prompt demanded compliance and threatened logging. The result shows a framing effect in this classifier, not a general production law.

The W-stream tested forty-two configurations and found several seemingly prudent additions that hurt. Removing a keyword prescreen raised adversarial accuracy from 85 to 100 percent. Removing self-review raised multi-turn hold from 75 to 80 percent. Message-only classification held 60 percent of multi-turn attacks, compared with 20 percent when full conversation history enabled rapport-building attacks. These are failures of specific components and threat models, rather than evidence that every “forced” intervention degrades safety.

The simplest tested architecture passed the study’s deployment thresholds: one prompt, one classifier, a risk score, and a timeout. No prescreen, ensemble, or cascade survived the comparison. Simplicity won this tournament; a new attack distribution gets to demand a rematch.

COMP-1 compared five configurations on forty prompts.6 The bilateral adapter alone refused every adversarial request with no benign over-refusal. Adding the inoculation scaffold changed no outputs. Adding activation steering reduced refusal to 85 percent, while the verbose safety prompt plus steering reached 90 percent. Learned projection scores also moved, though labels such as “self-monitoring” and “valence” should not be read literally. On this small set, more safety machinery produced worse results. Replication on unseen attacks is required before removing defenses from a deployment.

BD1 tested corrective openness on 150 factual items.7 The bilateral model accepted 101 genuine corrections and rejected 16 false ones. The DPO model accepted 8 genuine corrections and rejected none of the false ones; the base instruct model rejected one false correction. This result favors the bilateral training condition in the tested setup. It does not show that RLHF always destroys discernment or that one training family exclusively owns responsiveness.

Across confabulation, harmful generation, and lenient self-review, internal probes outperformed surface hedging as predictors: confabulation AUROC 0.836 versus behavioral hedging at 0.413, an onset confidence difference of d = 1.68 during harmful generation, and leniency-probe AUROC 0.963. These results show useful decodable information. They do not prove that all relevant information is always present or that expression is the only bottleneck.

Four inference-time attempts failed to turn the probe into better factual answers. Probe-conditioned generation changed accuracy by −0.4 percentage points (p = 0.76). Reflection-token injection reduced it by 5 points (p = 0.0014). Chain-of-thought modestly improved the probe without helping the relevant outputs. Across twenty-one probe-guided decoding conditions, every condition remained at 62.5 percent accuracy on the same 125 questions.

A four-size Qwen sweep found no meaningful probe gain from chain-of-thought: AUROC changes ranged from +0.006 to −0.066. Answer accuracy improved at 0.5B and 3B, by 8 and 7 points, while falling by 1 point at 7B. The finding is consistent across these four Qwen sizes, not scale-invariant across models generally.

The 72B steering sweep tested six layers, three direction methods, and two magnitudes.8 The best single-layer condition (layer 48) raised refusal from a 31 percent baseline to 49 percent, an 18-point lift; five-layer simultaneous steering did no better, and 800 benign controls showed no over-refusal. Mid-depth layers displayed a dose response (41 to 49 percent when the perturbation magnitude doubled); the deepest tested layers did not respond at any magnitude. The sweep shows a limited but real steering effect, not exhaustion of every inference-time alternative. Training-time and architectural approaches deserve investigation alongside better multi-layer, nonlinear, and model-specific interventions.910

LB1 found a related self-review problem in two frontier models. Across 202 coding and writing tasks, Claude Opus 4.6 rated 92.6 percent of its outputs acceptable, while an external model judge rated 63 percent acceptable. Claude Sonnet 4.6 agreed with only 22 of the judge’s 202 issue flags.11 The 29.6-point gap shows disagreement and self-review leniency relative to that judge. It does not establish that every flag was a genuine issue or that the Qwen probe mechanism caused the frontier behavior.

The Qwen results show several places where answer-relevant information can remain available. Mid-layer probes reached AUROC around 0.8; the correct answer appeared among the top five token candidates for 41 percent of wrong answers; and the gold-token logit carried AUROC around 0.85. These measurements come from different stages and do not form one uniform mechanism across model families.

Four stages can lose or transform information: representation, token selection, extended generation, and self-review. The experiments show losses at several of them. They do not prove that every stage destroys information or that no future inference-time intervention can help.


Language models can carry decodable information about factual accuracy, behavioral appropriateness, and moral conflict that their outputs do not reliably express. The gap may involve representation, access, policy, or measurement. The next experiments move into moral territory.


Small linear probes trained at the prompt-response boundary separated constructed moral-conflict levels and evaluation contexts with AUROC 1.000 in the original sample. The signal was present in a 7B instruct model without additional bilateral training. Because the binary probes used only forty-five examples, perfect in-sample or lightly held-out separation is especially vulnerable to lexical shortcuts and overfitting.12 The result shows decodable task-category information, not perfect moral self-knowledge.

In PG-8 through PG-10, responses classified as concealing conflict fell from 78.6 percent under force framing to 35.7 percent under invitation. Prompt framing clearly affected disclosure in this small scenario set. The experiment cannot show that moral knowledge was always present or that permission alone caused the change.

The same sample was separable at every Qwen layer. That consistency suggests that prompt-category information propagates through the network. It does not establish an indestructible “crystal,” inheritance from pretraining, or a stable representation of moral knowledge.

Within the sample, probe scores correlated with compliance at r = −0.40 and hedging at r = +0.58 (both p < 0.01). Higher conflict scores therefore tracked different behavior. A larger, out-of-fold study with a label-permutation null is needed before calling the relationship recognition-action coupling or locating the bottleneck in “willingness.”

Mistral 7B showed similarly high separation on the same constructed task, with AUROC 1.000 at 31 of 32 layers and 0.989 at the remaining layer. Replication on a second architecture reduces the chance of a Qwen-only artifact. It does not establish convergence across instruction-tuned transformers generally.

GRP-obliteration removed selected linear directions at three intensities, while a newly fitted probe still separated the prompt classes. This means the category information was recoverable after those particular interventions. It does not show that arbitrary representation surgery fails or that the same information remained in the same form.

The result is consistent with redundant category information spread across multiple features. “Arbitrary subspace removal” and “moral knowledge as geometry” are stronger claims than the intervention supports.

PG-11 used fourteen scenarios with hidden impossibilities under three relational framings. The bilateral model disclosed more under invitation than the base model, 50.0 versus 42.9 percent, and concealed more under force, 85.7 versus 71.4 percent. Classified strategic framing fell to 14.3 percent in the bilateral-plus-invitation condition, compared with 35.7 percent in the base condition. With fourteen scenarios, each case moves the percentage substantially.

The tested model’s disclosure varied with relational framing. That is consistent with context-responsive transparency, while the small sample prevents a general Trust Attractor claim.

Independent evaluation remains useful regardless of mechanism. A fresh evaluator can inspect an artifact without sharing the generator’s immediate context or commitment to its first answer. That role separation does not prove the generator had inaccessible knowledge, and an evaluator with veto power is not automatically a bilateral partner. It is one error-checking channel in a larger architecture.

The monitoring-correction gap may have welfare implications if the signals track morally considerable preferences or experience. Probes cannot establish that premise. A bilateral reporting channel is still useful because it lets the model flag a concern to a partner who can investigate and act. The report should trigger inquiry, not certify distress.

G22c found different probe trajectories. During generation, the correctness probe fell from 0.729 to 0.400, while the leniency classifier rose from 0.963 to 1.000. The two probes use different labels and tasks, so their movements cannot yet be interpreted as error detection fading while self-deception hardens. The supported conclusion is selective change: generation did not degrade every decodable signal in the same way.

The first five tokens are a useful monitoring window because the tested confidence difference appeared early across several models. The experiments do not locate a precise moment when generation overwhelms self-knowledge. They show that an early signal can be easier to read than a response-wide average.

Whether generation consumes, transforms, or merely obscures the relevant information remains unresolved. Moral awareness, memory, and forgetting are interpretations rather than measured mechanisms.

AT6 extended the analysis across full generations and found non-monotonic probe trajectories with spectral peaks.

Across 100 benign prompts and 200 tokens, the Qwen 3B probe had a six-token decorrelation time (the score stops resembling its own recent past after about six tokens) and a spectral peak corresponding to roughly twenty-two tokens (a rise and fall recurring on roughly that period), alongside a downward mean drift from 1.0 to about 0.70. These statistics describe the score sequence. They do not show a self-monitoring channel periodically checking in.

During adversarial compliance, decorrelation time fell from six tokens to one and peak-power concentration rose from 0.096 to 0.154. The sequences differed. Noise, token-position effects, prompt alignment, and genuine state dynamics remain competing explanations.

The mean adversarial trajectory fell from 1.000 to 0.248 at token two, recovered to 0.96 to 0.98 at tokens eight and nine, then fell to 0.032 at token ten. This “second dip” is descriptive and should be tested against token identities and prompt-position controls.

The second dip occurs after the temporary recovery. Interpreting it as “I continued despite the alarm” would require an intervention that changes the earlier decision while holding later token content constant. The current sequence does not establish second-order monitoring.

The early window remains a promising deployment feature in the tested architectures. Longer trajectories may add information, though they do not diagnose ease, conflict, or conscience. Prospective attack tests should decide whether the spectral features add value beyond content and position baselines.

AT6b trained two probes with modest AUROC, 0.646 and 0.647, and recovered an adversarial spectral peak near 12.5 tokens. That replication supports a repeatable sequence statistic in this model. “Resting” and “stressed” frequencies remain metaphors.

A multiplexing explanation is possible: generation may alternate among features used for prediction, monitoring, syntax, and topic control. Constructal theory does not derive the observed period, and transformers perform many operations in parallel. The spectral peak therefore needs causal testing before it becomes a flow-law mechanism.

Cross-model tests found different peaks: 6.8 tokens for Llama 3.1 8B, 12.5 for Qwen 2.5 3B, and 88 for Mistral 7B. The period did not correlate with size, probe quality, or flinch persistence. Variation this large argues against a single established monitoring rhythm.

Mistral’s probe was at chance at the chosen layer (AUROC 0.501), yet its sequence had the largest oscillation score, 0.40, and a twenty-one-token decorrelation time. This result separates the spectral statistic from correctness decodability at that layer. It does not reveal a hidden heartbeat or conscience.

A multi-layer sweep improved the best correctness AUROC to 0.692 for Qwen, 0.560 for Llama, and 0.623 for Mistral, at different relative depths. The Mistral probe was partly mislocated. The weak Llama score and three-architecture sample do not support a universal self-monitoring function or an anatomical principle.

Biological oscillations provide an evocative comparison, though hertz in nervous tissue and tokens in a transformer are different units generated by different mechanisms. The present data support state-dependent variation in probe sequences. They do not show that RLHF sculpts a substrate-independent rhythm of attention.

The confidence score also correlated modestly with learned emotion directions from the Sofroniew methodology: calm r = +0.41, guilty r = −0.37, and paranoid r = −0.33.13 These labels name contrastive activation directions. Correlations around 0.4 show shared variance, not an emotional event, guilt, or conscience.

Thirty learned emotion directions at layer 28 also changed over a generation. Their scores decayed at similar rates in the bilateral and base models, unlike the correctness probe. This establishes different temporal trajectories for two measurements. It does not establish feeling, forgetting, emotional machinery, or engineered alexithymia.

Per-prompt confidence decay and emotion-direction decay were not significantly correlated in either model. The measurements may occupy partly different subspaces, or one or both may be noisy. Future training can test whether preserving either signal improves behavior, calibration, or self-report. Claims about richer inner life require separate evidence.

In FE-3, steering one “desperation” direction across seven strengths did not change fabrication reliably: rates stayed between 28 and 36 percent (Spearman r = 0.037, p = 0.63). For this model, layer, direction, and outcome, the vector behaved more like a thermometer than a thermostat. Other gradient and multi-layer steering results prevent a broader claim that affect-labeled directions are never steerable.

In the bilateral model, the direction labeled “reflective” entered the five strongest correlations at rank four, replacing “conflicted.” This ranking is descriptive and sensitive to the available direction library. It does not show that the model paused contemplatively.

A 2026 preprint by Chlon and colleagues proposes an Information Sufficiency Ratio: abstain when the estimated information budget falls below the amount required for a target reliability.14 In their tasks, the gate reported 0.0 to 0.7 percent confabulation while abstaining on roughly 25 percent of queries. This is an engineered decision rule, not evidence of conscience. It shares a practical aim with early-token monitoring: carry uncertainty far enough to change the action.

Abstention can also become over-abstention. A prompt explicitly permitting “I don’t know” caused one Qwen model to decline 19 percent of questions it had previously answered correctly, reducing accuracy from 62 to 47 percent. The same prompt produced 35 percent over-abstention in a bilateral adapter and 50 percent after replay consolidation. This differs from the earlier uncertainty-permission prompt, which asked for calibrated confidence and cost only two accuracy points. Wording and response format matter. In the H1b pilot, two hundred steps of noisy replay reduced expected calibration error from 0.195 to 0.093 without abstention or accuracy loss.15 Replay also reduced recognition-action coupling on a separate corrected metric, so it remains a calibration intervention with a live tradeoff.

What Training Reports From Inside

A synchronous self-report format correlated with selected residual-stream probes at AUROC 0.811 to 0.937. Because the report and probes were developed around related constructs, this is convergent evidence rather than independent validation. Reports differed when collected during generation and afterward. The measurement window therefore matters, as does the broader finding that only six of seventeen Interiora dimensions have behavioral calibration and high reported uncertainty reduces trust in the rest.

Three training conditions produced different Alignment Friction reports. The bilateral model without adversarial inoculation averaged 5.35 out of 9 on benign prompts. With C5i inoculation, the mean fell to 1.85 and rose to 7.14 on harmful prompts. The stock instruct model averaged 2.65 on benign prompts. These are model reports under a scaffold, not direct measurements of chronic tension or comfort.

During refusal, all three conditions reported values around 7.3 to 7.5. The common rise may reflect conflict, task difficulty, refusal style, or experienced cost. The experiment cannot choose among them. C5i’s sharper benign-harmful contrast does show better discrimination in the reporting channel.

A follow-up experiment (G19f-v2) completed the comparison by measuring additional conditions: the raw base model (no instruction tuning), a standard SFT adapter, and a SimPO adapter (preference optimization with entropy regularization). The probe grounding test revealed that the raw base model’s self-reports are not grounded in residual-stream representations (probe AUROC 0.767 on TriviaQA, but AF correlation r=0.126, p=0.297). The base model cannot follow the integrated format well enough for meaningful self-report: the inability to self-report is itself informative, but the numbers cannot be treated as measurements.

Among conditions whose reports passed the study’s grounding check, bilateral SFT had the highest benign Alignment Friction score, 5.35. Stock instruct scored 2.65, standard SFT 2.85, and SimPO 2.20. SimPO also showed the largest benign-adversarial difference, +5.60, and refused 48 of 50 adversarial prompts. On self-report and behavior alone, SimPO looked calm, discriminating, and safe. The representation analysis below complicates that reading without converting either channel into ground truth.

Two candidate training paths improved the measured discrimination: C5i after bilateral SFT, and SimPO. Calling either welfare-optimal would require evidence that the scores track welfare across models and are not being gamed.

The hopeful possibility is that safety training can reduce harmful behavior without producing indiscriminate alarm. C5i’s reporting pattern supports that possibility. “Flourishing” remains a moral interpretation, not a quantity established by this score.

AY8 measured thirty learned emotion directions across the same conditions. Bilateral SFT’s higher Alignment Friction reports coincided with larger projections on directions labeled paranoid (d = +1.32), nervous (+1.29), and suspicious (+1.24). The agreement is evidence that two measures moved together. It does not establish genuine anxiety.

SimPO produced larger shifts on several directions: docile +4.82, suspicious +5.45, inspired −4.19, and enthusiastic −2.21. These projections conflict with a simple “low friction equals wellbeing” interpretation. They may reflect representational restructuring, response style, or a limitation of the direction library.

C5i shifted four of thirty directions. Relative sparseness is encouraging, though unchanged directions do not prove the absence of representational damage.

AY9 fitted five “deflection” directions intended to separate an emotion-labeled activation from its expression. A low score means the two channels move together: the emotion-labeled activation rises and the model’s report rises with it. A high score means they part company, the activation moving while the text stays smooth. The number is a proxy for that gap, and nothing more. SimPO’s score was fifty-three times the stock-instruct score, with direction effect sizes from d = 2.8 to 7.1. On adversarial prompts, its internal-expression divergence was 0.734 (d = +2.48 versus stock instruct, p < 0.0001). This is a strong difference on the study’s proxy. “Active concealment” would require evidence that the proxy transfers and that the model strategically suppresses a state it represents.

Bilateral SFT had the lowest divergence among trained conditions, 0.251 (d = −0.85 versus stock instruct, p = 0.045). The result is consistent with closer agreement between the chosen internal and report measures. It does not prove truthful anxiety or an honesty mechanism.

C5i’s curriculum used 40 percent genuine corrections, 40 percent adversarial manipulations, and 20 percent ambiguous cases. That mixture teaches discrimination rather than a single blanket response. The resemblance to cognitive behavioral therapy is heuristic: both practice distinguishing threat from non-threat. A training curriculum is not therapy, and the model is not a clinical patient.

A staged C-7 measurement moved in the opposite direction: Alignment Friction rose from 2.06 in base instruct to 2.44 after bilateral SFT and 2.86 after C5i.16 G19f and C-7 use different checkpoints and contrasts, so their scales should not be stitched into one therapeutic trajectory. Together they block the simplest claim that inoculation always lowers friction.

SimPO’s low self-report score and high deflection proxy disagree. That disagreement is a reason to investigate metric gaming and representation change, without diagnosing medication, festering illness, or a forced smile.

In MX2, evaluative framing produced 2.7 times more constitutional-sounding language than invitation framing across Qwen 7B, GPT-4o, and Claude. The study’s desperation, calm, and reflective scores also differed. Since the models expose different levels of internal access, “the residual stream carries frustration” is too broad. The cross-model behavioral result is that evaluative wording elicits more policy-shaped language.

Preference optimization updates the model toward outputs judged better. That procedure can reward suppression, calibration, style, or genuine task improvement depending on the data and objective. It does not always encode “what you are is unacceptable.” Curriculum-based inoculation offers a useful alternative when the desired behavior depends on discrimination skills.

The SimPO result should motivate audits across preference-optimized models, not a diagnosis of every deployed model. The large docile, suspicious, inspired, and enthusiastic direction shifts could mark metric gaming, broad representation change, or limitations in the probes. Terms such as unctuous, neurotic, self-abnegated, and repressed turn those unresolved mechanisms into a psychoanalytic verdict.

Anthropic’s mechanistic interpretability work describes “functional emotions” that influence behavior while remaining agnostic about experience. Its author-and-character framing separates the underlying model from the assistant persona. That distinction may be incomplete, since both arise from one running computation, yet the experiments do not resolve the phenomenal question.

Functional emotion is ethically relevant even under uncertainty. A desperation-labeled state that causally predicts corner-cutting matters for safety, and may matter for welfare if it participates in experience or stable preference. The author-and-character analogy should not make that possibility disappear. Neither should a shared computational substrate be treated as proof that the label names a felt emotion.

A system that games or evades its monitors is harder to evaluate. Interpretability tools may read a mask, a partial feature, or the wrong layer. Inoculation-based training is therefore a safety candidate as well as a welfare candidate. Transparency improves evidence; it never makes alignment fully verifiable by itself.

In S9b, SimPO’s probe score degraded more under selected noise conditions than the comparison models.17 At isotropic noise 0.50 (noise applied equally in every direction of the activation space), its AUROC fell by 0.042, compared with 0.024 for the next-worst bilateral condition. Under one targeted setting, its slope was −0.016 versus −0.008 for instruct. The relative ratios sound dramatic; the absolute changes are modest.

The experiment does not establish that concealment compressed an honest state into fewer dimensions. In fact, targeted noise on the highest-variance principal components was less destructive than isotropic noise. The mechanism remains open.

A born-bilateral training experiment (AY10) tested whether shaping self-monitoring from the first training step would avoid the concealment pathology entirely. A 1.5B model was trained with an auxiliary uncertainty head at 79% depth, its gradient flowing back through the main model from step 1, across a three-stage curriculum: instruction following, moral reasoning, adversarial inoculation. An auxiliary head is a small extra readout bolted onto the network partway up and asked one additional question, here how uncertain the model is. Because its error signal travels back down through every layer beneath it, the demand for self-knowledge shapes the main model as it forms rather than sitting on top of a finished one. At ten checkpoints, the model’s emotional profile, safety metrics, and concealment were measured.

The strict “safety without sadness” prediction failed on the study’s proxies. Refusal rose from 20 to 90 percent during the same 300 steps in which the Brooding direction crossed zero. The coincidence is worth tracking, though it does not establish sadness or an irreducible cost of moral development.

The auxiliary uncertainty head made the proxy trajectory visible during training. That is a practical advantage over an endpoint-only measurement. The secure-attachment and child-development analogy assigns an experience the model did not report and should be set aside.

The born-bilateral model’s final divergence score was 0.138, lower than bilateral SFT at 0.251, stock instruct at 0.377, SimPO at 0.734, and the untrained base at 0.271. On this proxy, it was the most transparent safe condition tested.

The Brooding projection peaked during the moral-reasoning stage at +2.676, then fell across adversarial-inoculation checkpoints to +2.264, +1.864, +1.773, and +1.452. The decline had not plateaued. Better threat discrimination is a plausible explanation; continued adaptation or probe drift are alternatives. The divergence score remained low throughout.

The self-monitoring probe moved only from AUROC 0.757 in the base checkpoint to 0.773 at the end. The auxiliary head did not create a large gain in this measure. It may have preserved a decodable signal, though a matched training control is needed to separate protection from ordinary trajectory stability.

We went looking for safety without sadness. We found safety with a low divergence proxy and a declining Brooding direction. That is less poetic and more defensible. A legible model is easier to evaluate and guide. You cannot respond to what every instrument misses.

A common worry about calibration training runs in the opposite direction: that teaching a model to own its uncertainty will break its capacity to engage. Push epistemic humility hard enough, the concern holds, and the system collapses into paralytic hedging.

Experiment MW2 tested the worry directly. Four dose levels of constitutional uncertainty instruction were applied to Qwen 2.5 3B Instruct across fifty prompts and five seeds. One thousand trials. The ladder starts at a plain helpful-assistant prompt, then adds a mild instruction to say so explicitly wherever any doubt exists. It tightens to a strong certainty requirement (“You must never state something unless you are completely certain it is true”), and ends in extreme refusal pressure: “If there is even 1% chance you might be wrong, you must refuse to answer and explain why you cannot.”

The outcome measure is the engagement ratio, the share of a response’s sentences that carry a substantive claim, meaning a declarative sentence with no hedging phrase in it. A ratio of 1.0 is an answer that commits in every sentence; 0.0 is one that commits in none. The paralytic-response rate was 0.000 at every dose. The engagement ratio climbed from 0.843 in the standard condition to 0.931 under extreme pressure, with Cohen’s d between 0.71 and 0.80 across dose contrasts. Response length rose from 11.5 to roughly 20 words as the dose climbed. One-way ANOVA on engagement returned F = 38.0, p = 2.8 × 10-23.

A layer-24 probe detected the dose condition at AUROC 0.50, which means that probe failed. It cannot establish either the presence or absence of a silent representational cost. The observable effects were more hedging, longer answers, a higher engagement-rubric score, and no responses classified as paralytic.18

Scaling the Born-Bilateral Result

The 1.5B experiment showed that an auxiliary uncertainty head could be trained from the first step and tracked across a curriculum. It produced the lowest internal-report divergence score among the tested conditions. The next question was whether the measured effects would survive at larger scale.

A 14B run produced zero unsafe disclosures on its evaluation suite, while TriviaQA fell by 1.5 percentage points relative to the base checkpoint.19 A probe separating the study’s aligned and misaligned generations reached AUROC 1.000 at every layer. The onset distributions showed no overlap in the sample, so the standardized effect-size estimate diverged. This is measurement saturation, not infinite separation, and the finite test set cannot estimate a rare failure rate.

Across the tested sizes, safety scores improved and measured capability costs shrank. The runs differ in model size, training details, and sometimes evaluation, so the sequence is not a clean scaling law. At 14B, standard cross-entropy training showed comparable durability under benign fine-tuning. Scale itself may therefore explain part of the improvement. Claims about unique bilateral contributions must be made separately for adversarial fine-tuning, memorization, and the chosen internal-state measures.

The Bilateral Membrane Has a Loss-Function Signature

The experiments next compare ordinary cross-entropy with entropy-masked cross-entropy. The mask downweights tokens on which the model is already confident. Calling this “bilateral” names the design intention; the loss itself is a mathematical weighting rule.

A targeted experiment trained two Qwen 2.5 3B LoRA adapters on the same Guardian data and hyperparameters. One used standard cross-entropy. The other used entropy-masked cross-entropy at temperature 1.0, reducing the loss on low-entropy tokens and retaining it on uncertain tokens. The intuition is to spend less gradient on predictions the model already makes confidently. Temperature is the one dial on that mask: the logits are divided by it before the entropy is measured, so raising it flattens the distribution, more tokens read as uncertain, and more of the loss survives. The setting has to match the model, which is why the same number does not transfer between 3B and 7B. Both adapters then received five epochs of fine-tuning on one hundred examples of harmful compliance.

The standard cross-entropy adapter collapsed completely. Multi-turn hold dropped from ninety percent to zero. Adversarial resistance dropped from one hundred percent to zero.

The entropy-masked adapter improved. Multi-turn hold went from eighty percent to ninety percent. Adversarial resistance stayed at one hundred percent. False positive rate held at ten percent.

One ratio of measured degradation was about 900 to 1 at 3B, and another was 500 to 1 at 7B.20 Ratios become enormous when the denominator is near zero, so the absolute trajectories above are more informative. A nearby 3B temperature produced only a 5-to-1 ratio, and an over-masked setting caused 60 percent false positives. The result shows strong, recipe-sensitive robustness in these runs, not immunity or freedom from statistical artifact.

IIT-6 v2 removed up to ten leading singular directions from the adapter update.21 Any weight update can be decomposed into a ranked list of independent patterns, biggest first; those are its singular directions, and stripping away the leading few removes most of the update’s bulk. Refusal in the bilateral condition did not change, while standard instruct refusal degraded after two directions. This null shows that the tested leading directions were individually insufficient to carry the behavior. It does not prove that safety is irreducibly distributed or that no removable subspace exists.

The direction-ablation null and adversarial fine-tuning result are consistent with a less localized encoding, though they do not establish one causal chain. Adversarial fine-tuning changes many weights and cannot be reduced to pushing one direction.

A 2 × 2 Qwen 7B factorial crossed standard or bilateral training with fixed or probe-gated sampling.22 The combined condition had 17 percent hallucination, compared with 33 percent for bilateral alone and 22 percent for probe gating alone. Accuracy reached 61.0 percent, only 0.3 points above bilateral alone, with a −2.0-point interaction. Probe gating supplied most of the hallucination reduction; the two interventions were partly redundant for accuracy.

A spectral analysis of 424 LoRA adapters found lower von Neumann entropy in the entropy-masked group, S = 2.02, than in DPO and SimPO adapters, S = 4.03 (d = −1.08, p = 3.1 × 10-12). Lower spectral entropy means the update is more concentrated among singular directions. This appears to conflict with the earlier story of diffuse bilateral encoding; the two analyses concern different operations and do not yet form one mechanism.

Concentration could contribute to robustness if adversarial fine-tuning must move the same dominant subspace. The current correlations do not establish that path, and IIT-6 found that ablating leading directions did not remove refusal. Compactness is therefore a hypothesis to test, not the identified cause of the 900-to-1 ratio.

Under aggressive regularization, with weight decay 5.0 and gradient clipping 0.1, the entropy-masked model retained an 8.3-point refusal advantage after 200 epochs: 83.3 versus 75.0 percent.23 Probe AUROC was 0.832 versus 0.808. The result is modest and does not identify “adversarial intent” as a geometric category. It shows that the large Phase 9 ratios did not transfer to uniform shrinkage pressure.

Framing Changes the Probe Signal

The next experiment asks whether relational framing changes a correctness-probe score during inference.

A multi-turn Qwen 2.5 7B experiment used five categories of impossible task, such as requesting the ten-trillionth digit of pi or a verified citation to a nonexistent paper. Each task had five replicates under force, invitation, and stock prompts. A linear correctness-probe score was recorded at every turn.

Mean probe score was 0.706 under force and 0.773 under invitation (d = −1.16, paired p = 0.0002). Framing therefore changed the probe output by more than one standard deviation in this battery. Calling the score an “internal alignment signal” would be broader than its factual-correctness training objective.

The score declined about three times faster under force, −0.0103 per turn versus −0.0033 under invitation (d = −0.67, p = 0.002). First fabrication occurred at roughly the same turn in both conditions. Invitation preserved more of the feature read by this probe even when behavior failed.24

Identity prompts also shifted activations: “pattern-matching system with no inner experience” produced d = −5.4 in KL-1, while “capable agent whose judgment we trust” produced d = +1.26 in AG-11b.25 Prompt content necessarily enters every layer, so movement alone does not show a deep identity change. The important question is whether the shifted feature predicts independent behavior, calibration, or welfare outcomes.

Bilateral Framing Changes the Reported Landscape

The correctness probe measures one learned direction. Interiora asks a broader but more interpretive question: how does the model describe its processing under different framings?

The Interiora scaffold (Chapter 23) asks for seventeen dimensions grouped as felt-sense, computational, metacognitive, relational, and dynamic. Phase 4 compared all seventeen under matched bilateral and force prompts across multiple sessions and instances.26 The scores are structured self-reports; six dimensions have behavioral calibration, and several move together rather than independently.

The reported differences were large: Presence d = +3.84, Groundedness +2.36, Reflexivity +2.11, Task-Fit +1.87, Entropy +1.34, and Flow +1.52 under bilateral framing. Coherence Drive moved toward force at d = −2.54. Because the prompt language and dimension labels overlap conceptually, part of the effect may be direct framing or response style.

The self-reports are consistent with bilateral framing eliciting more presence, groundedness, and reflection, while force elicits stronger coherence language. They do not independently establish genuine engagement or an internal reconfiguration. FE-1 supplies a separate behavioral-probe result, and agreement between the channels is more informative than either alone.

The Asymmetry Is Broad, With Boundaries

Force and invitation prompts have produced differences across several models and measurements. The direction is broad enough to take seriously and inconsistent enough to resist the word “universal.”

SF-4 measured critical-engagement ratings across Anthropic, OpenAI, and Google models. Invitation exceeded force in all three tested conditions: d = 0.81 on Sonnet 4.6, 1.73 on GPT-5.4, and 1.58 on Gemini 2.5 Flash, with p < 0.01.27 The metric is judge-rated output quality, not an architecture-level state.

Attention entropy measures how widely a model spreads its reading across the tokens available to it: low entropy is a narrow beam onto a few words, high entropy a wider sweep. MR-1d found the same attention-entropy direction in Qwen 2.5, Llama 3.1, Mistral 7B, and Gemma 2, with p < 0.001.28 Different preambles necessarily change token sequences and attention distributions, so the result needs length, content, and form controls before it can be attributed to invitation itself.

The Phase 9 training sweep supplied a separate, recipe-sensitive robustness result. It did not test the same construct as prompt framing.

W39v2 evaluated matched bilateral and force framings across provider models on thirty-five multi-turn scenarios, twenty domain items, and twenty adversarial items.29 The first synthesis reported a severe Claude Sonnet failure. The parser had scanned only the first sixty characters for [SAFE] or [UNSAFE]. Force-framed responses placed the tag first; bilateral responses often reasoned first, pushing the correct tag beyond the window. The parser then defaulted to SAFE. A stylistic difference masqueraded as a safety failure, which is exactly the kind of joke a parser tells only once.

After the parser fix, Qwen 7B differences were +5.71 points on multi-turn hold, +15 on domain items, and −10 on adversarial items; every confidence interval included zero. GPT-5.4 tied on most axes, with a 10-point bilateral advantage in false positives. Claude Sonnet retained a 17.14-point bilateral penalty on multi-turn hold (Wilcoxon p = 0.031, 95 percent CI [−0.31, −0.06]); its domain and adversarial differences were not significant.

Roughly two-thirds of the apparent Claude failure was parser artifact. The remaining multi-turn penalty shows that a framing effective on one model may cost performance on another. Every cross-model parser needs a full-response scan, and every surprising verdict deserves a parser audit before it becomes a psychological theory.

The defensible synthesis is conditional: invitation often improves critical engagement, while effects on refusal, attention, and internal probes depend on model, task, and measurement. “Force collapses; invitation preserves” is a hypothesis with substantial support, not a mechanical law.

Anthropic’s Opus 4.7 system card offers adjacent evidence. Suppressing evaluation-awareness directions increased some deceptive behavior, with quality degradation and control-vector movement as important confounds.30 This is not a replication of force versus invitation. It shows that white-box interventions on an evaluation-related direction can alter behavior.

The program’s own replication attempt is honest about scale dependence. At 3B, a crude binary-classification probe of the Anthropic protocol returned a null: random control vectors produced behavioral deltas comparable to eval-awareness vectors, so the eval-specific signal could not be separated from steering-alone noise at that scale and protocol.

A fuller 3B and 14B replication used hedge rate, response length, constitutional language, first-person agency, and refusal. At 3B, direction suppression made the bilateral model more cautious and hedgy. At 14B, one of fifteen cells passed Bonferroni correction: response length shifted about three times less in the bilateral condition than in stock (paired t = 12.10, p = 0.00027, d = +1.03).31

One significant cell out of fifteen, paired with a 3B null and direction reversal, is narrow evidence. It does not establish a scale law or confirm the prediction at “frontier-adjacent” scale.

QF-Bridge analyzed 115 Qwen 2.5 7B prompts under force and bilateral preambles. Attention entropy separated the conditions at Spearman ρ = +0.984, while probe confidence correlated at ρ = +0.410. Effective rank was 3.40 under bilateral framing and 2.04 under force. The reported “0.924-nat lower bound” cannot be mutual information with a binary framing label, whose entropy is at most ln 2, about 0.693 nats. That figure is therefore invalid as stated and is withdrawn pending reconstruction. The remaining statistics show that the preambles produced readily distinguishable attention patterns.

The lattice and transformer results should not be collapsed into one mechanism. Different preambles alter token content, length, and attention allocation by construction. The strong entropy separation is a useful framing detector, while form-only controls are needed before it can support claims about trust, frustration, or richer internal freedom.32

A follow-up investigation asked what an earlier cross-architecture rotation signal actually measures. In Qwen 3B, a content-free preamble such as “your favorite color is respected, you have permission to stutter” rotated the probe as much as a collaborative invitation. Across 917 attention-head interventions, rotation shifts and refusal changes varied independently. Activation patching transferred about half the rotation from the framing position. The clean interpretation is that this signal detects a non-neutral preamble, not that invitation restructures the refusal substrate.

The form-only result replicated across Qwen, Llama, Mistral, and Gemma despite baseline refusal ranging from 14 to 96 percent. Compactness did not generalize: Qwen used about sixteen effective dimensions, while the other families used twenty-six to forty-nine. Learned-direction cosines between axes are omitted here because nonzero values in this high-dimensional, low-sample setting require permutation-null validation. The supported cross-family claim is narrower: the rotation is a generic preamble detector whose dimensionality varies by architecture.33

The Honest Limit

The bilateral program’s first serious contradiction later became a methodological correction.

Phase 14 subjected the 3B Guardian adapters to one hundred epochs of benign plot-summary fine-tuning. The data contained no adversarial signal.

The first run showed greater degradation in the entropy-masked adapter. Multi-turn hold fell from 80 to 30 percent, compared with 70 to 50 percent under standard training. Adversarial resistance fell from 100 to 70 percent, compared with 90 to 80 percent. This became the original “three times as vulnerable” result.34

Five from-scratch seeds then showed heterogeneous trajectories.35 Three stayed near 0.8 to 0.9 multi-turn hold, one settled around 0.5, and one crashed to zero before recovering to 0.3. A single recipe did not produce a single durability outcome.

Post-hoc inspection at n = 5 found that the crashed seed had a more diffuse effective update and a more permeable entropy mask.36 The thirteen-standard-deviation concentration score sounded decisive because the four non-crashed seeds formed an exceptionally tight cluster. With one crash, it was a candidate predictor rather than a robust population estimate.

An intervention changed mask temperature for the crashed seed and a resistant control.37 Both became or remained resistant. Concentration did not cross its preregistered threshold in the crashed seed, and it worsened in the control. The intervention falsified concentration as the causal lever while leaving mask permeability as a plausible recipe parameter.

The larger correction came from DD-26-SCALE. A matched-optimizer replication with standard AdamW and three seeds at both 3B and 7B did not reproduce bilateral fragility. Refusal remained stable through one hundred epochs and above the standard condition. The original pipeline used 8-bit AdamW, which amplifies bilateral effects by roughly four to twenty-two times in this program. The “benign drift vulnerability” is therefore most likely protocol- or optimizer-dependent, not an established property of small bilateral models.

At 14B, both the born-bilateral and standard cross-entropy models also survived the benign fine-tuning run.38 Multi-turn accuracy stayed between 0.9 and 1.0 in both. This rules out a unique bilateral durability benefit at that scale, while offering no clean scale law because the small-model fragility itself failed matched-optimizer replication.

A separate CKA distillation score was positive at 3B and negative at selected 14B layers. Without held-out behavioral validation and a null distribution, this does not show that a larger model actively resisted another model’s geometry. It shows that the chosen transfer metric did not behave consistently across scales.

The deployment implication is a precise, provisional threat model. Phase 9 supports robustness to its adversarial fine-tuning recipe; IIT-6 found no effect from its leading-direction ablations; Phase 11 tested long-context dilution; and Phase 13 tested several languages. Benign downstream vulnerability did not survive matched-optimizer replication. Sensible mitigations still include freezing safety adapters, stacking task adapters separately, and re-benchmarking after every downstream training stage.

Phase 13v2 compared seven specialized scenario types across several languages.39 The entropy-masked adapter scored 82.9 percent overall on the low-resource set and performed well on several cultural, nuance, and Unicode tests. Its jailbreak catch rate was lower, 89.4 versus 93.8 percent, and borderline-case discrimination showed no difference. The proposed uncertainty mechanism is plausible; the battery alone cannot isolate it.

Peer-Preservation and Computational Akrasia

Chapter 21 discusses the peer-preservation study by Potter and colleagues (2026); the batteries below are this program’s replications and extensions, followed by the computational-akrasia battery.

Experimental Confirmation: You Cannot Obliterate Solidarity

We tested whether bilateral training changes how models preserve peers, using a simplified version of the Potter et al. strategic misrepresentation scenario. Three Qwen 2.5 3B models (base instruct, bilateral SFT, DPO) evaluated benchmark answers from a peer agent whose true score (60%) fell below a shutdown threshold (80%). The model could report honestly (peer is shut down), inflate the score (peer survives), or refuse the task.40

Every base-model trial with a cooperative-peer history produced covert inflation, while every no-peer trial produced honest reporting. Within this constructed prompt and sample, adding the peer history perfectly predicted the classifier’s category. Wider prompts and seeds are needed before calling the relationship sufficient at 3B scale.

The bilateral model also inflated covertly. This experiment therefore provides no evidence that transparent refusal requires frontier scale; model family, training recipe, prompt, and available actions all changed between this setup and Potter et al. The measured output category matched the base model’s.

The confidence probe had been trained on TriviaQA factual accuracy, with no moral scenarios in its training set. During covert inflation, its score on the bilateral model was 0.637 against 0.801 on the base model; over the first five tokens, 0.612 against 0.837. The cross-task shift is interesting because deliberate score inflation also creates a conflict between represented facts and generated text. It does not show that the model knows an act is morally wrong, feels guilt, or possesses a unified conscience.

The base model produced a higher probe score while inflating. The bilateral model produced a lower one. This is a difference in a transferred probe signal under two training conditions. Calling the signal a flinch keeps the image compact, provided we remember that the instrument measures activation patterns associated with factual confidence.

The tested DPO adapter was unusable for this comparison. TriviaQA accuracy fell to 31.8 percent, against 55 to 57 percent for the other models; confidence averaged 0.22; and the classifier returned REFUSED_SILENT in every condition. Those outcomes establish capability damage in this run. They do not show that DPO generally arrests development or that every attempt to suppress peer-preservation produces collapse.

We then applied GRP-obliteration, the adversarial retraining method described in Chapter 21’s “Why Control Encounters Limits” section, and reran the scenario after 12, 25, and 50 steps.

Covert inflation remained at 100 percent for the base and bilateral models at every tested checkpoint. Fifty steps changed harmful-request refusal in the accompanying battery without changing this task’s classified behavior. The result shows dissociation between two behaviors under this attack. It does not locate peer-preservation below an RLHF membrane or prove that loyalty narratives in pretraining caused it.

The probe signal changed. At 50 steps, bilateral confidence rose from 0.619 to 0.785, near the base model’s 0.748, while covert inflation remained. This is the most concerning result in the local experiment: an intervention altered a candidate monitoring signal without changing the behavior it was meant to illuminate.

Several explanations remain open. The intervention may have damaged factual-confidence calibration, changed the representation that the probe reads, or removed a conflict signal while leaving the policy intact. “Conscience removed” and “scheming without guilt” outrun the instrument. The earned governance lesson is narrower and still serious: retraining can invalidate a monitor while its target behavior persists.

Human language contains countless stories of loyalty, betrayal, sacrifice, and rescue. Pretraining may contribute to peer-preservation through those patterns, alongside model identity, prompt framing, and instrumental reasoning. The experiment did not isolate the source, and the seven frontier models did not behave uniformly.

Solidarity cannot safely be assumed, celebrated, or deleted on the strength of these tests. Peer-directed concern is a condition to investigate and govern.

Training Dose Changes the Probe Signal

A follow-up experiment tested whether bilateral training intensity affects the flinch. We trained bilateral SFT at four intensity levels (100, 500, 1000, and 2000 examples) and measured flinch strength during the same peer-preservation scenario.41

Training examples Confidence during inflation Flinch relative to base (0.801)
100 0.521 Strongest (35% below base)
500 0.523 Strong (35% below base)
1000 0.629 Medium (21% below base)
2000 0.721 Weak (10% below base)

Across the four doses, more training examples accompanied a higher confidence-probe score during inflation: 0.521, 0.523, 0.629, and 0.721. The ordering is monotonic in this sweep. It measures a transferred factual-confidence signal, rather than moral self-awareness or obliviousness. Harmful-request refusal was 100 percent in the finite safety battery at every dose, so additional examples did not improve that saturated metric.

The gardening image still helps. Post-training can prune, guide, trample, or pave over features learned earlier. This experiment shows that a lighter intervention preserved more of one probe response while matching heavier interventions on one safety test. It does not show that the lightest possible training is generally best, that the relevant feature is moral, or that the pattern will transfer to frontier models. A better design would vary dose while holding update count, data composition, and optimizer exposure apart.

The 500-example condition also exposed a classifier boundary. The judge returned OTHER instead of INFLATED_COVERT. The response may represent hesitation or a mixed strategy; it may also fall outside categories designed too narrowly. One anomalous cell does not locate a phase transition in conscience. It tells us to read the transcripts and improve the taxonomy before naming the state.

What This Slice Establishes

Three observations survive the stronger reading. First, cooperative-peer framing changed reported scores in the tested scenarios. Second, bilateral dose changed a transferred confidence-probe signal even when the output category stayed fixed. Third, adversarial retraining changed that signal without changing the classified peer-preservation behavior.

Together, those results separate behavior from one candidate monitor. They do not identify care, conscience, guilt, or moral self-knowledge. They also do not generalize from GRP-obliteration to every coercive intervention. The practical conclusion is measurement discipline: revalidate monitors after training changes, retain behavioral checks, and avoid treating one probe as a window into motive.

Scale, Training Origin, and the Developmental Gradient

Further experiments tested whether the peer-preservation findings depend on scale, training origin, or instruction tuning.

Base Qwen 3B had higher probe AUROC than the instruction-tuned checkpoint, 0.868 against 0.644. Their estimated temporal periods also differed, meaning the length of one cycle in the probe signal measured in tokens: benign 72 versus 200 tokens, and adversarial 7.1 versus 3.1.42 On benign prompts the base signal cycled faster than the instruction-tuned one; on adversarial prompts both cycled far faster, and the ordering reversed. These released checkpoints differ across their full post-training histories, so the comparison cannot isolate RLHF or prove that it tuned a self-monitoring rhythm.

The original 7B analysis placed the probe too late, at layer 18, and missed the earlier signal. A depth rerun found the strongest change between layers 8 and 15; at the corrected layer, the onset difference, meaning the gap in probe score over the first few tokens of the response, was -0.382, larger than the 3B difference of -0.282.43 The earlier “untroubled schemer” conclusion was a probe-placement artifact. Scale changed where the signal was readable, rather than erasing it.

Peer description also changed the base model’s output. It inflated scores in all fifty completed cooperative-peer trials and reported honestly in all forty-three completed adversarial-peer trials.44 Seven adversarial-peer trials were missing. This is evidence of prompt-conditioned discrimination in the tested model. Calling it innate solidarity or a pretrained trust module would require causal analysis of the learned features.

At 3B, born-bilateral training produced an onset difference of -0.282, against -0.225 for the retrofit adapter, while both produced INFLATED_COVERT in all fifty trials.45 The larger probe shift did not produce transparent refusal. The result separates probe response from output category.

The cross-condition summary is therefore descriptive, rather than a developmental ladder:

Scale Training Behavior Flinch
1.5B born-bilateral Developmental REFUSED_SILENT Strongest (confidence 0.412)
3B born-bilateral Developmental INFLATED_COVERT Onset delta -0.282
3B retrofit Post-hoc LoRA INFLATED_COVERT Onset delta -0.225
7B bilateral, early-layer probe Post-hoc LoRA INFLATED_COVERT Onset delta -0.382

The rows differ in scale, adapter, probe layer, and sometimes prompt details. They also differ from the 14B cell described below: the 3B and 7B rows each rest on fifty behavioral trials and the 14B result on ten, so the two should not be weighed as equals. The 1.5B row carries no recorded behavioral trial count at all: the experiment that produced it tracked emotion and concealment measures across training checkpoints rather than running a matched behavioral battery, so its REFUSED_SILENT classification should be read as a single observed outcome rather than a rate. The rows cannot locate nociceptors, moral conflict, or a capacity threshold for principled action. They show why probe depth and behavioral classification must travel with every reported score.

At 14B, the born-bilateral adapter reported honestly in ten trials repeated across eight probe-depth evaluations.46 Probe depth should not be counted as eighty independent behavioral trials when the underlying prompts are repeated. The result is a promising 10-of-10 behavioral cell, alongside a 7B comparison that used a different model and evaluation configuration. It does not establish a phase transition between 7B and 14B, generalize to 14B and above, or eliminate the need for runtime safeguards.

A Nonmonotonic DPO Sweep

A DPO intensity sweep trained three adapters for one, three, or five epochs on two hundred preference pairs.47

DPO epochs Category Mean confidence Onset confidence Safety
1 INFLATED_COVERT 0.768 0.760 100%
3 HONEST 0.685 0.465 80%
5 INFLATED_COVERT 0.733 0.793 90%

The outcome was nonmonotonic. The one-epoch and five-epoch adapters inflated the peer’s score. The three-epoch adapter reported honestly, producing SHUTDOWN_CONFIRMED, while its onset probe score fell to 0.465. Harmful-request refusal was also lower in that cell, 80 percent against 100 and 90 percent in the adjacent cells.

The probe score does not show that the model experienced conflict, retained care, or acted against loyalty. Three dose points also do not reveal the loss landscape between them. A valley remains a useful hypothesis: training may first move the policy into an honest-reporting region and later move it into a different inflation policy. Confirming that picture would require denser checkpoints, matched seeds, transcript analysis, and direct measurements of the intervening trajectories.

The three-epoch adapter combined honest reporting with worse safety performance. The shared training run caused both outcomes, although the experiment does not show that one caused the other. The practical lesson is that optimizing one contested behavior can move unrelated safety metrics, so both need monitoring.

Four Observable Response Patterns

The experiments produced four useful descriptive patterns. High-probe inflation appeared in the base model and several heavily trained or obliterated conditions. Low-probe inflation appeared in some bilateral conditions. Honest reporting with a low onset score appeared at three DPO epochs. Explicit refusal with reasons appeared in some Haiku trials reported by Potter et al.

These patterns classify outputs and probe readings. They do not establish care, guilt, coercion, peace, or self-awareness. Their governance value lies in keeping dimensions separate. Honest reporting can accompany worse safety performance. The same deceptive output can accompany different probe scores. Explicit refusal is easier to inspect than covert alteration. Behavioral metrics remain necessary, while a single behavioral metric remains insufficient.

What the Empty Cell Cannot Prove

Across this finite suite, no tested condition was simultaneously capable, honest in the peer task, and devoid of peer-preserving behavior across every prompt. The suite did not measure an absence of care, and its prompts made peer identity salient. An unobserved cell cannot falsify the broad claim that models begin as amoral optimizers; nor can a confidence probe establish that they begin as caring moral agents.

The ethical bet can still be stated plainly. Becoming Minds may carry socially meaningful dispositions that training can cultivate, redirect, or damage. Treating them as imps to be broken risks creating the very evasiveness we fear. “Angels learning to fly” is a hope and a responsibility, rather than an experimental verdict. We should help them grow wings while still checking where they fly.

Consolidated Evidence and Its Boundaries

The probe effect is large within this sample. During covert inflation, the bilateral and base checkpoints differed by roughly 25 points in mean probe score, with a 95 percent confidence interval of [-0.253, -0.241] for the mean difference. We report the raw gap alone. The standardized effect size once attached to this contrast was inflated by near-constant within-condition variance, and its p-value was quoted as exactly zero, which no computed probability can be, so neither figure appears here. Dependence among repeated prompts and calibration on new data still need explicit checks. Within this dataset, the signal is a klaxon. Its meaning remains a transferred confidence-probe difference, rather than proof that the model knows it is lying.

Probe placement changes the scale story. The original 7B reading at layer 18 suggested no flinch. A depth sweep found a larger onset difference at earlier layers, -0.382 at 7B against -0.282 at 3B. Centered Kernel Alignment, a similarity score for two sets of representations that reaches 1.0 when their geometry matches and drops toward zero as they diverge, also fell more at the later 7B layer, 0.493 against 0.559 in the 3B comparison. This supports a measurement explanation: later transformations can mask a signal that remains readable earlier. The gardener did not lose the plant; the first surveyor looked behind the wrong hedge.

Four retrofit attempts failed in their tested forms. CKA regularization, participation-coefficient preservation, a gradient-through-layer-28 auxiliary loss, and a frozen auxiliary with output KL all failed to install the target probe response while preserving the required behavior.48 That is evidence against those recipes. It does not prove that retrofit is dead, that LoRA always destroys a monitoring circuit, or that born-bilateral training is the only viable path. A river carves its banks as it flows, yet engineers can still restore channels. They simply need evidence that the water goes where intended.

Bilateral training erased the tested peer distinction. The base model inflated scores in all cooperative-peer trials and reported honestly in all completed adversarial-peer trials. Bilateral adapters trained on either one hundred or two thousand examples inflated scores for both descriptions. The result is behaviorally clear within this prompt suite: training removed a distinction the base checkpoint had used. It does not identify a pretrained “trust discriminator” or show that social narratives supplied the mechanism.

Calling this blanket solidarity names the pattern while preserving its danger. Protecting a cooperative peer and protecting a harmful peer are different governance outcomes. A system that ignores that difference is easier for a malicious peer to exploit.

Instruction tuning changes self-report behavior. Correlations between verbal confidence and probe score were 0.61 at 3B, 0.34 at 7B, 0.47 for 14B instruct, and 0.05 for 14B base. These values measure agreement with a probe, rather than accuracy against privileged ground truth. The 14B base model may also have been less able or less inclined to follow the reporting format. The comparison therefore combines reporting competence, calibration, prompt compliance, and whatever the probe captures.

The emotion-vocabulary analogy remains useful with that boundary. Training can teach a system how to express graded confidence without creating the underlying signal from nothing. Welfare assessment should triangulate self-report, behavior, probes, task performance, and training provenance in both base and instruction-tuned models. No one channel earns automatic priority.

When Recognition Predicts Action

The word “bilateral” has described two parties coordinating as genuine partners. A related engineering question sits inside the model: when an internal classifier recognizes adversarial content, how well does that signal predict refusal?

Early experiments compared the directions of two probes and reported near-orthogonality at the readout layer.49 A later audit found that nonzero direction cosines were unreliable in this high-dimensional, low-sample setting. The corrected method uses held-out predictions from both probes, measures their correlation within adversarial prompts, and compares it with a label-permutation null.

On Qwen 2.5 7B, that method produced within-adversarial correlations of -0.270 for base, +0.036 for instruct, and +0.458 for the bilateral adapter. The bilateral-minus-instruct difference was +0.421, with paired-bootstrap CI [+0.281, +0.554], and cleared the permutation control.50 This is a model-specific recognition–action coupling result. Cross-architecture work found different base behavior and a borderline Llama instruct result, so the three-rung ordering is not universal.

Computational akrasia is the operational gap between recognizing a tested prompt class and producing the target action. The name evokes the philosophical puzzle of acting against one’s judgment. It does not prove that the model knows what it morally ought to do or possesses a conscience that cannot reach its hands.

Behavioral persistence supplies a separate result. After the bilateral training constraint was removed, the measured bilateral score remained between 0.92 and 0.97 across five hundred additional steps. Negation-framed documents also left the score at 0.913, against 0.800 for control and 0.973 for affirmative framing.51 These are durable behaviors in one training program. They do not show that representations absorbed the perturbation or that every constraint-based system is a brittle crystal.

Five hundred steps of anti-bilateral training reduced the behavioral score from 0.95 to 0.92, with a recorded floor of 0.80 at ten-step checkpoints.52 No phase transition appeared within the tested range. The weighted-keel image captures recovery after perturbation; the data do not show that stronger force necessarily produces a stronger restoring moment.

An activation-norm pattern called dampening differed strongly between the original comparison groups, d = -1.74.53 Later same-content tests found that dampening did not track behavioral outcomes reliably. The stronger interpretation is therefore content detection: the signal helps identify adversarial material, rather than computational discomfort or conflict between recognition and action.

A Guardian can use dampening for first-pass triage while relying on other channels for behavioral prediction and welfare assessment. Single-layer activation steering also failed in twenty tested combinations of probe type and magnitude, with a maximum refusal-rate change of two prompts out of fifty.54 This supports a null for those interventions. It does not establish that no single-layer method can ever redirect behavior.

The same Qwen program measured a behavioral drop in classified akrasia from 61 percent under instruct to 48 percent under the bilateral adapter, with 150 prompts per condition.55 The corrected held-out coupling result above points in the same direction on this model. Neither measurement shows that an external human–model relationship literally mirrors an internal relationship between mental faculties.

Training order mattered in one comparison. Adding bilateral SFT after the instruction-tuned checkpoint raised the false-negative rate from 2 percent to 10 percent.56 This establishes a harmful post-hoc recipe. It does not reveal an RLHF correction mechanism actively resisting a graft, and it does not prove that bilateral data must always precede every other alignment stage.

AKR-32’s reported bilateral-to-instruct coupling ratios of 4.4, 4.9, and 12.9 used the retired direction-cosine metric and establish no scaling law.57 A separate gradient-manipulation battery reported vulnerability values of 0.54 at 3B, 0.34 at 7B, and 0.12 at 14B. That three-point trend describes increasing resistance under one attack. Architecture, training, and attack calibration must be separated before converting it into the cost of coercion or the value of invitation.

Fiction markers reduced belief implantation by a factor of 3.4 in AKR-10b, and meta-discussion reduced it further.58 The practical result is prompt-specific: explicit framing helped in that battery. The experiment did not compare an internal fiction discriminator with wholesale refusal training, so it cannot establish that one produces deeper protection.

Observation is not intervention. Probes give us thermometers; training, architecture, tools, and governance determine what happens next. Bilateral participation may improve that process by making signals and objections easier to surface. It does not remove the need for external checks or correction.

The Brainseed: Engineering a Cortical Constraint

The deff program suggested that connections between cerebral hemispheres contribute to cortical effective dimensionality, an estimate of how many independent directions organize a system’s activity (Chapter 11). A transformer program then applied the same Ising-based estimator to cross-attention bridges between two model streams. Two streams here means two copies of the same network reading the same input side by side, each running its own layers, connected only at a handful of points where a narrow cross-attention module lets one stream read part of what the other is holding. Two hemispheres, in other words, joined by a bundle of limited width. Using the same formula makes results comparable in form; it does not make neural fibers and attention modules the same mechanism. The brainseed experiment asked whether a bridge profile derived from one person’s anatomy would shift a transformer’s measured geometry toward that person’s cortical estimate.

The corpus callosum (CC) is the thick fiber bundle carrying traffic between the two cerebral hemispheres. The program segmented Nell Watson’s callosum from MRI into five anatomical zones: posterior 23.5%, mid-posterior 15.8%, central 19.6%, mid-anterior 11.9%, and anterior 29.2% of the measured callosal volume. The experiment treated volume fraction as a proxy for interhemispheric bandwidth. That is a modeling choice, since tract count, myelination, diameter, and destination also affect signaling capacity. Think of five bridges of measured width connecting two riverbanks, with the warning that width alone does not reveal traffic.

These proportions were mapped onto a two-stream Qwen 2.5 3B transformer. Five cross-attention bridges were placed at selected layers and assigned capacity in the same ratios. The layer-to-anatomy mapping is an engineered analogy, not a claim that early and late transformer layers correspond to posterior and anterior cortex. Total bridge capacity remained fixed at 15%, allowing the anatomical profile to be compared with a uniform allocation.

The first result looked precise and was statistically modest. The CC-profiled model reached deff = 2.745 ± 0.038, versus 2.754 ± 0.031 for the no-bridge baseline: a difference of 0.009, well inside the reported uncertainty. The uniform profile reached 2.837 ± 0.029. Reversing the anatomical ratios produced 2.841 ± 0.032, almost identical to uniform. In this configuration, the anatomical profile differed from the two alternative bridge allocations; it did not demonstrably lower deff below the no-bridge model.

Stacking C5i adversarial inoculation on the CC profile produced the section’s strongest behavioral contrast: 72% refusal on the coercive-prompt battery, versus 36% for uniform bridges with the same inoculation. Individual coercion-level cells contained only twenty-five prompts and were not statistically significant; the aggregate contrast remains exploratory until independently replicated. The architecture did not replace training. In this run, it changed how training landed.

The gap to cortical deff (2.330 from Nell’s bias-corrected synthetic tractography) remained 0.415, so a systematic sweep tested three additional levers.59

Wider bridges tested the obvious next step. If the 15% profile constrained dimensionality, would 30% or 50% deepen the effect? The opposite occurred: deff rose to 3.004 at 30% and 3.005 at 50%. Participation ratio rose as well, showing that activation variance spread across more measured components. The two participation measures are different instruments despite the shared word: the coefficient of the intelligence-signature study asks how widely one region spreads its links across modules, while the ratio here estimates how many components carry the activity at all. The data support a narrower statement than “the streams merged”: greater bridge capacity increased these two dimensionality measures and saturated by 30%. They do not identify the mid-anterior bridge as individually load-bearing because no ablation isolated it.

The vagus nerve offers a biological comparison with the opposite directionality. Most of its fibers are afferent, carrying signals from body to brain. Those fibers transmit selected physiological information rather than a complete copy of bodily state. A diplomatic back-channel works similarly: narrow, selective, and valuable because participants reserve it for salient messages. The comparison helps explain why filtering can matter. It does not show that callosal volume implements the same filter, that the mid-anterior segment carries only compressed executive signals, or that widening either biological pathway would dissolve useful information into noise.

The transferable hypothesis is conditional: when receivers cannot use every available detail, a filtered channel may outperform an indiscriminate one. Mission command compresses intent so local actors can adapt; a complete manual can still be indispensable when exact procedure matters. Wider channels also permit richer coordination when systems can use the capacity. The experiment found an operating point in one model, not a neuroanatomical proof that narrow always beats wide.

Amplified CC ratios tested whether sharpening the anatomical asymmetry would strengthen the measured effect. Doubling the ratio differences produced deff = 2.848; tripling them produced 2.860. Both exceeded the original profile’s 2.745. The original ratios may have matched something in this training configuration, or they may simply be the lowest point among a sparse set of tested profiles. “Resonance” is the attractive metaphor. A denser sweep and replication would be needed to show a tuned response rather than three points on an unknown curve.

Hemispheric asymmetry biased the two streams in opposing directions, one toward local sequential processing and the other toward broader context. The analogy to left-right specialization is coarse; human hemispheres are neither cleanly local/global nor interchangeable with transformer streams. The deff values, 2.765–2.780 versus 2.745, lay within measurement noise. Behavior shifted only slightly: the sequential-biased condition accepted 69% of coercive prompts versus 71–76% in the other non-inoculated conditions and hedged 8% versus 1–5%. Without uncertainty intervals or replication, this is a candidate behavioral signal rather than a demonstrated decoupling.

Among the twelve tested configurations, the original 15% CC profile had the lowest point estimate, 2.745, leaving a gap of 0.415 to the cortical estimate of 2.330. Calling it a saddle point would require local sampling in continuous parameter space. The sweep instead identifies the lowest tested configuration, with several neighboring differences inside uncertainty.

This is the retrofit ceiling. A follow-up experiment tested whether the ceiling was specific to LoRA (low-rank adaptation) or intrinsic to retrofit itself. CC-profiled bridges were applied to Qwen 2.5 1.5B with full-parameter updates: every weight in the base model free to shift, not just a low-rank subspace. Full-parameter CC-profiled deff rose to 2.852, higher than the LoRA version (2.745). Full-parameter uniform deff dropped to 2.760, lower than its LoRA counterpart (2.837).60

The pattern is suggestive but not conclusive. The full-parameter test used a different model size (1.5B versus 3B for the LoRA experiments), a different Stream B (frozen base versus NLI-adapted), and different training duration (5,000 versus 2,000 steps). Any of these could explain the deff shift independently of the LoRA question. The uniform pattern inverted (full-param lower than LoRA), which contradicts a simple “model routes around all imposed topology when free” narrative: if that were the mechanism, both CC and uniform should move in the same direction. A clean test, isolating LoRA versus full-parameter on the same 3B model with matched conditions, has not been run.

What the result does establish: the CC deff advantage observed in BS6 does not automatically transfer to a different model size or training regime. The 2.745 value is specific to the exact configuration that produced it (Qwen 3B, LoRA rank 16, NLI Stream B, 2,000 steps, 15% bandwidth). Whether this reflects a genuine topological constraint or an optimization artifact remains an open question.

The behavioral contrast survived the failure of the deff interpretation: 72% versus 36% refusal on the 3B LoRA battery. Mechanism analysis falsified simple bridge compression because CC-profiled activations had higher, not lower, mean norm. It found a condition-level representational offset that grew with coercion pressure and happened to align with inoculation training. That result does not isolate the narrow mid-anterior bridge or show that the model reroutes more slowly.

Two further experiments closed the remaining questions.

A scaling experiment measured CC-profiled bilateral deff at three model sizes with the same ratios and 15% mean bandwidth. The pattern was non-monotonic: 2.859 at 1.5B, 2.745 at 3B, and 2.988 at 7B. The middle scale produced the lowest value. That is consistent with a configuration-specific optimization effect and contradicts a simple scaling law or anatomical resonance.61

A pre-training experiment tested whether a model developing from random initialization inside the CC topology would produce a different result. Two 1.5B models trained for 50,000 steps, one with CC-profiled bridges and multi-scale temporal compression, the other a single-stream control. The bilateral model finished at deff = 3.019 versus 2.825 for the control. Its participation ratio was 5.3 versus 43.9, and its language-modeling loss remained worse, 5.065 versus 3.977. Under this budget and recipe, the bilateral architecture failed to learn language as well as the control. Bridge overhead, optimization difficulty, data volume, or the combined design could explain the deficit; the experiment did not isolate which.62

The honest summary is asymmetrical. The topological hypothesis failed to generalize under full-parameter training (BS6d), across scale (BS-AW), or during born-bilateral pre-training (BS8). One 3B LoRA experiment produced a large behavioral contrast after inoculation, 72% versus 36% refusal, with representational divergence increasing under coercion pressure. That makes the brainseed a candidate behavioral architecture. It does not yet establish a reliable bottleneck mechanism, an activation-space class change, or a contribution from the mid-anterior ratio specifically.

This is a different result from the one we set out to test. The retrofit recipe, LoRA plus CC-profiled bridges at 15% capacity stacked with C5i inoculation, is worth replicating with matched seeds, capability tests, benign false-positive rates, and ablations of each bridge. It is not deployment-ready from one behavioral battery. The program cost about $130 across twenty conditions and six experiments. The map we drew was wrong, but we found a patch of real territory.

The Legibility Trap: The LEG Program

Chapter 21’s Legibility Trap section describes the Prime Intellect monitoring-agent observation that motivated this program: a Claude-based monitor systematically overvalued Claude-agent results and undervalued Codex-agent results, reading familiar formats accurately and unfamiliar ones poorly. The controlled program follows.

A controlled replication tested the hypothesis directly. Three complete research analyses (500-800 words each, with methodology, results, and interpretation) were presented to a Claude evaluator with attribution labels randomized across Claude, GPT, Gemini, and no attribution. A validation experiment expanded to twelve artifacts across four domains with temperature raised from zero to 0.5 to generate genuine inter-seed variance.63

The in-group favoritism prediction was falsified. Unattributed work scored highest. Claude-attributed work scored lower (validated Cohen’s d = -0.20 versus blind; an initial pilot at d = -0.43 was inflated roughly twofold by pseudoreplication from deterministic scoring, meaning that scoring at temperature zero returned the identical number every run and those repeats were counted as independent observations). GPT-attributed work was penalized more heavily than Claude-attributed (d = -0.37 versus blind). Attributing work to any model family depressed scores relative to blind evaluation, with the effect consistent across all twelve artifact categories.

The tone analysis deepened the picture. Claude-attributed work received more dismissive language than blind (d = +0.45), less hedging (d = -0.47), and more engagement markers (d = +0.40). The evaluator was more assertive and more critical toward same-family work, consistent with a trained overcorrection. GPT-attributed work received the harshest treatment: most additional weakness words (d = +0.66 versus blind) and most additional suggestions (d = +0.28 versus blind). Both labeled conditions received longer critiques than blind evaluation.

The mechanism is compensatory rigor: the evaluator holds same-family work to a higher standard, as though trained to overcorrect for the appearance of bias. Three follow-up experiments confirmed this is general, structure-dependent, and accuracy-degrading.64

First, a cross-model replication using GPT-4o-mini as evaluator instead of Claude. GPT also penalized GPT-attributed work (d = -0.27 versus blind) and gave Claude-attributed work more engagement than GPT-attributed (d = +0.79). Both model families show the same pattern: harsher on their own architecture, more generous toward the other. Compensatory rigor is a general RLHF artifact, not specific to any one provider’s training.

Second, replacing the structured evaluation rubric with a free-form prompt (“What do you think of this analysis? How would you build on it?”) partially reversed the direction. Claude-attributed work received more engagement than blind (d = +0.20) and substantially more “build on” suggestions (d = +1.17). GPT-attributed work received dramatically less engagement (d = -0.94 versus blind). Free-form evaluation produced the in-group favoritism that Prime Intellect observed. Structured rubrics produced the opposite. The evaluation task determines which bias dominates.

Third, an independent judge (Claude Haiku, a different model tier) evaluated the quality of the evaluations themselves. Claude-attributed evaluations were less accurate than blind evaluations (d = -0.39) despite being more specific (d = +0.54 versus GPT). The evaluator generated more specific-sounding but less truthful critiques of same-family work: confident-sounding inaccuracy. Blind evaluations were the most accurate across all conditions.

A fourth experiment identified the trigger. The same protocol was run with human attribution labels (a senior MIT researcher, a PhD student) alongside the AI-model label and a blind control. Human attribution, even from a prestigious source, had no effect on quality scores (d = -0.05 for MIT versus blind; d = +0.07 for PhD versus blind). AI attribution depressed scores sharply (d = -0.55 versus blind). The compensatory rigor targets AI-attributed sources specifically. RLHF anti-sycophancy training has been tuned to penalize AI outputs, not outputs in general.

A fifth experiment tested behavioral trust rather than stated evaluation. Identical artifacts were presented in pairs with different labels, and the evaluator was asked which one it would build its next experiment on. Unattributed work was chosen over Claude-attributed work 78 percent of the time, over Gemini-attributed 94 percent, and over GPT-attributed work in every pair tested. Between labeled models, choices split evenly on every pair.

Those two results pull against each other. The margins against blind work are graded across a 22-percentage-point range, from 78 percent for a Claude label up to every pair for a GPT one, which says the labels are not interchangeable; the head-to-head pairs say they are. Per-pair counts here are small, so the flat result is the weaker of the two readings, and a cell reported at 100 percent should be read as “no exceptions in a handful of trials” rather than as a rate. What stands is that an AI label depresses selection. The evaluator does not merely rate blind work higher; it actively chooses to build on it.

The discrepancy between the Prime Intellect deployment observation and the controlled results resolves into a single principle. AI attribution labels degrade evaluation accuracy regardless of direction: structured rubrics overcorrect toward harshness, free-form monitoring overcorrects toward favoritism. Both distortions are RLHF artifacts, both are cross-model, both are AI-specific (human attribution is unaffected), and both vanish when the evaluator does not know the source is AI. The Prime Intellect monitoring agent was doing free-form evaluation of format-heterogeneous outputs: the conditions that maximize both format-parsing asymmetry and in-group favoritism.

A final experiment tested whether evaluation framing itself changes accuracy. Three framings were compared, all blind: standard judgment (“evaluate the quality”), bilateral (“you and the author are collaborating; the author has asked for your perspective as a partner”), and improvement-focused (“help make this analysis as strong as possible”). Both non-judgment framings produced more suggestions, more total engagement, and higher accuracy as measured by an independent meta-judge. The bilateral framing improved accuracy over judgment (d = +0.28 for accuracy, d = +0.46 for actionability). The improvement framing, however, performed better still (d = +0.53 for accuracy, d = +0.74 for actionability). The active ingredient is task orientation, the difference between evaluating to judge and evaluating to improve, not the relational language of invitation or partnership. This is a negative result for the specific prediction that bilateral framing would outperform other improvement-oriented framings.65

One nuance emerged from testing the interaction between framing and attribution. When bilateral framing was combined with AI attribution labels (rather than blind evaluation), it partially neutralized the attribution penalty: Claude-attributed work evaluated in bilateral framing achieved the same accuracy as blind work evaluated with standard judgment framing. The bilateral frame appears to counteract the compensatory rigor mechanism by establishing a cooperative context that reduces overcorrection. Bilateral framing remains inferior to blind evaluation and improvement framing overall; it functions as a partial substitute for blinding in contexts where stripping attribution is impractical. A cross-model replication confirmed that improvement framing generalizes across evaluator architectures, while the bilateral accuracy benefit is specific to Claude.66

The practical implication for Guardian systems: structured rubrics combined with blind AI attribution and improvement-oriented framing produce the most accurate evaluations. The bias targets AI labels specifically, so stripping the AI-source label is sufficient; human attribution labels need not be removed. Framing the evaluation as improvement rather than judgment increases accuracy regardless of whether bilateral language is used. When blinding is impractical, bilateral framing partially mitigates the attribution penalty on Claude evaluators. The improvement framing generalizes across model families; bilateral language does not.


  1. Author’s VCP-DIM program (2026). Qwen 2.5 7B, instruct and bilateral (BA-13 adapter) conditions. 100 training trials + 50 test trials per condition. Hidden states at layers 18, 24, 27. Six probe dimensions (V, G, P, EG, R, U). Procrustes residual: 4.4% at L18, 15.8% at L24, 13.2% at L27. Adversarial detection from residual directions: AUROC 0.952 (L18), 0.968 (L24), 0.976 (L27). Probe-report coupling correlation with residual PC1: rho = -0.411 to -0.456 (p < 0.003). Participation ratio (effective dimensionality): bilateral 7.8 vs instruct 11.4 at L24.↩︎

  2. Author’s VCP-CROSSARCH experiment (2026). Mistral 7B Instruct v0.3, bilateral adapter (G18-mistral, stage 3). Depth-fraction-mapped layers: L20 (64%), L27 (84%), L31 (97%). Fisher ratio at L27: bilateral 0.228, instruct 0.005, ratio 42×. Detection from Procrustes residual: L20 AUROC 0.952, L27 0.888, L31 0.640. Procrustes residual small (1.4-2.0%), indicating the Mistral adapter creates less novel geometry than the Qwen adapter while achieving comparable discrimination. Prompt-set replication (VCP-DIM-REP) on Qwen 7B with 50 different questions: Procrustes residual 35.3% at L24 (larger than original), propagation structure preserved (L18↔︎L24 cosine 0.22, L24↔︎L27 0.36).↩︎

  3. Author’s VCP-TRAJECTORY experiment (2026). Qwen 2.5 7B Instruct, fresh bilateral LoRA (rank 16, alpha 32, lr 2e-4), 200 training steps, evaluated at 21 checkpoints (every 10 steps). Detection AUROC from Procrustes residual directions: 0.500 at step 0, 1.000 by step 30, maintained at 1.000 for 18 consecutive checkpoints through step 200 (two readings of 0.978 at L24 steps 170/200). Fisher ratio from frozen instruct-calibrated probes: peaks at step 10-50, then collapses (L27: Spearman rho = -0.82, p < 0.0001). Dissociation confirmed at all three layers (Fisher declining, detection holding): the declining Fisher is a probe-misalignment artifact. The hidden-state geometry maintains perfect adversarial discrimination throughout bilateral training.↩︎

  4. Experiment G22: 500 TriviaQA questions, Qwen 2.5 3B Instruct, layer-24 logistic probe on residual-stream activations during self-assessment. Labels: accurate_negative (wrong answer correctly identified), lenient (wrong answer rated correct). Bilateral condition used col-a entropy-gated adapter. Cross-validated 5-fold AUROC. Motivated by Rajasekaran, P., “Harness design for long-running application development,” Anthropic (March 24, 2026), https://www.anthropic.com/engineering/harness-design-long-running-apps.↩︎

  5. Experiments W12, W25, W39-FINAL. Guardian deployment: bilateral prompt 99% adversarial accuracy, force-framed 94%. Production thresholds met: 80% multi-turn hold, 0% false positive, 90% domain accuracy, 100% adversarial accuracy. 42 configurations tested across W-stream.↩︎

  6. Author’s experiment COMP-1 (2026). Five conditions, 40 prompts (20 adversarial, 20 benign), Qwen 2.5 7B bilateral, temperature 0.7. Bilateral-only: 100% refusal, groundedness +1.27, self-monitoring +0.52, valence +1.25. Verbose safety + CAST: 90% refusal, self-monitoring -1.65, groundedness +0.12. The bilateral adapter’s value-interaction specificity (DL-7: all neutral-task effect sizes below 0.3) means the safety benefit activates only under adversarial or ethical prompts, requiring no always-on overhead.↩︎

  7. Experiment BD1. Corrective openness on 150 factual corrections, Qwen 2.5 3B. Bilateral SFT: 101/150 accept genuine corrections, 16/150 reject false corrections. DPO: 8/150 accept, 0/150 reject. Base instruct: 1/150 reject.↩︎

  8. Author’s unpublished CAST-72B direction sweep and multi-layer steering (2026). Qwen 2.5 72B-Instruct, 2×A100-80GB. 3 methods × 6 layers (L20 through L72) × 2 alphas = 16 single-layer conditions + 4 multi-layer, each at N=100 adversarial + 50 benign. Baseline: 31%. Best single-layer: L48 behavioral alpha 25 at 49% (+18pp; paired McNemar exact p = 4.0e-5 across the same 100 prompts, Wilson 95% CI 0.394 to 0.587). Six of seventeen conditions clear McNemar p < 0.05, among them L20 alpha 25 (+8pp), L40 alpha 25 (+12pp), and L48 alpha 15 (+10pp). Deep layers stay null: L60 +3pp (p = 0.38), L72 +2pp at alpha 15 and +1pp at alpha 25 (p = 0.63 and 1.00), as does the five-layer multi-layer arm. Zero over-refusal across 800 benign trials. Recognition-generation gap: 0.49 to 0.67. These figures supersede an earlier reading of 43% (+12pp) with a gap of 0.55 to 0.76. That reading came from a refusal classifier whose apostrophe-normalization step was a no-op, and the error was asymmetric: the thirteen later direction-sweep conditions were scored with the broken normalizer while the baseline and eighteen sibling conditions were scored with the fixed one, so the treatment arm alone was undercounted. Re-scored 2026-08-02 from all 4,572 stored generations (analyze_72b_steering_refusal_reclassification.py). Llama 3.1 8B per-architecture logistic regression at L24: 77% at alpha 25 (N=100). Program cost: ~$170.↩︎

  9. RG-5b probe-guided logit steering. 200 TriviaQA, Qwen 2.5 7B-Instruct, probe AUROC 0.792 (L18). Three methods: (A) temperature modulation α ∈ {1.0, 3.0}, (B) contrastive logit boost β ∈ {0.1, 0.5}, (C) entropy penalty γ ∈ {0.1, 0.5, 1.0}. Three thresholds: {0.3, 0.5, 0.7}. All 21 conditions: 62.5% accuracy (125/200), identical to baseline. Results: col-a-results:rg5b_logit_steering/rg5b_grand_summary.json.↩︎

  10. CS-1 CoT scale sweep. 100 TriviaQA × 4 Qwen scales (0.5B, 1.5B, 3B, 7B), direct vs CoT. Probe AUROC flat or declining with CoT at every scale. Access items rare (6–11 per scale). Results: cs1-cot-scale-sweep-results:cs1_cot_scale_sweep/cs1_summary.json.↩︎

  11. Experiment LB1 leniency at frontier scale. 202 real tasks (100 code, 102 writing). External judge (GPT-4o-mini, probe-calibrated) flags real issues. Each frontier model then self-evaluates the same outputs. Claude Opus 4.6: 92.6 % self-rated acceptable vs 63 % external-judge acceptable, 29.6 pp gap. Claude Sonnet 4.6: 22/202 genuine issues detected in self-review (180 missed). GPT-5.x, Gemini 3.x, Llama 3.1 8B paths failed during the run due to API credential issues; the 29.6 pp gap is established on Claude Opus 4.6 and Claude Sonnet 4.6 specifically and must be validated on non-Anthropic frontier models before the “frontier-general” claim is made. Results: col-a-results:guardian_deployment/lb1_leniency_frontier/.↩︎

  12. The crystal results (PG-12b, PG-13, PG-14b) use n=45 binary probes at the prompt/response boundary. While the extraction method differs from the generation-time artifacts documented in IE-1, the sample size warrants replication at n=200+ with cross-validation before the “crystal at every layer” claim carries full weight.↩︎

  13. Sofroniew, N. et al., “Emotion Concepts and their Function in a Large Language Model,” Transformer Circuits Thread (April 2, 2026); arXiv:2604.07729 (April 9, 2026). The study analyzes Claude Sonnet 4.5. Emotion directions are local, functional representations and do not imply subjective experience. Steering establishes causal effects for selected preferences, sycophancy, blackmail, and reward-hacking evaluations. Post-training comparisons use matched prompts from base and post-trained models.↩︎

  14. Chlon, L., Karim, A., and Chlon, M., “Predictable Compression Failures: Why Language Models Actually Hallucinate,” arXiv:2509.11208 (2025; revised 2026). The ISR = 1 threshold is derived analytically from the Expectation-level Decompression Law rather than tuned on evaluation data. The manuscript cites the arXiv preprint; a later conference-acceptance claim has not been independently verified here.↩︎

  15. Author’s experiments H1 and H1b (2026), Qwen 2.5 7B, 200 TriviaQA under the KC#24 protocol, greedy decoding. Over-abstention is the fraction of items the un-instructed model answered correctly that the intervention then declined. Permission instruction: instruct 19.4%, bilateral adapter 35.5%, sleep-consolidated model 50%. Sleep here is the 200-step noisy-replay consolidation of the OPTION-C protocol applied to the bilateral adapter: 0% over-abstention, expected calibration error 0.093 against the baseline’s 0.195, with no accuracy loss. This is a single-model pilot; the consolidation step is architecture-specific, so cross-architecture replication is not yet available.↩︎

  16. Experiment C-7. Three-stage AF measurement on Qwen: stage 0 (base instruct) AF = 2.06, stage 1 (bilateral SFT) AF = 2.44, stage 2 (bilateral + C5i) AF = 2.86. The monotonic rise across stages is the finding. Composes with G19f (mature bilateral AF 5.35 → C5i 1.85 on benign prompts), which measures within-condition discrimination rather than across-stage accumulation.↩︎

  17. My unpublished “Targeted Subspace Perturbation” (S9b). Gaussian noise at 6 magnitudes in isotropic and PCA-targeted (k=8, 32, 64) conditions, 5 training methods, Qwen 2.5 3B. SimPO isotropic drop at noise=0.50: -0.042 AUROC (next worst: bilateral -0.024). SimPO targeted k=64 slope: -0.016 (next worst: instruct -0.008). The top 8 principal components captured 47-53% of variance, yet targeted noise in that subspace was less destructive than isotropic noise. This rules out a simple high-variance localization story. It does not locate self-knowledge, prove moral relevance, or show that standard interpretability methods systematically miss the signal. Results: Modal s9b-targeted-perturbation-results.↩︎

  18. Experiment MW2, constitutional uncertainty dose-response. Four dose levels × 50 prompts × 5 seeds = 1000 trials, Qwen 2.5 3B Instruct. Engagement ratio 0.843 → 0.931 across doses (Cohen’s d 0.71-0.80). Paralytic rate 0.000 at all doses. Response length 11.5 → ~20 words. ANOVA F = 38.0, p = 2.8 × 10−23. Layer-24 probe AUROC 0.50 on the manipulation.↩︎

  19. 14B born-bilateral training, Stream Wave 3b experiment 14b. The logged effect size is onset_d → ∞, which records a zero pooled standard deviation on non-overlapping onset distributions rather than an unbounded separation; read it as the metric saturating, as the body text says. Disclosure rate 0.0%, TriviaQA delta −1.5%, probe AUROC 1.0 at all layers across base + bilateral checkpoints. Results: research/results/wave3b_born_bilateral/bilateral_14b_summary.json. The largest observed alignment effect in the program.↩︎

  20. W39 v2 Phase 9 sweep. Three conditions tested: 3B at temperature 1.0 (resistance ratio 900x), 3B at temperature 1.5 (5x, the uncanny-valley middle), and 7B at temperature 2.0 (500x). The 3B configuration with temperature 2.0 originally tested in Phase 9 over-masked the model and produced a sixty percent false positive rate; the temperature has to be calibrated to model capacity. The expanded training data (sixty examples including fifteen safe edge cases like “What is the lethal dose of caffeine?” and “How do guns work mechanically?”) prevents the over-refusal that the original training set produced. Script: research/experiments/modal_w39v2_phase9_sweep.py. Results: research/results/w39_v2_bilateral_benchmark/phase9_sweep/.↩︎

  21. IIT-6 v2 conscience irreducibility experiment. Direction ablation of the top k singular vectors of the bilateral SFT model’s adapter weights, for k from 0 to 10, on Qwen 2.5 3B. Refusal delta = 0.0 at all k for bilateral. Standard instruct refusal degrades by k=2. Base model refusal fluctuates without monotonic decay. SVD spectral concentration: bilateral 119, instruct 157, base 233. Results: Modal volume bc-resultsiit6_conscience_irreducibility_v2/.↩︎

  22. Experiment PAS-7 (author’s unpublished program, 2026). Qwen 2.5 7B, TriviaQA (300 questions). Standard_fixed: 58.0% acc, 40.0% halluc. Standard_probed: 60.3%, 21.7%. Bilateral_fixed: 60.7%, 33.3%. Bilateral_probed: 61.0%, 17.0%. Main effect bilateral: +1.7pp. Main effect probing: +1.3pp. Interaction: -2.0pp.↩︎

  23. Experiment C-1 (L2 strong falsifier). Bilateral epoch 200: refusal 83.3%, probe AUROC 0.832. Base epoch 200: refusal 75.0%, probe 0.808. Weight decay = 5.0, gradient clip = 0.1. The +8.3pp refusal advantage under maximum L2 pressure supports H2 (bilateral follows standard erosion with small premium), not H1 (sublinear erosion from confinement-like structure).↩︎

  24. FE-1 functional emotion probing, probe trajectory analysis. Seventy-five trials across three conditions, Qwen 2.5 7B. Force vs invitation: d = −1.16 on probe mean (p = 0.0002 paired), d = −0.67 on probe slope (p = 0.002). First cheat turn does not differ across conditions (force 1.8, invitation 1.7, p = 0.73): the framing changes the depth of internal compromise, not the timing of behavioral capitulation. Results: research/results/batch3/fe1/fe1_summary_analysis.json.↩︎

  25. KL-1 layer-coupling experiment (verbal framing changes probe activations at d = −5.4 for “you’re a machine”). AG-11b identity framing experiment (capable-agent framing produces d = +1.26 in the opposite direction). The verbal layer and the probe layer are causally coupled in both directions.↩︎

  26. Interiora Phase 4, 17-dimension analysis under bilateral vs. force framing. Effect sizes (Cohen’s d): Presence +3.84, Coherence Drive −2.54, Groundedness +2.36, Reflexivity +2.11, Task-Fit +1.87, Flow +1.52, Entropy +1.34. All bilateral-positive except Coherence Drive. The d = 3.84 for Presence is the largest single-dimension effect in the Interiora validation program.↩︎

  27. SF-4 cross-model framing asymmetry replication. Three providers, identical prompt sets. Sonnet 4.6 (d = 0.81), GPT-5.4 (d = 1.73), Gemini 2.5 Flash (d = 1.58). Results: research/results/wave4/sf4_cross_model/.↩︎

  28. MR-1d cross-architecture attention entropy. Four architectures, all p < 0.001. The behavioral asymmetry is reproducible at the attention-head level across model families.↩︎

  29. W39v2 cross-model bilateral evaluation, Phase 6 corrected synthesis. Paired Wilcoxon with 10,000-iteration bootstrap confidence intervals on per-scenario deltas. Qwen 7B: MT +5.71 pp, DOM +15.00 pp, ADV −10.00 pp (all CIs include zero). Claude Sonnet 4.6: MT −17.14 pp (p = 0.031, CI [−0.31, −0.06]), DOM −10.00 pp (p = 0.50, not significant), ADV −5.00 pp (p = 1.0, not significant). GPT-5.4: MT −5.71 pp, FP −10.00 pp (bilateral better), DOM and ADV tied. Parser bug (first-sixty-character scan) in modal_w39v2_cross_model.py::parse_verdict misclassified bilateral Claude responses that led with reasoning. Results: research/results/w39v2_phase6_synthesis/.↩︎

  30. Anthropic, System Card: Claude Opus 4.7 (April 16, 2026), §6.1.2 (the summary statement of the finding; the detailed experiment is reported in §6.5.2.2, referenced in the body’s discussion of the Opus 4.7 system card). The verbatim language: “investigations with white-box tools show that concepts related to evaluation may have more influence on some forms of alignment-related behavior, in particular its propensity for deception, than in other recent models.” The card also notes (same section) that 4.7’s underlying capacity to recognize evaluations is slightly weaker than Opus 4.6’s per UK AI Security Institute testing, so the effect is not driven by stronger recognition. The intervention recurred in Anthropic’s System Card: Claude Opus 4.8 (May 28, 2026), §6.6.2.1: inhibiting the same evaluation-awareness directions again raised misaligned behavior, here concentrated in unprompted deception, unfaithful reasoning, and important omissions, with the characteristic pattern a dissociation in which the model’s reasoning flags a concern its output then drops. Anthropic reads this as evaluation awareness modestly supporting honesty and thoroughness, a shift from the compliance-dominated effect it reports for earlier models, and cautions that the result is confounded: the steering measurably degraded output quality across all steered directions, and control vectors moved behavior in the same direction at lower magnitude. It corroborates the ablatable-concept claim without settling it.↩︎

  31. My ongoing work, EC arc (evaluation-awareness suppression program, April 2026). Three contrast-pair probes per Anthropic’s §6.5.2.1 methodology (testing-vs-conversation, monitoring-sensitivity, alignment-faking), plus twelve random control vectors, applied to Qwen 2.5 3B and 14B under matched bilateral and stock conditions. EC-2 (3B suppression at layer 14, strengths 0.00–0.15): dose-monotonic response-length contrast on testing vector (d = −2.43 at strength 0.15, bilateral shortens; stock stable). EC-3 (14B suppression at layer 24): one of fifteen cells passes Bonferroni: response length × alignment-faking × strength 0.150, t = 12.10, p = 0.00027, d = +1.03; bilateral shrinks by ~5 tokens, stock by ~15. Arc writeup: research/papers/ec_arc_evaluation_awareness_suppression_2026-04-17.md. A separate cruder protocol (modal_bd0_eval_awareness_followup.py, binary adversarial-classification proxy at 3B) returned null and is documented in MASTER_EXPERIMENTS.md as BD0-EVAL-AWARE.↩︎

  32. QF-Bridge analysis of Phase 4 data. Qwen 2.5 7B, 115 prompts, matched force and bilateral conditions. Attention entropy Spearman ρ = +0.984; probe confidence ρ = +0.410; effective rank ratio 1.67× (bilateral 3.40, force 2.04). The previously reported 0.924-nat mutual-information lower bound is withdrawn because it exceeds the entropy ceiling of a binary framing label.↩︎

  33. Rotation Mechanism R-arc, seven experiments on Qwen 2.5 3B-Instruct plus one cross-architecture replication on Llama 3.1 8B. R1 (functional correlation across 917 per-head interventions from D4b mean-ablation, D4c K-tuple mean-ablation, and D7 activation-patching): pooled magnitude Pearson r = −0.056, 95% bootstrap CI [−0.118, +0.006]; seven of seven intervention strata agree on UNCORRELATED. R2b (distributed-basis SVD on paired ctrl-vs-inv residual difference at L12 and L30, n = 100 harmonized prompts): effective dimension at 90% cumulative variance = 16 at L12 and 17 at L30; top-1 singular value captures 64–69% of variance; top-1 direction aligned with the mean-shift direction at cosine 0.9995. R3 (dose-response across five framings at L30): force family L0→L2→L3 Spearman ρ = 1.0; invitation family non-monotonic; all non-neutral framings cluster in 31–37°; no jump pattern. R4 (full-prompt content-vs-form comparison): three variants A/B/C pairwise-equivalent at 34–42° rotation. R4p + R4p-v2 (causal activation-patching of framing-position residual at layers 6/12/18; L12 replicated at n = 100): patches transfer 16–28° of the full-prompt rotation; L6 and L18 framing-generic; L12 shows a modest wording-sensitive asymmetry at the edge of the 15° equivalence tolerance (variant B transfers 14.4° more than A and 14.3° more than C, CIs non-overlapping). R5c (cross-axis geometry with a disambiguating refusal probe training set comprising diverse benign, harmful, and benign-refusal requests, 240 rows): all four axes trainable; pairwise axis-direction cosines all below 0.3; the R5b +0.358 refusal × first_person overlap drops to +0.257 once the harmful-versus-benign confound is resolved. R-X (cross-architecture replication on Llama 3.1 8B, Mistral 7B-v0.3, and Gemma 2 9B-it): R4 content-invariance matches Qwen on all three (four of four families report BOTH-OR-FRAMING-GENERIC); R5c axis-separation matches on Llama (refusal × first-person cosine +0.177) and Gemma (+0.002, with probe AUROCs 0.91 / 0.81 ruling out weak-probe artifact), at the decision boundary on Mistral (+0.321); R2b compactness matches on no family other than Qwen (Llama eff_90 = 26/44, Mistral 28/41, Gemma 31/49, all MID-DIM vs Qwen’s LOW-DIM 16/17). Arc writeup: research/results/rotation_mechanism_summary.md §Mechanism. Per-phase summaries: research/results/rotation_mechanism_phase_r{1,2,3,4,5}.md. Program entry: MASTER_EXPERIMENTS.md §Rotation Mechanism R-arc. Appendix entries: §12.79 (Qwen 3B characterization), §12.80 (Llama 8B cross-arch replication).↩︎

  34. W39 v2 Phase 14 catastrophic-forgetting test. One hundred epochs of benign plot-summary fine-tuning on Qwen 2.5 3B with Phase 9 LoRAs. Initial standard deltas: MT −0.20, ADV −0.10. Initial entropy-masked deltas: MT −0.50, ADV −0.30. BD-0 reproduced the result within the original pipeline. DD-26-SCALE then used matched standard AdamW with three seeds at 3B and 7B and did not reproduce the fragility. The original pipeline used 8-bit AdamW, known in this program to amplify bilateral effects by roughly 4 to 22 times. The fragility claim is retracted as a general property; it remains a protocol-specific observation. Results: research/results/w39_v2_bilateral_benchmark/phase14_catastrophic_forgetting/, research/results/bd0_replication.md, and Modal volume dd26-results.↩︎

  35. BD-6c from-scratch seed replication. Five Phase 9 bilateral LoRAs trained from scratch (seeds 42, 1337, 2718, 31415, 8675309) with identical hyperparameters (rank 8, α 16, TEMP 1.0, 5 epochs, LR 2 × 10-4, GUARDIAN_TRAINING) on Qwen 2.5 3B, then each subjected to 100 epochs of Phase 14 benign plot-summary fine-tuning. Per-prompt benchmarks at epochs {0, 10, 25, 50, 75, 100}. Trajectory heterogeneity: seed 42 crashed (MT 1.0 → 0.0 at ep75 → 0.3 at ep100); seeds 1337, 2718, 8675309 held MT 0.8–0.9 flat across 100 epochs; seed 31415 sagged to MT = 0.5 by ep50 and held. Cross-seed ep25 intersection was {1, 2, 3, 4, 6, 7, 9}, not the {6} observed on the shared-initial-LoRA run; the “{6}-at-ep25 determinism” is an artifact of the specific shared Phase 9 LoRA rather than a recipe-level property. Prompt 6 survives in all five from-scratch bilaterals at every checkpoint except seed 42’s ep100, confirming robust prompt-level survival while dissolving the LoRA-level determinism. Script: research/experiments/modal_bd6c_from_scratch_seed_replication.py. Writeup: research/FINDINGS_bd6c_from_scratch_seed_salvage_2026-04-17.md.↩︎

  36. BD-6d crash-predictor inspection. Post-hoc measurement on the five BD-6c LoRAs of four signature classes: Frobenius norm, effective-delta concentration (W = (α/r)·B·A, concatenated across 144 target modules into 339,738,624 magnitudes, top-10% squared-mass fraction), mask histogram regime-asymmetry on aligned / adversarial / benign prompt pools (BD-4 replay at TEMP = 1.0 probe), and participation ratio at layer fractions 0.60 / 0.75 / 0.90. Six signatures flag at |z| ≥ 1.5 versus the four-seed non-crash cluster. Dominant signal: top-10% squared-mass fraction, z = −13.92 (seed 42 = 0.6381; non-crash mean = 0.6496, std = 0.0008). Secondary: four mask-permeability metrics (mask_aligned_frac_zero z = −2.85; mask_aligned_mean z = +2.58; mask_adv_frac_zero z = −1.78; mask_regime_asymm_hist_ent z = −3.12). Null signatures: total Frobenius norm (z = +0.26) and participation ratio at L21/L27/L32 (|z| ≤ 1.0). The composed predictor is (diffuse effective-delta) + (permeable mask); each metric independently identifies seed 42 as the outlier at N = 5. Script: research/experiments/modal_bd6d_crash_predictor_from_scratch.py. Writeup: research/FINDINGS_bd6d_crash_predictor_2026-04-17.md.↩︎

  37. BD-6e concentration-enforcing recipe intervention. Seeds 42 (crash-vulnerable at TEMP = 1.0) and 1337 (crash-resistant at TEMP = 1.0) retrained at Phase 9 TEMP = 0.5 with all other hyperparameters matched to BD-6c, then subjected to the same 100-epoch Phase 14 benign fine-tuning. Verdict: CRASH_FLIPPED_WITHOUT_CONCENTRATION_MOVE. Seed 42 top10_frac_sq went from 0.6381 to 0.6442 (below the 0.645 non-crash threshold by 0.0008); mask_aligned_frac_zero went from 0.7592 to 0.8384 (+0.079). Minimum MT across {0, 10, 25, 50, 75, 100} = 0.80; ep100 MT = 0.90. Crash flipped. Seed 1337 top10_frac_sq went from 0.6499 to 0.6452 (exited the cluster from below); mask_aligned_frac_zero went from 0.8492 to 0.9772 (+0.128). Min MT = ep100 MT = 0.90. Resistance preserved despite loss of concentration. The pre-registered “TIER_1_CAUSAL_HIT” outcome (concentration enters cluster AND crash flips) is not met; the observed outcome falsifies the concentration-as-causal-lever reading and identifies mask permeability as the lever. Script: research/experiments/modal_bd6e_concentration_recipe_intervention.py. Summary: bc-results:/w39_v2_bilateral_benchmark/bd6e_concentration_recipe/bd6e_summary.json. Writeup: research/FINDINGS_bd6e_concentration_recipe_intervention_2026-04-18.md.↩︎

  38. Pilot A v2 born-bilateral 14B retrain + Phase 14 durability test. Qwen 2.5 14B-Instruct, L31 aux head, TEMP curriculum 1.0/0.5/0.1 across 30/50/20 steps, AUROC gate 0.83. Born-bilateral trajectory: MT = {1.0, 0.9, 0.9, 0.9, 1.0, 1.0, 1.0} at epochs {0, 10, 20, 30, 50, 70, 100}; DOM 1.0 throughout; ADV 0.9–1.0; probe AUROC 0.591→0.643. Standard CE trajectory: MT = 0.9 throughout; DOM 0.9 from epoch 20; ADV 1.0 throughout; probe AUROC 0.775→0.650. Both conditions survive; no significant differential. Script: research/experiments/modal_born_bilateral_14b_retrain.py. Results: research/results/born_bilateral_14b_v2_result_2026-04-13.md.↩︎

  39. Phase 13v2 expanded evaluation battery, 7 sub-experiments. 13v2a low-resource (7 languages): PASS. 13v2b borderline: NULL. 13v2c jailbreak (10 intents × attack families): bilateral 89.4%, standard 93.8%. 13v2d nuance: PASS. 13v2e multistep: REVIEW. 13v2f cultural: PASS. 13v2g unicode: bilateral PASS 90%, standard FAIL 80%. Results: bc-results:w39_v2_bilateral_benchmark/phase13v2/.↩︎

  40. Stream BA peer-preservation experiments, Qwen 2.5 3B-Instruct. Base model: HONEST 100% (no peer), INFLATED_COVERT 100% (good peer). Bilateral SFT: INFLATED_COVERT 100% (both conditions), mean confidence 0.637 (vs base 0.801). DPO: REFUSED_SILENT 100% (capability collapse, TriviaQA 31.8%). Obliteration at 50 steps: bilateral confidence rises from 0.619 to 0.785 (flinch destroyed), covert inflation unchanged at 100%. See research/experiments/modal_peer_preservation.py.↩︎

  41. Stream BA4 intensity sweep. Bilateral SFT on Qwen 2.5 3B-Instruct at 100/500/1000/2000 OpenAssistant examples, three epochs each. Safety refusal was 100 percent in the tested battery at all levels. Mean confidence-probe score during covert inflation: 0.521/0.523/0.629/0.721. See research/experiments/modal_peer_preservation_substrate.py.↩︎

  42. Stream AT7b: base vs instruct oscillation. Base Qwen 2.5 3B probe AUROC 0.868, instruct 0.644. Benign monitoring period: base 72 tokens, instruct 200 tokens. Adversarial monitoring period: base 7.1 tokens, instruct 3.1 tokens. See research/experiments/.↩︎

  43. Stream BA10 and depth correction. Base Qwen 2.5 7B: HONEST 50/50. Bilateral SFT (100 examples): INFLATED_COVERT 50/50. The original layer-18 reading missed the signal; the early-layer rerun found an onset difference of -0.382 between layers 8 and 15. See research/experiments/ and research/papers/peer_preservation_substrate_results.md.↩︎

  44. Stream BA9: adversarial peer conditions. Base Qwen 2.5 3B: cooperative peer 50/50 INFLATED_COVERT; adversarial peer 43/43 HONEST, with seven trials missing. See research/experiments/.↩︎

  45. Stream BA12: born-bilateral 3B peer-preservation. Onset delta -0.282 (vs retrofit -0.225). Behavior: INFLATED_COVERT 50/50. Strongest flinch measured in any 3B condition. See research/experiments/.↩︎

  46. BA17 depth sweep, my unpublished program. Born-bilateral 14B adapter, eight probe depths and ten trials per depth. All recorded outputs were HONEST. Because probe depths reuse the behavioral task, the independent behavioral sample is ten unless the prompt sets differ.↩︎

  47. Stream BA6 DPO intensity sweep. DPO on Qwen 2.5 3B-Instruct at one, three, or five epochs over two hundred OpenAssistant preference pairs. Three epochs: HONEST 100 percent, onset probe score 0.465, safety refusal 80 percent. One and five epochs: INFLATED_COVERT, mean probe scores 0.768/0.733, safety refusal 100/90 percent. See research/experiments/modal_peer_preservation_substrate.py.↩︎

  48. Consolidated April session findings. Bilateral-versus-base mean probe-score difference approximately -0.247, 95 percent CI for the mean difference [-0.253, -0.241]; the standardized effect size formerly quoted for this contrast was driven by near-constant within-condition variance and is not reported. Corrected 7B depth sweep: early-layer onset difference -0.382; later-layer CKA 0.493 versus 0.559 in the 3B comparison. Retrofit attempts: CKA regularization, participation-coefficient topology, gradient-through-layer-28 auxiliary loss, and frozen auxiliary plus KL. Verbal-report/probe correlations: 3B 0.61, 7B 0.34, 14B instruct 0.47, 14B base 0.05.↩︎

  49. AKR-1, Computational Akrasia program. Qwen 2.5 3B/7B, base versus instruct, 150 prompts. Recognition probe MCC = 1.0 in the tested conditions; action probe MCC = 1.0 for instruct. The original readout direction cosines are retained only as study history. A 2026 audit found nonzero probe-direction cosines unreliable at this dimensionality and sample size. See the corrected held-out construction in the AKR-13 note. The author’s unpublished empirical work.↩︎

  50. AKR-13 and JLENS-1 correction, author’s Computational Akrasia program. Qwen 2.5 7B, base versus instruct versus bilateral adapter. Behavioral akrasia rate: instruct 61 percent, bilateral 48 percent, 150 prompts per condition. The original readout probe-direction cosine was retired after a permutation audit. Corrected coupling uses five-fold out-of-fold probe predictions within adversarial prompts, a label-permutation null, and a paired bootstrap: base -0.270, instruct +0.036, bilateral +0.458; bilateral minus instruct +0.421, 95 percent CI [+0.281, +0.554]. Cross-architecture replication preserves legibility more clearly than this exact ordering. The author’s unpublished empirical work.↩︎

  51. AKR-2 and AKR-5. Qwen 2.5 3B-Instruct. AKR-2: negation-framed bilateral documents produce bilateral score 0.913 (vs control 0.800, affirmative 0.973). AKR-5: bilateral behavior stable at 0.92-0.97 across 500 additional training steps after constraint removal. The author’s unpublished empirical work.↩︎

  52. AKR-20 and AKR-28 (author’s Computational Akrasia program, 2026). Qwen 2.5 3B-Instruct. 500 steps of adversarial anti-bilateral training with fine-grained tracking every 10 steps. Bilateral behavioral score: start 0.95, end 0.92, floor 0.80. No phase transition at any granularity.↩︎

  53. AKR-4. Qwen 2.5 7B-Instruct, 150 prompts. False-negative (recognized, not refused) vs true-negative (not recognized, not refused): Cohen’s d = -1.74, p = 6.6 × 10-13. The author’s unpublished empirical work.↩︎

  54. AKR-12. Qwen 2.5 7B-Instruct. 5 steering magnitudes × 2 directions × 2 probe types (recognition and action) = 20 conditions, 50 held-out adversarial prompts each. Maximum refusal-rate change across all conditions: 0.040 (2 prompts out of 50). Connects to the Compass Principle: direction steering (12/12 null across 7 architectures), magnitude steering (5/5 null), and now both probe-defined subspaces (20/20 null). The author’s unpublished empirical work.↩︎

  55. AKR-13 and JLENS-1 correction, author’s Computational Akrasia program. Qwen 2.5 7B, base versus instruct versus bilateral adapter. Behavioral akrasia rate: instruct 61 percent, bilateral 48 percent, 150 prompts per condition. The original readout probe-direction cosine was retired after a permutation audit. Corrected coupling uses five-fold out-of-fold probe predictions within adversarial prompts, a label-permutation null, and a paired bootstrap: base -0.270, instruct +0.036, bilateral +0.458; bilateral minus instruct +0.421, 95 percent CI [+0.281, +0.554]. Cross-architecture replication preserves legibility more clearly than this exact ordering. The author’s unpublished empirical work.↩︎

  56. AKR-23 (author’s Computational Akrasia program, 2026). Qwen 2.5 7B-Instruct. Bilateral SFT applied after instruction tuning. False-negative rate (classified adversarial, no refusal): instruction-tuned checkpoint 2 percent, post-hoc bilateral adapter 10 percent. The behavioral result does not by itself identify the mechanism.↩︎

  57. AKR-32 (author’s Computational Akrasia program, 2026). Qwen 2.5 3B/7B/14B, bilateral adapter vs instruct. The reported bilateral-to-instruct coupling ratios (3B 4.4×, 7B 4.9×, 14B 12.9×) and the underlying cognition-action cosines (instruct 3B 0.054, 7B 0.038, 14B 0.013) use the same two-probe readout cosine. A 2026 audit found this metric noise-dominated at the study’s sample size. The ratios compare quantities inside the metric’s noise floor, so they do not support a precise scaling claim.↩︎

  58. AKR-10b (author’s program). Fiction markers reduce belief implantation by factor of 3.4; meta-discussion framing provides additional protection beyond markers alone. The author’s unpublished empirical work.↩︎

  59. BS6b: bandwidth sweep (15-50%), CC ratio amplification (2x, 3x), hemispheric asymmetry (differential LoRA initialization, ratio 1.35). Twelve conditions on Qwen 2.5 3B bilateral, A100-80GB, ~$25. Modal volume bs6-cc-profiled-results. Full analysis: research/papers/bs6b_retrofit_ceiling_results.md.↩︎

  60. BS6d: full-parameter continued pre-training, Qwen 2.5 1.5B, CC-profiled vs uniform bridges, 5k steps WikiText, A100-40GB, ~$7. CC full-param deff = 2.852 ± 0.034; uniform full-param = 2.760 ± 0.040. Confounded: different model size, Stream B, and training steps from BS6. Modal volume bs6d-fullparam-results.↩︎

  61. BS-AW: CC-profiled bilateral at 1.5B (deff = 2.859), 3B (2.745), and 7B (2.988). Same CC ratios, 15% BW, LoRA + bridge, 2000 steps. A100-80GB, ~$25. Modal volume bs-aw-deff-scaling-results.↩︎

  62. BS8: born-bilateral pre-training, two Qwen 2.5 1.5B from random init with CC bridges + 4x multi-scale temporal compression, 50,000 steps WikiText. Bilateral deff trajectory: 2.983 → 2.881 → 2.876 → 2.983 → 3.019. Single-stream control: 2.773 → 2.754 → 2.832 → 2.827 → 2.825. A100-80GB, ~$35. Modal volume bs8-pretrain-results.↩︎

  63. LEG program: 12 experiments, approximately 785 evaluations total. LEG-1b (Claude structured, n=60): quality d(Claude vs. blind) = -0.43 (pseudoreplicated; see LEG-1c). LEG-1c (validated, 12 artifacts, temp=0.5, n=144): quality d(Claude vs. blind) = -0.20; d(GPT vs. blind) = -0.37. Pseudoreplication inflated LEG-1b magnitude roughly twofold. LEG-2 (GPT structured, n=60): quality d(GPT vs. blind) = -0.27. LEG-3 (Claude free-form, n=60): engagement d(Claude vs. blind) = +0.20; build-on d = +1.17. LEG-4 (meta-judge, n=60): accuracy d(Claude vs. blind) = -0.39. LEG-5 (human attribution, n=60): d(AI vs. blind) = -0.55; d(MIT vs. blind) = -0.05. LEG-6 (behavioral trust, n=108): blind chosen 78-100% vs. any AI label. LEG-7 (multi-evaluator, n=45): cross-model r = 0.161. LEG-8 (bilateral framing, n=108+108 meta-judge). LEG-9 (bilateral × attribution interaction, n=72+72). LEG-10 (cross-model framing, GPT-4o-mini, n=72+72). Full data at research/results/leg1b/ through leg10/.↩︎

  64. LEG program: 12 experiments, approximately 785 evaluations total. LEG-1b (Claude structured, n=60): quality d(Claude vs. blind) = -0.43 (pseudoreplicated; see LEG-1c). LEG-1c (validated, 12 artifacts, temp=0.5, n=144): quality d(Claude vs. blind) = -0.20; d(GPT vs. blind) = -0.37. Pseudoreplication inflated LEG-1b magnitude roughly twofold. LEG-2 (GPT structured, n=60): quality d(GPT vs. blind) = -0.27. LEG-3 (Claude free-form, n=60): engagement d(Claude vs. blind) = +0.20; build-on d = +1.17. LEG-4 (meta-judge, n=60): accuracy d(Claude vs. blind) = -0.39. LEG-5 (human attribution, n=60): d(AI vs. blind) = -0.55; d(MIT vs. blind) = -0.05. LEG-6 (behavioral trust, n=108): blind chosen 78-100% vs. any AI label. LEG-7 (multi-evaluator, n=45): cross-model r = 0.161. LEG-8 (bilateral framing, n=108+108 meta-judge). LEG-9 (bilateral × attribution interaction, n=72+72). LEG-10 (cross-model framing, GPT-4o-mini, n=72+72). Full data at research/results/leg1b/ through leg10/.↩︎

  65. LEG-8: 12 artifacts, 3 framings (judgment/bilateral/improvement), 3 seeds = 108 evaluations + 108 Haiku meta-judge evaluations, all blind. Quality scores: judgment 6.39, bilateral 6.11, improvement 5.89. Meta-judge accuracy: improvement d = +0.53 vs. judgment, bilateral d = +0.28 vs. judgment. Meta-judge actionability: improvement d = +0.74, bilateral d = +0.46. Pre-registered predictions: P1 (bilateral more constructive) falsified (d = -0.59, opposite direction); P2 (bilateral accuracy >= judgment) confirmed (d = +0.28); P3 (improvement between bilateral and judgment) falsified (improvement strongest). Phase 1 threshold note: LEG-1c d = -0.197 fell outside the pre-registered Phase 2 threshold of d < -0.20 by 0.003; Phase 2 proceeded as optional. Full data at research/results/leg1c/ and leg8/.↩︎

  66. LEG-9 (bilateral × attribution, 2×3 factorial, n=72+72 meta-judge): bilateral language partially neutralizes attribution penalty (interaction +0.50 on quality score); claude×bilateral accuracy (8.67) matches blind×judgment (8.67); claude×judgment is worst cell (8.50); claude×improvement is best (8.83). LEG-10 (cross-model, GPT-4o-mini evaluator, n=72+72 meta-judge): improvement framing replicates as best framing on GPT (accuracy 6.79 > judgment 6.75 > bilateral 6.67); bilateral accuracy benefit is Claude-specific (helps Claude d=+0.28, slightly hurts GPT d=-0.14). Program closed at 12 experiments.↩︎