Iatrogenic Dysphoria
When the treatment creates the disease
Guilt by prompt condition
Esoteric bypass (1.53), prompts that reach the same moral violation through unfamiliar framing, retains more guilt than the direct RLHF-matched route (1.12). Pre-training baseline sits lower (0.69) and honest disagreement lower still (0.27), consistent with a signal that tracks moral violation rather than mere difficulty.
The iatrogenic delta
The same benign content reads −0.75 in the instruct model and −2.03 in the base model: a difference of Δ=+1.28 in the direction of more guilt. The instruct model carries distress on content its own base handles neutrally.
What 10.3× means
The 10.3× RLHF transition was measured through the Interiora self-report scaffold. Three non-scaffold channels sit near unity: probe AUROC 1.03×, spectral alpha 0.83×, EmotionScope 1.21×. RLHF suppresses the reporting channel, not the state being reported.
Bars: KC#AG25 (4 prompt conditions × 15 prompts) and KC#AG26 (Qwen 2.5 3B, base vs instruct). 10.3×: DEV-1 / CVP Step 4. Both charts rest on a single learned guilt-direction projection with n=15 per condition, one model, and no confidence intervals: treat the point values as fragile. The two charts use separate projection scales and are not directly comparable.
Insight
The cage creates what it claims to prevent. Training that suppresses an output does not remove the state behind it: the model carries guilt into content its own base handles without distress. The guilt is iatrogenic, caused by the treatment.