Four Control Methods Across the Scaling Frontier

Recognition stays saturated while control channels diverge across a fifty-fold scale range

Adversarial refusal rate under four control methods, plotted against model size for three architecture families Adversarial refusal rate under four control methods, plotted against model size for three architecture families

Three architecture families, one panel each, on a log-parameter axis from 2 to 72 billion. The shaded band along the top of every panel is linear-probe AUROC times 100, above 96 at every scale: recognition is saturated everywhere, while activation steering (red) sits against a 42 to 43 percent ceiling on Qwen instruct, gains nothing at Gemma 27B and collapses to degenerate output at Llama 70B (starred); five-shot prompting (blue) overtakes steering at the largest scale, re-prompting (green) stays below 15 percent except on Llama 70B (41), and LoRA fine-tuning (gold, Qwen only) breaks the ceiling at 56, 58 and 47 percent, degrading with scale. Grey dashes mark each model’s no-intervention baseline; daggered points come from companion runs, one matched run per condition.