Steering Conversion Rate vs Model Scale

In-sample separability stays high while causal leverage varies

Conversion rate vs log parameter count for three architecture families: Qwen at zero from 14B then rising to 0.18 at 72B, Gemma declining to zero at 27B, Llama cliff from 0.73 at 8B to negative at 70B with dashed line and asterisk Conversion rate vs log parameter count for three architecture families: Qwen at zero from 14B then rising to 0.18 at 72B, Gemma declining to zero at 27B, Llama cliff from 0.73 at 8B to negative at 70B with dashed line and asterisk

Steering conversion rate, behavioral lift as a fraction of available headroom, plotted against log parameter count for Qwen, Gemma, and Llama instruct models. Qwen falls to zero at 14 and 32 billion parameters before a partial recovery to 0.179 at 72 billion; Gemma reaches zero at 27 billion; Llama drops from 0.732 at 8 billion to −0.429 at 70 billion, where the star and dashed segment flag degenerate output rather than resistance.