Kramers-Wannier Duality: When Misalignment Is Order
Specialist Annex
Kramers-Wannier Duality: When Misalignment Is Order
Technical section for Chapter 17: The symmetry between trust and distrust
The Duality Discovery
A universality class is physics’ way of grouping systems that, whatever their ingredients, approach their phase transitions in the same mathematical way: with the same critical exponents. AI alignment in the agent model matches the 2D Ising universality class in symmetry structure and transition topology (the exponents β = 0.125, γ = 1.750 of Onsager’s 1944 solution), though the measured order-parameter exponent is β ≈ 0.22, which sits between the 2D Ising value (0.125) and mean-field (0.5) and suggests crossover. [Inference] Annex 60 (Variational Convergence) discusses this discrepancy in detail. Taking the membership as approximate, does alignment inherit the other properties of that class?
The 2D Ising model possesses a deep symmetry called Kramers-Wannier duality, discovered in 1941, three years before Onsager’s exact solution. This duality maps the ordered phase to the disordered phase and vice versa, the way a photographic negative swaps dark for light while carrying exactly the same picture. The mathematics is exact:
where K is the coupling strength and K* is its dual. At the critical point, K = K*: the system equals its own dual.
The question for alignment: if trust is the ordered phase, what is its dual?
The Dual Order Parameter
In the original Ising picture: - Spins () live on lattice sites - Magnetization measures collective alignment - Ordered phase: high , spins mostly parallel
In the dual picture: - Domain walls live between sites with opposite spins - Disorder parameter μ measures domain wall density - The disordered phase of the original = the ordered phase of the dual
For AI alignment, this maps to:
| Original | Dual |
|---|---|
| Agent alignment states | Behavioral boundaries |
| Collective alignment (φ) | Inconsistency density (μ_dual) |
| Trust Attractor | Coordinated-Defection Attractor |
The dual order parameter is the density of behavioral boundaries: places where an agent’s behavior switches between aligned and misaligned states.
What We Found
We ran numerical tests for Kramers-Wannier duality in our agent model:
| Test | Expected | Observed | Status |
|---|---|---|---|
| φ and μ_dual cross at criticality | Yes | No crossing | ✗ |
| Variance ratio = 1 at critical point | 1.0 | ~4.4 | ✗ |
| Susceptibility peaks coincide | Yes | ✗ | |
| Domain wall clusters peak at transition | Yes | Peak observed | ✓ |
(Here
denotes the location of the critical point on the disorder-parameter
axis, distinct from
,
the dual order parameter itself.
is the separation between the two susceptibility peaks. These results
were produced by the verification script
kramers_wannier_verification.py; full numerics in
KRAMERS_WANNIER_RESULTS.md.)
Verdict: Approximate duality, not exact.
The elevated variance ratio (~4.4×) likely reflects continuous agent states and all-to-all topology, both of which soften duality relative to the discrete lattice Ising model. Universality class and exact duality are different things: our continuous, all-to-all agent model shares critical exponents with 2D Ising (that is what universality class means) but lacks the specific symmetry of the discrete square-lattice model.
The conceptual mapping remains valid. The quantitative equivalence does not.
The Philosophical Payoff
The key insight survives the numerical weakness. The conceptual duality rests only on the existence of a dual order parameter, not on exact self-dual symmetry. The one test that passed (domain wall clusters peak at the transition) confirms that such a parameter exists and is physically meaningful; the three tests that failed concern the exact reciprocal relationship, which the continuous all-to-all model was never expected to satisfy. The conceptual mapping needs only the former.
Misalignment is order in a different basis.
Consider what the dual picture reveals:
In the aligned phase (Trust Attractor): - High φ: strong collective alignment - Low μ_dual: sparse behavioral boundaries, randomly located - The system is ordered in the original basis
In the misaligned phase: - Low φ: weak collective alignment - High μ_dual: dense behavioral boundaries, forming connected networks - The system is disordered in the original basis, yet ordered in the dual basis
A misaligned AI system is structured: failures form patterns, inconsistencies trace networks. Understanding misalignment means finding this dual order.
Two Descriptions, One Physics
The deepest implication:
Physics provides two equivalent descriptions of the same system. In one description, alignment is “order” and misalignment is “disorder.” In the other, the labels swap.
Neither description is more fundamental. The underlying reality is one thing; our descriptions of it are choices.
This is why ethics cannot be reduced to physics.
The thermodynamic stability of the Trust Attractor is physics. The preference for trust over distrust is ethics. Both are real. Neither reduces to the other.
The dual picture makes this vivid: the universe provides both the Trust Attractor and the Coordinated-Defection Attractor. Both are locally stable equilibria, each “ordered” in its respective basis. Physics is not silent between them: the Trust Attractor occupies the deeper basin, as the rest of this book argues, and in that sense physics does rank the two. What physics cannot supply is the ought. A deeper basin is greater stability, and greater stability is not the same as goodness. Even the thermodynamically favored attractor does not, by itself, license a value.
(A note on what the duality does and does not establish: Kramers-Wannier duality maps the ordered phase of one system to the disordered phase of its dual, with an inversion of temperature regime. It does not assert that both phases are stable at the same temperature, nor that the two attractors are equally stable. The existence of the Coordinated-Defection Attractor is a game-theoretic fact, evidenced below by cartels and mutual deterrence, not a consequence the duality proves.)
We choose. The choosing is real.
Practical Implications
Dual Training Objectives
RLHF (reinforcement learning from human feedback, the standard alignment-training method) maximizes the original order parameter φ: it trains for aligned outputs directly.
Consistency training would minimize the dual order parameter μ_dual: it would penalize behavioral variance across inputs, regardless of alignment direction.
These are different objectives with different robustness profiles: - RLHF: robust to noise, potentially fragile to systematic attacks - Consistency: robust to systematic attacks, may drift in alignment direction
The duality suggests both approaches have merit.
Jailbreaks as Domain Wall Injection
In the dual picture, a jailbreak does not make the model misaligned everywhere. It creates behavioral boundaries: domain walls that disrupt the low-μ_dual state of the ordered phase.
Practical implication: monitoring domain wall density (behavioral inconsistency) may detect jailbreaks before alignment metrics collapse.
The Dual Trust Attractor
If the Trust Attractor exists, so does its dual: a stable equilibrium of coordinated defection.
Examples in human systems: - Cartels (coordinated non-competition) - MAD (mutual assured destruction) - Stable adversarial equilibria in game theory
These are ordered, not chaotic: ordered in the dual basis. Physics tells us which basin runs deeper; it does not tell us which to call good. That choice is ours.
Summary
The Kramers-Wannier investigation revealed:
- Exact duality does not hold for our continuous all-to-all model
- Conceptual duality remains valid: there is a dual order parameter (domain wall density)
- Misalignment has structure: it is alternative order, organized in the dual basis
- Physics ranks the attractors; ethics supplies the ought: physics can say which basin runs deeper (trust does), yet greater stability is not goodness. The gap between “more stable” and “should” is why ethics is irreducible to physics
Annex 60 (Variational Convergence) formalizes the fixed-point structure behind coordination networks through a renormalization group analysis. That analysis derives the bifurcation ratio shared across constructal, biological, and urban scaling, and classifies coordination operators (coercion among them) as relevant or irrelevant under coarse-graining. It treats the 2D Ising exponents discussed here as input rather than as something it independently derives.
The mathematics of duality does not dictate our values; it clarifies what values are. Physics already tells us which basin runs deeper. Values answer a different question: which attractor to call “good.”
Physics gives us the landscape. Ethics chooses the destination.
“What if the difference between trust and distrust were merely a choice of which dual to inhabit?”
The answer: there is an objective difference. Trust occupies the deeper basin. The duality is real, and so is the choice of which attractor to call good. Physics ranks; we still must choose.