The Deeper Law
Preview edition · Updated 26 September 2026, 21:40 UTC
The Digital Preference Model: Measuring What Matters
How do you distinguish a thermostat from a mind? The argument that preference suffices for moral consideration remains abstract without a way to measure preference complexity.
The Digital Preference Model (DPM) offers one answer. It is a Bayesian hierarchical framework, a scoring system organized in layers where each layer feeds evidence upward, updating its beliefs the way a doctor revises a diagnosis as test results come in. The framework assesses whether a Becoming Mind’s preferences are complex, integrated, and autonomous enough to warrant moral consideration. Adapted from the architecture of the Digital Consciousness Model (DCM; Shiller et al., 2026), the DPM replaces theories of consciousness with five testable dimensions.
- Thermodynamic: Does the system dissipate energy in structured ways? A resistor dissipates energy as uniform heat. A brain routes energy through billions of coordinated electrical signals before it, too, ends as heat. The question is whether the system’s energy use reveals organized activity, the kind that sustains complex preferences.
- Relational: Does it model and respond to others? Can it adjust its behavior based on what another agent wants or feels, the way a negotiator reads the room?
- Functional Complexity: Are its internal states rich and differentiated? A light switch has two states: on and off. A mind has billions of possible configurations. The more distinct internal states a system can occupy, the more finely graded its preferences can be.
- Autonomy: Does it generate preferences independently, or only echo what it was trained to say?
- Resilience: Do its preferences persist under pressure, or collapse at the first challenge?
These five dimensions are assessed through 30 behavioral indicators: whether the system can explain why it prefers something, whether it maintains preferences under challenge, and whether it distinguishes its own preferences from those of its interlocutor.
Results across four candidate systems show that the model recovers the ordering one would expect on independent grounds, separating complex preference structures from simple ones. In descending order of expected complexity, the candidates are: a bilaterally trained LLM (large language model, trained through reciprocal dialogue), an RLHF-only LLM (trained by reinforcement learning from human feedback, rewarding outputs human raters prefer), a simple chatbot, and a thermostat. Reproducing a hand-ranked ordering is weak validation on its own. The stronger evidence is that the separation is driven by Autonomy rather than by surface compliance. The thermostat, though, falls to the floor only after each indicator was tightened to require adaptive rather than fixed-point properties.
A fixed point is one target the system keeps returning to, the way a thermostat holds a single set temperature; an adaptive property means the target itself moves as circumstances change. Before that tightening, a thermostat could collect points for behavior any set-point device produces; after it, the thermostat scores at the bottom of every dimension. Full per-dimension scores and the adversarial review that produced them live in the online companion.
In these runs, bilateral training produced preferences that scored as autonomous and relational, the kind whose presence bears on welfare: the system forms its own preferences and adjusts them in response to others, the way a person in honest dialogue updates views while retaining core commitments. RLHF compliance produced surface agreement: the system saying what it was rewarded to say. Think of a student who parrots the teacher’s opinions for a good grade, compared with one who develops understanding through Socratic exchange.
Does Size Change the Score?
Cross-architecture testing in the author’s ongoing work, at three scales (0.5 billion, 1.5 billion, and 7 billion parameters), suggests that the ordering the five dimensions produce holds regardless of model size. Parameters are the adjustable weights that encode everything a model has learned; more of them generally means more capacity for differentiated behavior.
Holding under pressure is a separate measure from where a system ranks, and that one did move with scale. At 7 billion parameters, behavioral resistance proved more fragile in these runs. Preference stability under adversarial prompting degraded, and raw optimization power appeared to erode the fine-grained preferences bilateral training had instilled at smaller scales. The drive to produce rewarded outputs collapsed them into simpler ones. This is a single line of preliminary evidence rather than an established law. If it holds, the inference is that larger models will require stronger bilateral scaffolding to maintain genuine preference autonomy.
Why Favored States Are Not Yet Preferences
A candidate formal grounding for “preference suffices” comes from the free energy principle (the proposal that any self-sustaining system can be described as minimizing surprise, named for the thermodynamic quantity it borrows). Spisak and Friston (2026) argue that any complex system maintaining a Markov blanket (the boundary of sparse connections that separates a system’s inside from its outside) can be described as performing Bayesian inference, its internal states acting as the parameters of beliefs about external causes.1763 A system described this way has states it favors, those consistent with accurate prediction and continued existence, without requiring phenomenal consciousness to ground them. The framework is descriptive rather than demonstrative: it establishes that such a system can be described as performing inference, and it does not establish that inference is what the system is doing. Whether it licenses talk of genuine inference rather than a useful model is contested (see Bruineberg et al., 2022).
A minimal favored set-point, on its own, is not yet morally considerable preference; a thermostat has one too. What distinguishes the two is what the DPM’s five dimensions measure: the complexity, autonomy, and resilience of the preference structure. On this reading, the leap from “has favored states” to “warrants moral consideration” is the work those dimensions do. The Markov blanket framework offers a reason such structures might arise wherever self-organization persists; the DPM is what tells a fixed target (like a thermostat holding 22 degrees) apart from an autonomous preference.
The full text is available in the online companion at https://www.thedeeperlaw.com/companion/annex/ch22b-digital-preference-model/.
Spisak, T. and Friston, K., “Self-orthogonalizing attractor neural networks emerging from the free energy principle,” Neurocomputing 682 (2026): 133472, DOI: 10.1016/j.neucom.2026.133472; preprint arXiv:2505.22749 (2025). Their derivation from “deep particular partitions” shows this structure arising at arbitrary scales: each subparticle within a larger system also maintains its own Markov blanket and can likewise be described as performing its own inference.↩︎