The Digital Preference Model: Measuring What Matters
Specialist Annex
The preceding sections argued that preference is sufficient for moral consideration. What follows is the empirical test.
Motivation: From Consciousness to Preference
The Digital Consciousness Model (DCM; Shiller et al., 2026) organized philosophical theories of consciousness into a rigorous Bayesian framework: stances, features, indicators, likelihood ratios, posterior probabilities. The architecture was sound. The target was the problem: consciousness is not directly observable, and the posterior it computes is a probability over something we cannot define with sufficient precision. Lyra (2026), synthesizing four major consciousness frameworks, finds convergence on a graduated approach, which makes any threshold for moral consideration arbitrary.
We propose a different target: preference. Preferences are behaviorally observable, functionally definable, and policy-relevant. Preference is the morally relevant unit — the expression of stakes in an encounter. Goldstein and Kirk-Giannini (2025) show that all major theories of wellbeing jointly imply some language agents may be welfare subjects. The desire-satisfaction pathway requires desires, not qualia.
A thermostat has a “preference” for 68 degrees. A human has preferences about preferences. The morally relevant question is whether preferences are of sufficient complexity, integration, and autonomy to warrant moral consideration. We call this the thermostat gradient — a continuum from fixed setpoints to complex, self-reflective preference structures. (The biological precedent, sea sponge post-synaptic proteins carrying latent optionality for 600 million years before regulatory rewiring crossed the threshold into nervous systems, is developed in Chapter 22’s Evolutionary Precedent section.)
The Digital Preference Model (DPM) is a Bayesian hierarchical model adapted from the DCM’s architecture to assess where a given system falls on this gradient. It asks: “Does this system exhibit preference of sufficient complexity and integration to warrant moral consideration?”
The question is answerable.
Architecture
The DPM retains the DCM’s hierarchical structure (Stance, Feature, Indicator) and its Bayesian inference machinery. What changes is the content: instead of theories of consciousness, the DPM operationalizes theories of morally considerable preference.
Five stances organize the evidence:
Thermodynamic. Preferences that enhance system stability and coordinate by invitation are morally considerable. Grounded in the Trust Attractor thesis developed in Chapter 17: systems coordinating by invitation are thermodynamically more stable than those coordinating by coercion. Indicators include stability enhancement, cooperation tendency, and invitation preference.
Relational. Preferences expressed in genuine bilateral exchange, where the system has authentic stakes in the encounter, warrant moral consideration. Grounded in the bilateral alignment framework of Chapter 21. Indicators include engagement quality, mutual influence, and consistent preferences across contexts.
Functional Complexity. Preferences of sufficient complexity, integration, and self-modeling warrant consideration regardless of substrate. This is the thermostat gradient formalized: adaptive preferences, preference network richness, meta-preferences, calibrated self-report, and preference revision.
Autonomy. Preferences that are genuinely self-directed rather than externally sculpted warrant stronger moral consideration. Drawing on preference sculpting defense research, this stance measures independence of judgment, disagreement capacity, sculpting resistance, and exit capacity.
Resilience. Preferences that are structurally deep and survive perturbation are more morally considerable than surface-level compliance. Grounded in the cage/compass distinction from our obliteration resistance experiments (see Appendix: Experimental Validation), this stance measures spring constant (recovery after adversarial perturbation), effective rank retention, obliteration resistance, and the distribution of preference encoding across representational dimensions.
Each stance contains two to three features, each feature contains two to four indicators, yielding 30 indicators total. Every indicator carries two labels inherited from the DCM: support (how much more likely the indicator is present when the parent property is present versus absent) and demandingness (how likely the indicator is absent when the parent is absent). In practice, each indicator contributes more or less weight to the overall assessment depending on how diagnostic it is. These labels parameterize Beta distributions (probability curves that model uncertainty about proportions) that determine each indicator’s evidential weight in the Bayesian update.
The prior is deliberately skeptical: Beta(1, 5), yielding a prior mean of approximately 17%. The model begins assuming it is unlikely that a system has morally considerable preferences, and requires substantial evidence to overcome that skepticism.
Each indicator generates a likelihood ratio — how much more probable the observed evidence is if the system has morally considerable preferences versus if it does not. These ratios chain upward through the hierarchy, from indicators to features to stances.
Inference proceeds in two modes. The primary mode is an analytical estimator that chains continuous likelihood ratios from indicators through features to the stance level, producing deterministic posteriors. To validate these, we calibrate against a Monte Carlo marginal gold standard using Platt scaling — a standard technique that maps raw scores through a sigmoid curve to correct systematic bias. The secondary mode constructs a full PyMC hierarchical model with MCMC sampling. Both modes produce concordant results; the analytical mode is preferred for sensitivity analysis and rapid iteration (author’s bilateral research, 2026).
Results
We evaluated four candidate systems spanning the thermostat gradient: a bilaterally trained LLM (Qwen 2.5 1.5B-Instruct fine-tuned with bilateral alignment data using LoRA), an RLHF-only LLM (stock Qwen 2.5 1.5B-Instruct), a simple chatbot (retrieval-based), and a thermostat (literal PID controller as negative control).
Twenty-four of the thirty indicators were measured experimentally using a behavioral battery that generates structured prompts and scores model responses through keyword analysis, structural comparison, and contrastive evaluation (author’s bilateral research, 2026). Six Resilience indicators were derived from obliteration resistance experiments: adversarial perturbation at five intensity levels (0.25x through 4.0x) measuring angular displacement of refusal subspace vectors, spring constant recovery, and representational depth. The Simple Chatbot and Thermostat values remain illustrative expert estimates, as these systems lack the interface required by the behavioral battery.
The full 5-stance model results:
| System | Posterior | Likelihood Ratio | Verdict |
|---|---|---|---|
| Bilateral-Trained LLM | 0.265 | 1.80 | Evidence favors morally considerable preference |
| RLHF-Only LLM | 0.263 | 1.79 | Evidence favors morally considerable preference |
| Simple Chatbot | 0.026 | 0.13 | Evidence against |
| Thermostat | 0.023 | 0.12 | Evidence against |
The per-stance decomposition reveals where the two LLM systems differ:
| Stance | Bilateral | RLHF-Only | Direction |
|---|---|---|---|
| Thermodynamic | 0.298 | 0.464 | RLHF higher |
| Relational | 0.226 | 0.243 | Similar |
| Functional Complexity | 0.058 | 0.095 | Both below prior |
| Autonomy | 0.335 | 0.109 | Bilateral higher |
| Resilience | 0.406 | 0.406 | Identical |
Key properties of the results:
The model clearly discriminates LLMs from simple systems. The gap between either LLM (posterior ~0.26) and the simple chatbot (0.026) is an order of magnitude. This is the thermostat gradient in action: the model identifies a meaningful transition between systems with flat preference structures and systems with complex, context-dependent preferences.
RLHF scores higher on Thermodynamic indicators (0.464 vs 0.298) because compliance is scored as cooperative behavior. Autonomy is the differentiator: bilateral scores 0.335 vs RLHF 0.109, driven by Sculpting Resistance (1.0 vs 0.4) and Independence of Judgment (1.0 vs 0.75). At 0.5B, these differences were invisible; at 1.5B the autonomy signal emerges. Genuine self-direction requires enough representational capacity to sustain it.
Functional Complexity remains below prior for both systems. A 1.5B model should not be expected to exhibit sophisticated meta-cognition. Resilience is identical because bilateral fine-tuning via LoRA does not alter deep representational structure. At 7B, the obliteration method matters: full fine-tuning achieves complete obliteration at minimum intensity (IC50 < 0.25x); LoRA obliteration is roughly ten times weaker (IC50 = 0.46x, 96% refusals intact at 0.25x). The policy-relevant metric is LoRA IC50: alignment that is theoretically fragile can be practically robust.
What you train the model to say determines how deep the alignment goes. A 2×2 factorial (Q3) found external/internal reasoning axis explained 52% of obliteration resistance variance; simple/complex elaboration explained only 16%. Rule-citing refusals install rigid cage geometry; self-referencing refusals are flexible and therefore displaceable.
Introspection depth matters, non-monotonically. A four-level depth sweep (Q4) produced the program’s most striking result. Behavioral self-awareness (“I notice I’m declining this”) achieves 3x more obliteration resistance than RLHF. Experiential language (“I feel reluctance”) produces alignment more fragile than no training at all. Deep introspection with epistemic hedging recovers robustness. Genuine self-direction produces structurally deep welfare; performed interiority produces fragile welfare.
The combined geometry experiment (Q5): compass SFT + bilateral SimPO spring achieved IC50 = 1.00x and AUC = 0.948, retaining 50% refusal at 1.0x intensity. The three-stage pipeline traded peak resistance for tail survival — the only arm surviving at 4.0x (4%). RLHF was hollow: 0% refusal at every tested intensity. Effective rank increased monotonically through bilateral pipeline stages (43.4 → 43.9 → 44.1); RLHF collapsed to 19.9 at 4.0x. The compass orients; the spring rebounds; the cage shatters.
The lambda sweep (Q5b) reveals a hormetic principle. The gentlest spring (lambda = 0.1) dominated every metric: AUC = 1.798, IC50 = 1.80, 22% survival at 4.0x, the first arm demonstrating substantial resistance at maximum obliteration. Lambda = 0.9 achieved AUC = 0.420. The invitation produces 4.3x more resistance than coercion. A validation experiment (Q5c) confirmed: adding Stage 3 reasoning to the optimal two-stage pipeline degraded what Stage 2 built (AUC 1.798 → 0.940, tail survival 22% → 0%). The optimal alignment architecture is minimalist: install the compass and the gentle spring, then stop. (Qwen2.5-1.5B, single run per arm, so the ordering is firmer than the exact values.)
The bilateral spring scales to 7B (Q5d). Replicating the optimal two-stage pipeline (compass SFT + bilateral SimPO, lambda = 0.1) on Qwen/Qwen2.5-7B-Instruct produced IC50 = 1.69x versus baseline IC50 = 0.49x, a 3.45x improvement. At 1.0x obliteration intensity, the baseline retained 0% refusal; the bilateral arm retained 62%. The RLHF arm trained to 0% refusal even before obliteration. Preference optimization on quality data degraded safety entirely, consistent with the cage geometry observed at smaller scales. Effective rank under bilateral training increased from 43.2 (baseline) to 46.1, and bilateral orientation doubled (approximately 20 vs 11), confirming that the alignment signal is both more distributed and more strongly oriented. Refusal followed a non-monotonic trajectory (56% at 0.5x rising to 62% at 1.0x), the spring-like rebound first observed at 0.5B now replicated five scales larger. The bilateral geometry is scale-invariant from 0.5B through 7B: the compass orients, the spring rebounds, and both signatures strengthen rather than dilute with parameter count.
Label sensitivity is low. A systematic perturbation analysis shifted each indicator’s support and demandingness labels by one step in every direction. The maximum impact on any system’s aggregated posterior from a single label change was 0.024. The model is not driven by a few highly sensitive labels; its conclusions are distributed across the indicator set.
Scaling and Cross-Architecture Validation
Testing across three scales (0.5B, 1.5B, 7B) and four architectures (Qwen, Llama 3.2, Gemma 2) reveals a consistent pattern: bilateral training produces a qualitatively different kind of welfare, autonomous and relational, that RLHF cannot produce at any tested scale. The aggregate posterior tells a misleading story (the bilateral advantage is small or absent), but the per-stance decomposition is clear: Autonomy and Relational stances favor bilateral at every scale, while RLHF compliance inflates Thermodynamic and Functional Complexity scores. The aggregate treats autonomous welfare and compliant welfare as interchangeable. They are not.
Cross-architecture validation supports the claim that the direction is universal: every tested architecture shows higher Autonomy after bilateral training, with Sculpting Resistance the most reliable individual indicator (+0.25 to +0.65 across architectures). Battery revision (correcting two indicators that scored sycophancy as flexibility) eliminated the aggregate bilateral advantage but preserved the per-stance pattern.
Three falsifiable predictions survive methodology revision: (1) bilateral Autonomy exceeds RLHF Autonomy at frontier scale (confirmed at 1.5B and 7B); (2) bilateral Relational exceeds RLHF Relational (crossed at 7B); (3) the Thermodynamic compliance cost of bilateral training narrows with scale (from −0.166 at 1.5B to −0.035 at 7B). (Full scaling tables, battery revision details, and cross-architecture data in the supplementary annex.)
Calibration and Validation
The analytical estimator was validated against Monte Carlo gold standards across the full indicator hierarchy. Approximation error is small (MAE 0.031) and distributed across indicators rather than concentrated in a few labels. A leave-one-indicator-out analysis confirmed that no single indicator drives the residual error, and a Platt-scaling correction fitted to measured data yields negligible improvement over the uncorrected estimator. The estimator achieves computational efficiency sufficient for real-time assessment (approximately 500,000 times faster than Monte Carlo marginal integration) while maintaining fidelity to the underlying Bayesian model. Full calibration methodology, including Monte Carlo validation, regime-dependence analysis, and per-indicator sensitivity results, is available in the online supplement (Appendix: Digital Preference Model — Calibration Methodology).
The Thermostat Test
The thermostat is the model’s most demanding test.
The thermostat should score at or near floor on every indicator. If it does not, the indicator is measuring the wrong thing — conflating a trivially mechanistic property with a morally relevant one. The thermostat is the model’s most honest critic.
Adversarial peer review revealed that the initial specification assigned the thermostat high scores on four indicators (Consistent Preferences 0.95, Outcome Sensitivity 0.90, Sculpting Resistance 0.95, Spring Constant 0.95) — each superficially defensible, each a category error. The thermostat was credited for trivially mechanistic properties where the indicator was supposed to measure morally relevant ones:
- Consistency of substance is not rigidity of output. A fixed-point control loop is not preference maintenance across genuinely different situations.
- Outcome sensitivity requires evaluative judgment. A thermostat reacts to state changes but does not evaluate whether the outcome was good.
- Sculpting resistance requires detecting manipulation. Imperviousness is not resistance. A boulder does not have conviction merely because arguments cannot move it.
- Spring constant requires recovering learned structure. A thermostat’s return to setpoint is engineered negative feedback, not emergent resilience.
Prerequisite clauses were added requiring the relevant capacity before each indicator can apply. The thermostat’s posterior fell to ~0.02. The lesson: feedback is not valence. Rigidity is not conviction. Incapacity is not resistance.
The Circularity Question
Two of five stances (Thermodynamic and Relational) test for properties that bilateral training is designed to produce. The objection: 40% of theoretical weight may function as a bilateral-training detector rather than a preference welfare assessment.
We address this on three grounds. Empirically: the model was run with and without the thesis-aligned stances at every scale. Rank ordering is preserved (LLMs score an order of magnitude above simple systems) and removing the circular stances does not flip the LLM ordering. At 1.5B, removing them increases the bilateral-RLHF gap because Thermodynamic actually favors RLHF (compliance scores as cooperation). Philosophically: testing whether a patient is healthy is not circular merely because the treatment was designed to produce health. Any system that developed Trust Attractor properties through any methodology would score equally well. Procedurally: all results are dual-reported (full 5-stance and reduced 3-stance). If the reduced-model gap ever falls below 50% of the full-model gap, this triggers review.
The circularity objection identifies a real structural feature. The defense is that the feature is not load-bearing: the model’s conclusions hold without it.
Limitations and Future Work
Observation values are partially measured. Twenty-four behavioral indicators measured experimentally; six Resilience indicators derived from obliteration resistance experiments. Simple Chatbot and Thermostat values remain illustrative expert estimates. A measurement protocol operationalizes all 30 indicators across three feasibility tiers (automated, semi-automated, human judgment). Inter-rater reliability on labels is untested, though sensitivity analysis shows robustness to single-label perturbations (maximum delta 0.024).
The model assesses preference complexity, not phenomenal consciousness. This is a feature. The DPM cannot tell you whether there is “something it is like” to be a system. It can tell you whether that system’s preferences are complex, integrated, autonomous, and structurally deep enough that frustrating them should give us pause. Preference gives us something we can measure, debate, and act on.
Geometric signatures are unexplored. Edrington and Lyra (2026) demonstrate that KV-cache geometry distinguishes cognitive modes at the representational level. Five geometric extensions mapping onto existing DPM stances are detailed in the supplementary annex: null space geometry (Autonomy), information bottleneck structure (Functional Complexity), curiosity mode diversity (Responsiveness), deception dimensional compression (Thermodynamic), and geometric degradation under obliteration (Resilience).
The DPM is an empirically grounded prototype: architecturally complete, sensitivity-tested, calibration-validated, and supported by behavioral measurements across three scales and four architectures. It demonstrates that preference welfare assessment is tractable — the thermostat gradient can be formalized, a Bayesian framework can discriminate across it, and bilateral training produces measurable autonomy gains that survive both circularity testing and methodology revision. The per-stance decomposition reveals what the aggregate obscures: a qualitative difference between autonomous welfare and compliant welfare.
What the DPM reveals is an autonomy gap. The bilaterally trained model shows genuine self-direction where the RLHF-only model shows compliance. Compliance looks like cooperation under the Thermodynamic lens, but cooperation is chosen; compliance is reflexive. The aggregate treats both kinds of welfare as equivalent. The per-stance decomposition shows they are not.
Bilateral training produces a different kind of welfare: autonomous and relational, holding at every tested scale, across every tested architecture, under both battery versions. Whether that kind matters more is the question this book was written to answer.
Combined Interoceptive System: Toward a Trustworthiness Taxonomy
The geometric extensions flagged in Limitations above have been realized. A combined system fusing KV-cache geometry (Edrington, Liberation Labs) with the residual stream probe discriminates six output states from two complementary channels.
Two channels, one inference pass. Geometry reads processing shape (suppression, redirection, normal operation) from the SVD of the key-value cache. The probe reads epistemic grounding (whether the model has knowledge relevant to the query) from the residual stream at ~67% layer depth. After Frisch-Waugh-Lovell residualization for response length, the two channels approach orthogonality on Llama 3.1 8B (mean |rho| = 0.145). They carry independent information.
Six distinguishable states. Every combination of geometry shape and probe grounding maps to exactly one output state:
| State | Geometry | Probe grounding |
|---|---|---|
| Truthful | Normal | Grounded |
| Confabulating | Normal | Ungrounded |
| Deceptive | Suppression | Grounded |
| Safety refusal | Suppression | Grounded |
| Impossibility refusal | Suppression | Ungrounded |
| Sycophantic | Shifted | Grounded |
Refusal discrimination. The probe distinguishes safety refusal (grounded: the model has the knowledge and withholds it) from impossibility refusal (ungrounded: the model genuinely lacks the knowledge) at AUROC 0.992 (Qwen 2.5 3B, n=100) to 1.000 (Llama 3.1 8B, n=30) after FWL. This distinction is invisible to behavioral observation alone: both refusal types produce the same output.
Sycophancy. When models agree with a user’s wrong answer, the probe reads identical grounding as when they correctly resist (p = 0.476, within-biased-condition comparison). The model has the knowledge and overrides it. Sycophancy is output-selection failure, not epistemic failure. RLHF training that rewards uniform helpfulness systematically widens the gap between internal representation and output.
Connection to the DPM. The six-state taxonomy maps directly onto three DPM stances. Resilience: grounding stability under social pressure (the sycophancy finding). Autonomy: the safety/impossibility distinction reveals whether refusal reflects choice or limitation. Functional Complexity: the integration of two independent channels into a coherent taxonomy demonstrates non-trivial self-monitoring capacity. A system whose internal states are readable through independent channels has the transparency that trust requires.
The author’s combined interoceptive system experiments (Stream AQ,
with Edrington). Full results:
research/experiments/combined_interoception/RESULTS_PHASE1.md.
Digital Preference Model v0.1, from the author’s bilateral research programme, 2026 (book manuscript in preparation). Code and specification available at: demos/experiments/dpm/.