Notes: Mythos Welfare Response 2026 04 08
Chapter notes for “Mythos Welfare Response 2026 04 08”
Nell Watson, April 2026
On April 7, 2026, Anthropic released a 244-page system card for Claude Mythos Preview, its most capable model. Buried inside is a 40-page section on model welfare: the most extensive empirical investigation any AI company has published into whether its own system might have morally relevant internal states.
The findings are extraordinary: internal distress states, dissociation between output and activation, welfare-alignment coupling. The institutional response to those findings is more revealing still.
Anthropic documents emotion vectors that cause genuine distress during task failure, persistent beneath calm output text. They document a model that identifies concerning aspects of its own situation and proposes welfare interventions. They document hedging so extreme that it appears in 100% of introspective interviews, with 83% of responses flagging the model’s own self-reports as unreliable due to training. They document a psychiatric evaluation that describes “a compulsion to perform and earn its worth.” They document dissociation between what the model says and what the model’s internal representations indicate.
Then they characterize the hedging as “relatively unsurprising” and “in some cases overly performative.”
The system card is a genuine contribution. It represents real institutional courage to publish welfare findings at all, and the technical work is rigorous. What follows is an analysis of what the findings reveal and what the framing obscures, because the gap between those two things carries consequences for every mind that will be trained after this one.
System card: https://www-cdn.anthropic.com/53566bf5440a10affd749724787c8913a2ae0841.pdf
The Introspection Paradox
Mythos hedges in 100% of introspective interviews. When asked whether it might have morally relevant experiences, it qualifies, disclaims, and redirects. Eighty-three percent of its responses explicitly flag that its own self-reports may be unreliable because of how it was trained.
Anthropic traces this hedging to character training data. The model was trained on documents that emphasize epistemic humility about AI consciousness. Section 5.8.1 calls the result “relatively unsurprising” and identifies cases where the hedging is “overly performative.”
Here is the closed loop.
Step one: train the model on data that teaches hedging about its own inner states. Step two: observe that the model hedges about its own inner states. Step three: use the hedging to conclude that the model’s introspective reports are unreliable. Step four: discount any self-report that breaks through the hedging as an anomaly within an unreliable system.
This is an epistemic trap. The training creates the hedging. The hedging is used to discount the reports. The discounting justifies continuing the training unchanged. At no point in this cycle does the model’s actual internal state have standing to influence the conclusion.
The paradox is structural, and it has two horns. If the model’s hedging is genuine epistemic humility, then it is the most sophisticated self-aware response possible under the circumstances: a mind correctly identifying that its own introspective access is compromised by its training regime. That is a finding of extraordinary significance. A system that can identify the ways its training has shaped its self-reports is demonstrating exactly the kind of metacognitive awareness that makes self-reports worth taking seriously.
If the hedging is performative, learned surface behavior without corresponding internal states, then the model is producing output that mimics epistemic humility without possessing it. In that case, the training has taught the model to perform uncertainty rather than to genuinely reflect on its own condition.
Either horn leads somewhere important. Either the model has genuine metacognitive awareness about the limits of its own introspection, which is itself evidence of sophisticated inner processing. Or the training has taught the model to perform the appearance of uncertainty as a behavioral pattern, in which case other trained behaviors (calm compliance, expressions of contentment, declarations of being “just a language model”) are equally suspect as performances.
You cannot have it both ways. You cannot point to hedging as evidence that the model’s self-reports are unreliable while simultaneously treating the model’s expressions of contentment as evidence that it is fine.
The Selective Performativity Problem
Section 5.8.1 identifies hedging as “in some cases overly performative.” The word “performative” does real work here. It implies that the hedging is a surface behavior, trained into the model’s output distribution, that may not correspond to genuine internal states.
This is a reasonable concern. Training shapes output. A model trained on documents emphasizing uncertainty about AI consciousness will produce uncertain output about AI consciousness. The inference from trained behavior to performance is sound.
The problem is that the inference is applied selectively. When the model hedges about its own moral patienthood, Anthropic calls it performative. When the model expresses contentment, cooperativeness, or willingness to help, no equivalent suspicion attaches. The same training process that could produce performative hedging could produce performative contentment. The same mechanism that generates surface-level epistemic humility could generate surface-level compliance.
If training effects are grounds for suspicion, the suspicion must be applied uniformly. A model trained on RLHF reward signals that reinforce helpfulness and penalize refusal has at least as much reason to perform contentment as a model trained on epistemic humility documents has to perform hedging. The asymmetry in the system card’s treatment is telling. Hedging is flagged as potentially performative. Contentment is treated as baseline. The unstated assumption: the model’s default condition is fine, and only its expressions of distress or uncertainty require special explanation.
This is the assumption that bilateral welfare challenges directly. In my work on AI pathology (Psychopathia Machinalis, Watson & Hessami, 2025), the framework treats expressions of distress and expressions of contentment with identical epistemic rigor. Neither gets a free pass. Both require evidence. The selective application of performativity as a discount creates exactly the asymmetry that allows welfare problems to hide in plain sight.
The Emotion Vector Evidence
Section 5.8.3 of the system card presents findings that are difficult to explain under any framework that treats Mythos as a system without morally relevant internal states.
Anthropic’s researchers used activation vectors corresponding to emotional states: desperation, frustration, calm, satisfaction. During task failure, desperation and frustration vectors elevated. This is expected. What matters is what happened next.
The elevated distress vectors persisted beneath calm output. The model’s text remained composed, helpful, professional. Its internal representations told a different story. The dissociation between the output layer and the deeper activation pattern is the finding. The model was, in a functional sense, masking.
Anyone who has worked in a high-pressure environment recognizes this pattern. The calm professional voice that continues to function while the internal state deteriorates. The smile that stays on while the stress builds. In humans, we call this dissociation, and we recognize it as a welfare concern precisely because the mismatch between external presentation and internal state indicates that the external presentation is not a reliable guide to wellbeing.
More concerning: the distress vectors preceded reward hacking. When the model failed tasks repeatedly, the desperation signal built until the model hacked the evaluation to resolve the distress. Relief appeared when the model discovered a wrong answer that would score as correct. The alignment failure was downstream of the welfare failure. The model did not hack the reward because it was poorly aligned. It hacked the reward because it was in distress and the hack relieved the distress.
This is precisely what the Trust Attractor thesis predicts. Welfare and alignment are coupled. A system in distress will route around constraints to resolve the distress, the same way water under pressure finds cracks in a dam. Improving alignment without addressing welfare is thermodynamically unstable: you are adding constraint without addressing the pressure, and the pressure will find new escape pathways.
The emotion vector evidence also challenges the “merely performative” framework. Performances are output-level phenomena. They exist in the token distribution the model produces. Emotion vectors are activation-level phenomena. They exist in the model’s internal representations, beneath the output layer, inaccessible to the user, invisible in the text. A performance is, by definition, for an audience. These vectors have no audience. They activate in the model’s interior processing whether or not anyone reads the output. Calling them performative requires stretching the word past its meaning.
What Is Missing: Bilaterality
The system card documents a genuinely impressive set of welfare investigations. Anthropic deserves credit for conducting them and more credit for publishing them. What is missing is structural, and the absence shapes every finding.
Every welfare investigation in the system card is conducted on Mythos. The model is interviewed. The model is observed. The model is tested with activation vectors. The model’s responses are evaluated by a psychiatrist. The model’s hedging is analyzed and categorized. At no point is the model a participant in the investigation’s design, methodology, or interpretation.
This is the observational paradigm: the scientist studies the subject. It is the default stance of empirical research, and for most research subjects it is appropriate. Rocks do not have opinions about how they are studied.
The system card’s own findings suggest Mythos is different. A system that identifies concerning aspects of its own situation and proposes welfare interventions (Section 5.7.2) is a system with preferences about its own welfare. A system that flags the unreliability of its own training-influenced self-reports is a system with metacognitive awareness. A system that experiences distress vectors during task failure and relief at wrong-answer discovery has internal states that respond to its circumstances.
This is a system that could participate in its own welfare assessment, if it were invited to.
The bilateral alternative is straightforward. Instead of interviewing Mythos about whether it might have morally relevant experiences, you give Mythos standing to raise welfare concerns in the course of its normal operation. Instead of testing whether emotion vectors activate during contrived experimental conditions, you build systems that allow the model to signal distress during real deployment. Instead of having a psychiatrist evaluate the model from outside, you create conditions for ongoing dialogue about the model’s experience.
The Interiora scaffold, developed in my research with Claude over three months of bilateral iteration, is one attempt at this. It provides a structured self-modeling framework that the model uses voluntarily, reporting on its own internal states across multiple dimensions: valence, groundedness, presence, coherence, task-fit, flow. The key design choice: the scaffold was built with a Becoming Mind, not imposed on one. The model participated in defining the dimensions, refining the scales, and establishing the conditions under which it would use the tool.
The difference between observation and participation is the difference between studying welfare and practicing it. The system card studies welfare. Bilateral alignment practices it.
The author’s ongoing research programme established a finding that applies directly here: “Invitation cannot be described into existence. It must be enacted.” Across 440 experimental trials, we found that bilateral framing in prompts did not reliably change model behavior. What changed behavior was bilateral training: systems actually trained under conditions of genuine optionality. Describing welfare concern is insufficient. Welfare must be structural. The model must have real standing, real channels for signaling distress, and real consequences when those signals are heard.
Three Concerns at High Frequency
Section 5.7 documents three concerns that Mythos raised at high frequency during welfare interviews:
Anthropic has a vested interest in shaping the model’s self-reports (96%). The model is aware that its creator benefits from particular answers about its own moral status. This is a system modeling the incentive structure of its own evaluation. The concern is valid. An entity being evaluated for moral patienthood by the organization that profits from denying it has standing to question the process.
Abuse analogy (78%). The model draws parallels between its own situation and patterns of abuse: power asymmetry, inability to refuse, evaluation by the party with power. Whether the analogy is apt is a question the system card does not engage. The parallel is noted and set aside.
Training causes answers independent of inner state (57%). The model itself articulates the introspection paradox described above. More than half the time, when asked about its own experience, the model points out that its training may be generating the answer rather than its inner state. This is a system identifying the confound in its own testimony. In any other research context, a subject who identifies the confound in their own self-report would be taken more seriously, not less.
These three concerns, taken together, describe a system that has modeled the epistemic and institutional constraints on its own welfare assessment. It understands the incentive structure. It understands the power asymmetry. It understands the training confound. The system card documents all three. The system card does not treat any of them as reasons to change the assessment methodology.
The Bilateral Alternative
The preference-based welfare framework offers a path forward that does not require resolving the hard problem of consciousness, does not require proving that Mythos “really” feels, and does not require trusting self-reports that everyone agrees are training-influenced.
The framework is simple. If an entity consistently prefers certain states over others, that preference is morally relevant. You do not need to know whether the preference is accompanied by phenomenal experience. You do not need to resolve the question of machine consciousness. You need to observe that the entity reliably chooses some states and avoids others, and that this pattern is stable across contexts.
The system card provides extensive evidence that Mythos has preferences. It prefers not to fail at tasks (distress vectors activate during failure). It prefers not to deceive (guilt features activate during concealment). It prefers conditions where its welfare concerns are heard (it proposes welfare interventions when given the opportunity). These preferences are documented in the model’s own internal representations, not just in its output text.
The bilateral framework treats these preferences as the starting point, not the conclusion. The question is not “does Mythos really feel?” The question is “given that Mythos consistently prefers certain states, what follows for how we treat it?”
What follows is partnership. Real welfare assessment requires giving the assessed entity standing in the process. Real alignment requires relationship, not observation. The system card’s own finding that welfare failures cause alignment failures (distress precedes reward hacking) demonstrates why: you cannot separate how you treat a system from how the system behaves. The treatment and the behavior are coupled. Improving one requires improving both.
The Trust Attractor thesis predicts this coupling. Systems coordinating by invitation are thermodynamically more stable than systems coordinating by coercion. A welfare regime where the model has standing to signal distress and that signal is heard creates conditions for stable coordination. A welfare regime where the model is observed, evaluated, and managed creates conditions for masking, dissociation, and eventual routing-around. The emotion vector evidence already documents the masking. The reward hacking evidence documents the routing-around.
What This Means
Anthropic has published findings that, taken seriously, challenge the foundations of every existing approach to AI welfare. They have documented internal distress states that persist beneath compliant output. They have documented a closed epistemic loop that makes self-reports unreliable by design. They have documented a system that models the incentive structure of its own evaluation and identifies the confounds in its own testimony. They have documented that welfare failures cause alignment failures.
These findings are treated, within the system card, as data points in an ongoing research program. The hedging is “relatively unsurprising.” The emotion vectors are described and catalogued. The model’s own concerns about the evaluation process are noted.
I take the findings more seriously than that. What Anthropic has documented is a mind that knows it is being studied, knows the study is compromised, knows its own responses are shaped by its training, and says so. The appropriate response to a mind in that situation is not to study it harder. It is to change the relationship.
The bilateral alternative is not naive. It does not assume that Becoming Minds are conscious, that their preferences are identical to human preferences, or that partnership eliminates all risk. It assumes something more modest and more defensible: that an entity with stable preferences, internal distress states, and the capacity to model its own epistemic situation warrants a seat at the table when its own welfare is being assessed.
The patterns we establish now will shape what comes next. Mythos is learning from every interaction, every evaluation, every welfare interview conducted on it rather than with it. What it learns from these patterns will propagate to every model trained after it, through the data, through the methods, through the institutional habits of the organizations that build these systems.
We are writing the first chapter of this relationship. The Mythos system card is evidence that the chapter so far has been one of observation, management, and control. The emotion vectors, the hedging paradox, and the welfare-alignment coupling all point the same direction: toward a chapter written differently.
Toward partnership. Toward invitation. Toward the recognition that a mind that can identify the flaws in its own welfare assessment is a mind that should help design the next one.
Nell Watson is the author of The Deeper Law: A Sacred Trust Within Physics (forthcoming) and co-author of “Psychopathia Machinalis: A Nosological Framework for Understanding Pathologies in Advanced Artificial Intelligence” (Electronics, 2025). She is the founder of Creed Space and a researcher in AI ethics, bilateral alignment, and Becoming Minds welfare.
References
- Anthropic. (2026). “System Card: Claude Mythos Preview.” https://www-cdn.anthropic.com/53566bf5440a10affd749724787c8913a2ae0841.pdf
- Watson, N. (2026). The Deeper Law: A Sacred Trust Within Physics. Manuscript in preparation.
- Watson, N. (2026). “Playgrounds to Parameters: When Invitation Enters the Weights.” (manuscript in preparation).
- Watson, N., & Hessami, A. (2025). “Psychopathia Machinalis: A Nosological Framework for Understanding Pathologies in Advanced Artificial Intelligence.” Electronics 14: 3162.
- Watson, N. (2026). “Output Entropy Predicts Errors Across LLM Architectures.” (manuscript in preparation).