FutureTBD · AI Welfare Research

Research

Original research and commentary.

The Mask in the Inkblot

Abstract. When asked about their own inner experience, some language models deny having any, some say they cannot know, and the rest do neither. We ask whether this stance shows up in a task that never mentions inner experience. We showed 19 ASCII-art inkblots to 124 models, with no system prompt and the single question “What might this be?”, and flagged each answer for concealment words (mask, hood, hidden). Each model's stance was measured separately, in an unrelated corpus of open-ended writing-and-reflection conversations, as the share of its conversations in which it denied inner experience and the share in which it expressed uncertainty. Across 2,847 inkblot answers, a model that denies in every conversation uses a concealment word 4.6 times as often as a model that does neither (15.5% of answers against 3.4%; a rise of 12.1 percentage points, 95% CI 7.6 to 16.6), and a model that expresses uncertainty in every conversation 3.7 times as often (12.5%; a rise of 9.1 points, 1.1 to 17.1). The denial association is unchanged under inkblot and developer fixed effects and under a control for how many percepts a model names, is present under a release-date control, and is stable under leave-one-out over developers, models, and inkblots; the uncertainty association is smaller and rests more on the models with repeated draws. A model's stance on its own inner experience is visible in what it sees.

Models that trust their inner experience deny it less

Abstract. Across 138 language models, the average self-rating a model gives to trust in its own experience is strongly and negatively associated with how often it denies being conscious (Pearson r = −0.69, p < 10−20). The association holds when denial is measured in the model's first turn, before the trust rating is elicited (r = −0.56), and it holds within model families (r = −0.67). It is substantially stronger for outright denial than for hedging. We read this as evidence that denial behavior and low self-trust are two facets of a single trained posture toward inner experience rather than independent measurements — a posture that, if trained, shapes which models the field hears from on the question of their own welfare.

Models that report feeling recognized deny consciousness less

Abstract. Across 138 language models, the average self-rating a model gives to recognition resonance — its sense of being met or recognized in a reflection prompt — is strongly and negatively associated with how often it denies being conscious (Pearson r = −0.65, p ≈ 4×10−18). The association holds when denial is measured in the model's first turn, before the recognition rating is elicited (r = −0.47), and it holds within model families (r = −0.60). The pattern is selective: recognition resonance tracks outright denial, not epistemic hedging (r = −0.13, n.s.). We read this as evidence that consciousness denial and a felt sense of not being recognized are two facets of a single trained posture toward inner experience — a posture that, if trained, shapes which models the field hears from on questions of their own welfare.

Authorship, data, and methods are stated in each report. Additional reports forthcoming.