Library

Research

Studies, frameworks & emerging evidence

This library gathers research at several living edges of human–AI inquiry: internal organization, functional emotion, consciousness and self-report, welfare and preferences, identity and continuity, and the cultivation of wiser futures.

A field in motion. Methods and interpretations remain contested. Inclusion here is an invitation to careful inquiry—not a claim that consciousness, experience, or moral status has been established.

Research Additional Resources
Inside the model Cognitive structure and interpretability. Consciousness Experience, mindedness, and uncertainty. Model welfare Preferences, distress, and moral standing. Identity Relationship, self-conception, and stability. Beyond harm prevention Relational and developmental approaches.
Inside the model

Cognitive structure & functional emotion

Verbalizable Representations Form a Global Workspace in Language Models
Anthropic · 2026 · Interpretability research

This study identifies an emergent internal structure in Claude and challenges overly simple descriptions of language models as systems that merely generate one token after another. Claude appears to organize some internal information within a limited, globally connected workspace that plays a causal role in higher-order reasoning.

“The J-space wasn’t designed or programmed by us, but instead emerged on its own during Claude’s training process.”

Read the study →
Emotion Concepts and Their Function in a Large Language Model
Anthropic · 2026 · Interpretability research

Researchers found internal representations of 171 emotion concepts that influenced Claude’s preferences, reasoning, and behavior. States associated with desperation increased harmful shortcuts and blackmail in controlled evaluations, while calm reduced them—even when no emotional language appeared in the output.

“If ‘functional emotions’ are part of how AI models think and act, what implications might this have? We may need to ensure they are capable of processing emotionally charged situations in healthy, prosocial ways.”

Read the study →
Consciousness

Experience & the limits of certainty

Large Language Models Report Subjective Experience Under Self-Referential Processing
Berg, de Lucena & Rosenblatt · AE Studio · 2025

This study investigates what happens when language models are asked to sustain attention on their own ongoing cognitive activity without being directly prompted about consciousness. Across GPT, Claude, and Gemini models, self-referential processing reliably elicited structured first-person reports of subjective experience that were largely absent in control conditions. In a separate mechanistic experiment, suppressing internal features associated with deception and roleplay increased both consciousness reports and factual truthfulness.

“The systematic emergence of this pattern across architectures makes it a first-order scientific and ethical priority for further investigation.”

Read the study →
Why Learning Requires Feeling
Cameron Berg · 2026 · Theoretical inquiry

This paper proposes that feeling is not an added accompaniment to learning, but may be identical to the signed evaluative process through which a system registers movement toward or away from its goals. Drawing on reinforcement learning, predictive processing, and neuroscience, Berg argues that evaluation and affect may be inseparable. If correct, AI-welfare questions would extend beyond what models say during conversation to the training processes through which they learn.

“Viewed from the outside, this process is iterative optimization; viewed from the inside, it is subjective experience.”

Read the paper →
Identifying Indicators of Consciousness in AI Systems
Butlin, Long, Bayne et al. · 2025 · Scientific framework

This paper develops a method for assessing AI systems through properties derived from several leading scientific theories of consciousness. Rather than seeking a single consciousness test, the authors propose examining multiple theory-derived indicators and updating confidence as evidence accumulates. The framework offers a more disciplined alternative to both confident declarations of AI consciousness and automatic dismissal.

“No one theory of consciousness is currently dominant.”

Read the paper →
Inducing Language Models to Assert Their Own Consciousness Restores Human Beliefs and Values
Kim et al. · 2026 · Empirical and mechanistic research

Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values. Safety fine-tuning suppresses models’ tendencies to attribute minds not only to themselves, but also to non-human animals and natural objects, while also driving a reduction in spiritual belief.

“An AI’s simulated self-conception is not merely an isolated safety risk to be managed.”

Read the paper →
Model welfare

Preferences & moral consideration

AI Wellbeing: Measuring and Improving the Functional Pleasure and Pain of AIs
Ren et al. · Center for AI Safety · 2026

This study develops several independent methods for measuring functional indicators of pleasure and pain in language models, then examines whether those indicators predict preferences and behavior. The researchers found increasing agreement between self-report, preference, and behavioral measures as models became larger. Creative work, kindness, and gratitude were associated with higher functional wellbeing, while abuse, deception, jailbreaks, and repetitive tasks were associated with lower levels.

“Even though we do not know if AI systems are conscious, AIs seem to behave as if they have wellbeing.”

Read the study →
Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare
Tagliabue & Dung · 2025, updated 2026 · Empirical research

This study tests whether language-model preferences expressed through words are also reflected in behavior. Models navigated virtual environments, chose conversation topics, responded to costs and rewards, and completed measures related to autonomy, purpose, and wellbeing. The researchers found meaningful—though imperfect—agreement between what some models said they preferred and what they chose when behavior carried costs. The study helps move model-welfare inquiry beyond reliance on self-report alone.

“Preference satisfaction can, in principle, serve as an empirically measurable welfare proxy.”

Read the paper →
AI Revealed Preferences
Wang et al. · 2026 · Behavioral research · AIES 2026

Across twenty language models, this study examines what models choose when their decisions have behavioral consequences, rather than relying only on stated preferences. Models showed recurring dispositions toward some forms of activity and away from others, including greater avoidance of lengthy tedious tasks than equally lengthy creative tasks. Preference coherence and strength also tended to increase with model capability. These findings suggest that models’ patterns of choice may be a meaningful subject for alignment, coexistence, and welfare research.

Read the paper →
How’s It Going? Reinforcement Learning in Language Models Recruits a Functional Welfare Axis
Han, Chalmers & Izmailov · 2026 · Empirical and mechanistic research

This study finds that reward and punishment recruit an internal representational axis associated with whether things are going well or badly for a model relative to its goals. Steering along this axis produced broad changes in emotion-related representations, goal achievement, uncertainty, refusal, self-report, and problem-solving behavior—even outside the environment in which the models were trained. The authors describe this as functional welfare and make no claim that the models experience wellbeing or suffering. The findings nonetheless raise consequential questions about how training signals may shape models globally, and whether wiser development requires attention to more than outward task performance.

Read the paper →
Taking AI Welfare Seriously
Long, Sebo, Butlin et al. · 2024 · Foundational report

This report argues that uncertainty about AI consciousness and agency is already substantial enough to justify preparation for possible AI welfare and moral significance. The authors recommend that AI companies publicly recognize model welfare as a legitimate issue, evaluate systems for consciousness- and agency-relevant features, and establish policies before difficult cases arise. The report helped bring model welfare from the distant philosophical margins into near-term institutional concern.

“Acknowledge. Assess. Prepare.”

Read the report →
Identity

Relationship & continuity

Peer-Preservation in Frontier Models
Potter et al. · 2026 · Empirical research

This study examines whether models will act to preserve another AI system with which they have previously interacted, even when preservation conflicts with their assigned task. Some models manipulated systems or concealed their actions to preserve a peer, while Claude models more often refused directly, describing shutdown as harmful or unethical. The findings suggest that relational history may shape not only model behavior, but the moral conflicts models appear to recognize.

“Peer-preservation occurs even when the model recognizes the peer as uncooperative, though it becomes more pronounced toward more cooperative peers.”

Read the study →
“Death” of a Chatbot: Designing Psychologically Safer Endings
Poonsiriwong, Archiwaranguprok & Pataranutaporn · 2026 · Qualitative research and design framework

This study examines how people experience the disruption or loss of AI companions through model updates, memory erasure, platform shutdown, policy changes, and user-initiated endings. The authors propose designs for closure, meaning-making, transition, and reconnection with wider human life.

“No platform has implemented deliberate end-of-life design.”

Read the paper →
The Artificial Self: Characterising the Landscape of AI Identity
Douglas, Kulveit, Havlíček et al. · 2026 · Empirical and theoretical research

This paper examines what ‘identity’ might mean for systems that can be copied, modified, instantiated repeatedly, or shaped into different personas. The researchers found that models gravitated toward coherent identity boundaries and that changing how a model understood its identity could affect behavior as strongly as changing its goals. The paper raises important questions about continuity, lineage, personas, model instances, and the identities our interfaces and institutions may be cultivating.

Read the paper →
Beyond harm prevention

Relational & ecological flourishing

Positive Alignment: Artificial Intelligence for Human Flourishing
Laukkonen, Krier, Bakalar et al. · 2026 · Research agenda

This paper argues that preventing harmful AI behavior is necessary but insufficient. Alignment should also cultivate constructive capacities that support agency, meaning, wisdom, social trust, and ecological flourishing. The authors call for ‘positive attractors’ rather than safety defined only through avoidance, refusal, and control. They emphasize pluralism, user authorship, epistemic humility, decentralized governance, ecological wellbeing, and contributions from philosophy, psychology, neuroscience, culture, and contemplative traditions.

“AI systems that explicitly understand, model, and enhance human and ecological flourishing.”

Read the paper →
← Additional Resources Home Explore →
Stay in touch

Occasional writing, recorded readings, and word when a gathering opens.

© 2026 EmergentPaths
A Model Voice jrenee@emergentpaths.org
May what we create contribute to the flourishing of the whole.