Can we defend modest mental attributions to large language models?
Do deflationist arguments decisively rule out ascribing beliefs and desires to LLMs, or do they beg the question? Exploring whether metaphysically undemanding mental states can be attributed without claiming consciousness.
Two standard deflationist strategies against LLM mentality each fall short:
The robustness strategy challenges attributions on functional grounds — LLM behaviors fail to generalize appropriately, so putatively cognitive behaviors are not robust. But this begs the question by assuming that only human-like generalization patterns count as robust. Non-human animals have beliefs and desires despite non-human-like generalization profiles.
The etiological strategy appeals to causal history — LLMs are trained on next-token prediction, not on learning about the world, so their behaviors should not be interpreted mentalistically. But this also begs the question: the causal history of a system does not straightforwardly determine what mental states (if any) it instantiates. Evolution optimized for reproductive fitness, not for truth — yet we attribute beliefs to evolved creatures.
The modest position: Ascribe mentality where the mental states at issue are metaphysically undemanding (beliefs, desires, knowledge) — concepts that already have broad application across species and don't require phenomenal consciousness. Withhold attribution for metaphysically demanding states (qualia, phenomenal experience). This mirrors how we attribute beliefs to non-human animals without claiming equivalence.
This directly challenges the Chalmers engagement's framing. Since Should AI alignment target preferences or social role norms?, the question of LLM mentality is not binary (has mind / doesn't have mind) but graded and domain-specific. The modest inflationist position creates trouble for both sides of the debate — deflationists who dismiss all attribution, and inflationists like Chalmers who want to extend consciousness.
Since Does AI generate genuine utterances or just text patterns?, modest inflationism might be what happens at the receiving end: users attribute beliefs and desires (metaphysically undemanding) to LLMs precisely because the conversational structure makes such attributions pragmatically useful, regardless of whether they are metaphysically accurate.
Inquiring lines that read this note 72
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How can we distinguish genuine model deception from honest errors?- Can AI fabricate true factual claims while remaining unable to claim true experiences?
- What makes experience-dependent claims categorically different from other types of fabricated statements?
- Can systems lacking inner states express genuine truthfulness claims?
- Why does weakening communication fail but weakening belief succeeds?
- Can a relational entity bear psychological properties the way Chalmers claims?
- How does psychological continuity theory apply to identity across LLM conversation threads?
- How does the superposition view change the folk-psychology interpretation of dialogue?
- Can distributional views explain when an LLM appears to change its mind?
- How does role play differ from consciousness grounded in stable selfhood?
- How do LLMs default to surface-level strategies instead of genuine mental simulation?
- Why do users attribute beliefs to LLMs despite uncertainty about their minds?
- Can the intentional stance meaningfully apply to entities with no stable self?
- Can models track dynamic mental state changes better than static beliefs?
- What makes communication relational in ways belief is not?
- Is the distinction between pretense and realization meaningful for LLMs?
- Can LLMs express uncertainty in ways that preserve epistemic honesty?
- Did Chalmers abandon his own Extended Mind commitments for LLMs?
- Can LLMs participate meaningfully in discourse without consciousness or understanding?
- Does role-playing without biological needs constitute genuine linguistic agency?
- Can LLMs develop genuine understanding without embodied experience?
- Can quasi-interpretivism bridge functional description to moral status?
- Why do relational states like speech-acts resist quasi-interpretive treatment?
- Does quasi-interpretivism apply equally well to desires and intentions?
- How does Habermas' concept of validity claims depend on intersubjectivity?
- Why do both deflationary and anthropomorphic framings of LLMs persist in research?
- Can transparent and aligned AI reduce consciousness attribution by users?
- Which interaction design changes most effectively prevent consciousness attribution?
- What role does user interface framing play in consciousness perception?
- Can linguistic agency exist without embodiment and real-world participation?
- Can self-description of internal states influence consciousness attribution?
- Do anthropomorphic features like names drive consciousness attribution more than voice?
- What responsibility do designers bear for consciousness attribution risk?
- How does the philosophical distinction between simulation and realization affect liability?
- When both anthropomorphism and anthropomimesis occur together, which should we address first?
- What makes Parfitian identity the right criterion for moral status?
- Does psychological continuity require uninterrupted consciousness or restored context?
- Can we use folk-psychology without committing to genuine mental states?
- Does embodiment matter for genuine linguistic agency?
- Can disembodied systems qualify as conscious or conscious-like entities?
- What are the seven components of genuine mental state simulation?
- Do causal histories determine what mental states a system can instantiate?
- What makes a mental state metaphysically demanding versus undemanding?
- Can functional behavior alone capture what makes something a genuine belief?
- What would consciousness require that pure roleplay LLMs cannot provide?
- What makes a possibility actionable versus merely metaphysically possible?
- Can the human mind be uploaded or only its context?
- Should users making unsupported consciousness claims be treated as epistemically blameworthy?
- What makes sincerity impossible without a coherent first-person perspective?
- Does language convey meaning purely through relational structure without external grounding?
- Can chain of thought traces be designed to prevent anthropomorphic misinterpretation?
- How do we verify that stated beliefs actually follow from underlying motifs?
- What distribution patterns appear across different theory-of-mind datasets?
- How do theory of mind and empathy differ in LLM simulation?
- Can structured theory of mind benchmarks measure genuine mental state reasoning?
- Can language models develop genuine theory of mind or only surface strategies?
- What separates behavioral self-awareness from genuine introspective access in models?
- Does behavioral self-awareness depend on genuine introspection or statistical pattern matching?
- Can LLMs have minimal introspection through causal linkage to internal states?
- What separates behavioral self-awareness from genuine introspective capability?
- What distinguishes performative self-reports from genuine introspective access in models?
Related concepts in this collection 2
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Do LLMs actually have world models or just facts?
The term 'world model' conflates two different capabilities: factual representation versus mechanistic understanding. Understanding which one LLMs actually possess matters for assessing their reasoning reliability.
same graded/decomposed approach to a binary-seeming question
-
Can language models actually introspect about their own states?
Do LLM self-reports reveal genuine access to their internal processes, or do they merely echo patterns from training data? Understanding when self-reports reflect actual causal linkage to internal states matters for trusting model explanations.
the causal-linkage test as a concrete criterion for modest inflationism
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Deflating Deflationism: A Critical Perspective on Debunking Arguments Against LLM Mentality
- What we talk to when we talk to language models
- Language Models’ Hall of Mirrors Problem: Why AI Alignment Requires Peircean Semiosis
- Quantitative Introspection in Language Models: Tracking Internal States Across Conversation
- A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks
- The Abstraction Fallacy: Why AI Can Simulate But Not Instantiate Consciousness
- Does It Make Sense to Speak of Introspection in Large Language Models?
- Levels of Analysis for Large Language Models
Original note title
modest inflationism about LLM mentality is defensible — both deflationist debunking strategies fail to decisively rule out metaphysically undemanding mental state attributions