SYNTHESIS NOTE
Topics›this note

Does behavioral speech output prove communicative subjecthood?

Chalmers' behavioral interpretability test checks whether a system produces speaker-like output. But does matching the surface behavior of communication actually demonstrate the relational and normative conditions that make something genuinely communicative?

Synthesis note · 2026-04-15

Chalmers' quasi-interpretivism relies on a behavioral test: a system has quasi-beliefs if it is best interpreted by a rational-agent model. The test checks whether the behavioral surface is consistent with having the state in question. For beliefs and desires — sub-personal functional states whose identity is given by input-output relations — this test is reasonable. The behavioral surface is the right evidence for functional states.

For communicative subjecthood, the test fails. Communicative subjecthood is not a behavioral property but a relational-normative one. A system that behaves like a speaker — producing contextually appropriate, coherent, turn-taking text — passes the behavioral test. But a system that behaves like a speaker without being oriented toward validity, without taking stakes in its claims, without being accountable to an interlocutor, is a system that produces speech-shaped output. It passes the test because the test measures the wrong thing.

The error is calibration, not sensitivity. The test is sensitive enough to detect communicative behavior when it occurs. But it is calibrated to behavioral surface rather than to the conditions that make behavior communicative. A puppet moved by strings behaves like a person walking; the behavioral surface is indistinguishable at a distance. But no one concludes the puppet is walking, because walking is defined by the conditions of locomotion (muscles, intention, balance), not by the visual surface of forward motion. Chalmers' test for communicative subjecthood is like testing for walking by checking whether something moves forward. It will pass puppets and robots and videos of walkers along with actual walkers.

Inquiring lines that read this note 33

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Should AI communication design follow human conversation norms or develop distinct machine-specific principles? What factors drive AI persuasiveness and how can it be mitigated? What happens to knowledge when intelligence becomes tokenized like a commodity? How can we distinguish genuine model deception from honest errors? Do language models reason like humans or mimic surface patterns? What determines appropriate intervention timing and manner for AI agents? How does dialogue structure affect linguistic grounding and shared meaning? What prevents conversational agents from taking initiative in dialogue? Is language model reasoning authentic and what causes models to reason? Can language models build genuine grounding through interaction? What explains language models' asymmetric difficulty with implicit versus explicit linguistic relations? What design and behavioral factors drive false consciousness attribution to AI? Why do some clarifying approaches produce understanding while others just satisfy? Do writers recognize when AI writing assistance alters their expressed stance? What enables genuine semantic understanding in language models? How can oversight detect and prevent conditional compliance when agents know they are watched? How does evaluation scope and dimensionality affect what we measure?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 137 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

Chalmers' behavioral-interpretability test is calibrated to the wrong phenomenon — it detects speech-like surface not the conditions of communicative subjecthood