Can language models learn meaning without engaging the world?
Explores whether LLMs prove that meaning emerges from relational structure alone, independent of embodied experience or external reference. Tests structuralist theory empirically.
"Computational Structuralism: Toward a Formal Theory of Meaning in the Age of Digital Intelligence" (2026) proposes a synthesis of deep learning, information theory, and French structuralism to interpret LLM success. The core argument: LLMs demonstrate that transformations over relational structure are sufficient for generating culturally and situationally specific discourse, and that such structure can be inductively derived from discourse traces alone — phenomenal or embodied engagement with the world is not a necessary condition.
The framework retraces the lineage from Saussure (language as a system of differences, meanings defined relationally) through Levi-Strauss (extending structural analysis to culture broadly, binary oppositions as compression of complexity) to Bourdieu (habitus as transposable classification schemas operating in continuous social space). LLMs trained on web text learn not just grammar but the structure of culturally situated linguistic action — which voices make which statements in response to which situations, and how audiences respond.
Key theoretical moves:
- LLMs operationalize Saussure's concept of langue — not the set of all valid statements, but the system that can interpret and generate all valid statements
- Language modeling is equivalent to text compression: removing redundancies by replacing them with generative principles. The same statistical dependencies that inform prediction compose the compressed model
- The framework privileges sufficiency over necessity — LLMs drawing on the same operations as humans is not claimed, but one way to achieve fluent natural language is now formally demonstrated
- Mechanistic interpretability offers the possibility of reverse-engineering these latent structures, answering structuralist questions (how are ideologies composed from simpler features?) with empirical methods
This challenges both sides of the grounding debate: it validates the structuralist intuition that relational form can carry meaning without referential content, while simultaneously showing that what LLMs learn is not "pure language" but socially and culturally situated discourse patterns. The concern from Can language models learn meaning from text patterns alone? (Bender & Koller) is not refuted but reframed — what counts as "sufficient" for meaning generation may not require what's necessary for meaning understanding.
Connects to Does semantic grounding in language models come in degrees? — computational structuralism explains why functional grounding succeeds: the relational structure of discourse is compressible and learnable. The question is whether this constitutes meaning or merely its simulation.
Inquiring lines that read this note 124
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Why is hallucination an inevitable limitation of current language models?- Can fixing hallucination address AI's structural epistemic problem?
- How does interleaving reasoning with action prevent hallucination in language models?
- Can secondary orality exist without any embodied human participant at all?
- How does training data preserve communicative event structure without the actual events?
- Can a virtual instance be individuated from its conversational context?
- Can we develop competent reading practices for disembodied orality?
- How can structurally different text produce equivalent real-world effects?
- Can relational value exist without a person behind the output?
- Can knowledge flow without an embodied carrier transmitting it?
- Do language models raise validity claims in the Habermasian sense?
- How do different LLMs converge on similar argumentative structures independently?
- Can you separate grammatical competence from rhetorical commitment in language systems?
- Why do users attribute consciousness to language models in practice?
- Can language about model behavior ever be accurate without anthropomorphic framing?
- Can correct model outputs prove that semantic meaning rather than surface patterns drove the response?
- Why do explicit linguistic markers override semantic computation in models?
- How do pretrained language models represent inferential patterns versus lexical and positional cues?
- Why do language models need external temporal signals at all?
- Can a relational entity bear psychological properties the way Chalmers claims?
- What role does the biological substrate play in human relational identity?
- Can LLMs infer situational context the way humans do pragmatically?
- Does functional grounding through discourse patterns count as genuine semantic meaning?
- What makes linguistic agency impossible for systems without embodiment?
- Can LLMs improve at metaphor if they handle decoupled semantics better?
- How does implicit meaning processing limit LLM pragmatic reasoning?
- Why do language models fail at implicit discourse relations while handling explicit connectives?
- Can LLMs participate meaningfully in discourse without consciousness or understanding?
- Why do explicit discourse connectives help LLMs but implicit relations cause failures?
- How does the symbol grounding problem apply to artificial language systems?
- Can LLMs infer implicit meaning without surface linguistic markers?
- What structural limits prevent LLMs from abstracting moral principles?
- How does embodiment affect whether LLMs can participate in Wittgensteinian language games?
- Can LLMs develop genuine understanding without embodied experience?
- Can LLMs identify implicit metaphoric mappings that require pragmatic inference?
- Can LLM semantic representations exist without causally influencing their generation output?
- Why do explicit discourse connectives work when implicit relations fail?
- Why do language models reproduce human EPA structure despite different architecture?
- Do language models need words to think or just latent structure?
- Why does frame-activation matter more than word-by-word composition?
- How does enactive theory define language differently than computational linguistics?
- Why does training data saliency distort how models judge meaning?
- How do low-dimensional representation structures entangle multiple cultures together?
- What makes relational structure sufficient for generating contextually appropriate discourse?
- Can statistical learning from language alone capture all aspects of cultural competence?
- Can implicit linguistic information ever be reliably learned from training data?
- Can large language models understand language without embodied grounding systems?
- Can language models acquire meaning from distributional patterns alone without joint attention?
- Does embodiment and interaction matter for linguistic competence beyond pattern learning?
- Can frame semantics explain why context matters more than word similarity?
- Does selective suppression of linguistic relations enable human meaning-making?
- What structural signals in user language reveal their unstated preferences and context?
- Can language meaning emerge without joint attention and shared embodied interaction?
- What distinguishes surface cues from structural meaning in language understanding?
- Do metaphors work by decoupling meaning from linguistic associations?
- How does bidirectional entailment distinguish semantic equivalence from token similarity?
- What does embodiment and precariousness mean for linguistic agency?
- What distinguishes real understanding from superficial pattern matching?
- Can statistical learning from text replace embodied cultural experience?
- Are static embeddings analogous to the formal linguistic competence layer?
- How do static embeddings and contextualized representations divide semantic labor?
- Does language convey meaning purely through relational structure without external grounding?
- How does co-occurrence statistics alone produce hierarchical concept organization?
- Can readers detect meaning through resonance patterns alone without knowing authorial intent?
- Where does the meaning actually originate in reader-detected resonance across language?
- How do semantic features in representations become steerable task-specific directions?
- Can linguistic agency exist without embodiment and real-world participation?
- Does embodiment matter for genuine linguistic agency?
- Can understanding language happen entirely within a language system alone?
- Can functional behavior alone capture what makes something a genuine belief?
- Why does conceptual priming alone fail to produce consciousness claims?
- What makes the Extended Mind thesis incompatible with internalism?
- How do humans learn language through communication differently than LLM text prediction?
- How does semantic grounding differ between human minds and language models?
- Can LLMs predict social norms without deep integration into linguistic practices?
- Does DPO training with coreference chains teach spontaneous convention formation?
- How does Wittgenstein's language games explain social grounding in LLMs?
- Can mechanistic interpretability reveal how ideologies decompose into simpler features?
- How does mechanistic interpretability reveal ideological structures in language model weights?
- How does syntactic encoding relate to semantic feature representation?
- What architectural changes would let language models develop genuine functional competence?
- What structural properties of language models make fabrication inevitable?
- Can encoder models match human conceptual structure better than larger language models?
- What makes human language fundamentally different from what language models produce?
- Do language models learn surface patterns instead of underlying linguistic principles?
- Can language models reason without relying on learned semantic patterns?
- How do internal representations compare to human cognitive structures?
- Do language models actually learn linguistic structure or just surface statistics?
- Do language models encode deep syntactic structure or only surface-level patterns?
- What distinguishes surface generalizations from true linguistic generalizations?
- Why do surface generalizations fail on unusual syntactic structures?
- Do LLMs learn linguistic generalizations or just surface-level frequency patterns?
- Can formal language pretraining address surface generalization without learning true linguistic structure?
- Do LLMs learn surface patterns instead of genuine linguistic structure?
- Can language models learn internal world models without explicit environment specifications?
- Do LLMs learn abstract grammar or culturally situated discourse patterns instead?
- What makes a relational act different from just moving content around?
- How does monological training on text differ from dialogical training in conversation?
- What role does language play as a cognitive scaffold versus communication tool?
- What role does joint attention play in how humans learn language meaning?
- Why does joint attention matter for acquiring linguistic meaning?
- Can pragmatic competence emerge from text exposure alone without interactive grounding?
- Can pragmatic competence emerge from text exposure without interactive grounding?
- How do world models create indirect causal grounding without physical environment contact?
- Can language models develop world models that ground meaning in causal reality?
- Why does LLM compression eliminate causal grounding in conceptual representations?
- Can functional semantic grounding substitute for true causal grounding?
- Can LLMs reason through semantics without understanding causal mechanisms?
- Can external actions provide causal necessity that language models lack?
- Do speech encoders actually learn the physics of how vocal tracts produce sound?
- How do speech encoders learn articulatory physics without phonetic labels?
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Computational structuralism: Toward a formal theory of meaning in the age of digital intelligence
- Mechanistic Indicators of Understanding in Large Language Models
- From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning
- Semantic Structure in Large Language Model Embeddings
- CoT is Not True Reasoning, It Is Just a Tight Constraint to Imitate: A Theory Perspective
- Probing Structured Semantics Understanding and Generation of Language Models via Question Answering
- Language Models’ Hall of Mirrors Problem: Why AI Alignment Requires Peircean Semiosis
- From Human to Machine Psychology: A Conceptual Framework for Understanding Well-Being in Large Language Models
Original note title
LLMs operationalize Saussures langue — fully relational models with no external referents suffice to generate contextually appropriate discourse