SYNTHESIS NOTE
Topics›Sentiment Semantics Toxic Detections›this note

Can reasoning happen at the sentence level instead of tokens?

Does moving from token-level to sentence-level reasoning in embedding space preserve the capability for complex reasoning while enabling language-agnostic processing? This challenges assumptions about how LLMs must operate.

Synthesis note · 2026-02-23 · sourced from Sentiment Semantics Toxic Detections

Current LLMs operate at the token level — every reasoning step is a next-token prediction. Meta's Large Concept Model (LCM) challenges this by operating at the sentence level, reasoning in an abstract embedding space (SONAR) where each "concept" corresponds to a sentence.

The architectural difference is fundamental. The LCM:

The hierarchical structure adds a planning layer. The LCM predicts a sequence of concepts auto-regressively until it produces a "break concept" — analogous to a paragraph break. At that point, a Large Planning Model (LPM) generates a plan to condition the LCM for the next sequence. This two-level architecture (sentence-level prediction + paragraph-level planning) is designed to produce more coherent long-form output than flat token-level generation.

The comparison to JEPA (LeCun, 2022) is instructive: both predict representations in embedding space rather than raw observations. But where JEPA emphasizes learning the representation space via self-supervision, LCM focuses on accurate prediction within an existing embedding space (SONAR). The embedding quality is assumed, not learned end-to-end.

This connects to the latent reasoning thread through a different mechanism. Can models reason without generating visible thinking tokens? achieves reasoning without tokens via recurrent depth in continuous space. LCM achieves it via sentence-level embeddings. Both challenge the assumption that verbalized token-by-token generation is necessary for reasoning, but from different angles: depth-recurrent models reason within a single token's representation; LCM reasons between sentence-level units.

The practical implication: if reasoning can happen at the concept level rather than the token level, then the verbalized chain-of-thought paradigm is not the only path to sophisticated reasoning. The question is whether sentence-level granularity captures enough structure for complex reasoning tasks, or whether some tasks require finer-grained (sub-sentence) reasoning steps.

Inquiring lines that read this note 56

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What causes reasoning models to fail or wander off track? Why do embedding systems fail to capture task-relevant relationships? Does chain-of-thought reasoning reveal genuine computation or imitate patterns? Can models improve accuracy without degrading reasoning quality? What structural properties of attention create systematic model biases? Can reasoning scale in latent space without tokens? How does policy entropy collapse constrain scaling of reasoning-focused RL? What enables genuine semantic understanding in language models? How effectively can language models perform reasoning, especially combined with symbolic methods? Why can recurrent transformers achieve reasoning capabilities that standard transformers cannot? Do language models reason through causal mechanisms or semantic associations? Why do token-level mechanisms matter for learning to reason? Do reasoning traces faithfully reflect actual model reasoning? What explains language models' asymmetric difficulty with implicit versus explicit linguistic relations? What reasoning architectures enable models to solve complex problems efficiently? How should retrieval systems handle complex multi-step reasoning? Does encoded knowledge in language models actually influence their outputs? How do neural networks achieve compositional generalization at scale? How do soft reasoning mechanisms explore multiple paths without explicit training?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 138 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

Large Concept Models enable sentence-level reasoning in a language-agnostic embedding space — hierarchical abstraction beyond token-level processing