SYNTHESIS NOTE
Topics›Knowledge Graphs›this note

Can knowledge graphs teach models deep domain expertise?

Explores whether organizing knowledge as structured graph paths, composed from simple to complex, can enable language models to develop genuine domain superintelligence rather than surface-level pattern matching.

Synthesis note · 2026-02-23 · sourced from Knowledge Graphs

Language models acquire general abstractions through top-down self-supervised learning on vast corpora, but this approach captures surface-level regularities rather than deep domain expertise. Bottom-up curriculum learning from knowledge graphs offers an alternative: KG paths naturally encode compositional reasoning chains where atomic triples (e.g., "Methane Contains Element Carbon") compose into multi-hop paths that build toward higher-order understanding (e.g., methane's bonding structure through C-H bonds → sigma bonds → single covalent bonds).

The pipeline synthesizes 24,000 reasoning tasks from a medical KG, paired with structured thinking traces derived from diverse medical primitives. Fine-tuning QwQ-32B on this curriculum produces QwQ-Med-3, which significantly outperforms state-of-the-art open-source and proprietary reasoning models across 15 medical domains on the ICD-Bench evaluation suite.

The key architectural insight: KG topology naturally induces the bottom-up curriculum — beginning with atomic relations and composing them into increasingly complex reasoning chains. This mirrors how human students build expertise through pedagogical structure (foundational → advanced chapters), not encyclopedic browsing. Previous neuro-symbolic and probabilistic graph inference approaches attempted similar hierarchical reasoning from primitives but failed to generalize beyond synthetic regimes; LMs provide the generalization capability that symbolic systems lacked.

The broader implication challenges the AGI-as-breadth paradigm: domain-specific superintelligence may be achievable through relatively small models (32B) fine-tuned on structured domain knowledge, composing into broader intelligence through interacting specialist agents — analogous to how human society acquires expertise through collaborative specialization.

This connects to:

Inquiring lines that read this note 52

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What reasoning architectures enable models to solve complex problems efficiently? Why is hallucination an inevitable limitation of current language models? How much do training data properties shape model reasoning? What capability trade-offs arise from domain specialization through fine-tuning? How can humans maintain meaningful oversight as AI systems become increasingly autonomous and complex? Do knowledge graphs offer advantages over embeddings for multi-hop retrieval? How should retrieval systems handle complex multi-step reasoning? Why can't prompting alone inject genuinely new knowledge into models? When do semantic similarity approaches miss structural retrieval failures? Can models improve accuracy without degrading reasoning quality? What compositional reasoning failures limit large language models despite scale? Why does adding new knowledge through fine-tuning degrade existing capabilities? Can prompt-based context override biases that were embedded during pretraining? Is language model reasoning authentic and what causes models to reason? How do standardized protocols improve multi-agent coordination and reliability? How should systems decide whether to retrieve or reason alone? Do language models learn genuine understanding or just surface patterns? How does decomposing tasks improve reasoning and prevent failure propagation? Why do token-level mechanisms matter for learning to reason? Why do embedding systems fail to capture task-relevant relationships? What role does sparsity play in model behavior and scaling decisions? What trajectory-level metrics beyond task success best evaluate agent performance? Is reasoning capability latent in base models or created by post-training? What do systematic disagreements between annotators reveal about ground truth?

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

Knowledge graph curriculum enables bottom-up domain superintelligence by composing primitives into complex reasoning chains