Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence
AI models are increasingly used in scientific discovery and human decision-making. Yet how AI models work and what risks they pose remain poorly understood. As AI development becomes faster and more automated, research on the mechanisms underlying AI remains largely manual, widening the gap between model capabilities and our ability to understand and control them. To bridge this gap, we introduce Mechanist, an agentic system that uses AI as a scientific instrument for the autonomous discovery of mechanisms underlying AI. To ground novel mechanism hypotheses, we construct a scientific knowledge graph of 13,000 studies on AI mechanisms, alongside a multidisciplinary database of 43 million papers spanning 26 fields. For reliable experiment execution, we curate a library of 32 foundational methods for mechanism analysis, causal intervention, and validation. Compared with Claude Code and existing AI-scientist systems, Mechanist generates higher-quality mechanism hypotheses and executes experiments more reliably. Across four case studies, Mechanist autonomously discovers new model behaviors and their underlying mechanisms, and translates these discoveries into mechanism-guided interventions and interdisciplinary design.
Introduction. Artificial Intelligence (AI) models are rapidly evolving from tools for daily chat into intelligent systems that guide problem formulation, experimental design, and human decision-making [1, 2, 3]. The growing success of AI models raises a fundamental question [4, 5, 6]: how do these AI models acquire knowledge about the world, form beliefs 7, reason, and act? It also remains unclear whether AI models harbour latent risks or exhibit unreliable behaviors in critical domains such as healthcare, finance, and chemical manufacturing [9, 10, 11]. More critically, AI development is accelerating and becoming increasingly automated, outpacing progress in understanding and controlling the mechanisms underlying AI[1, 4]. Discovering the mechanisms of model intelligence can reveal how models operate internally, elucidate the principles underlying their behavior, identify risks early, and enable adaptive control over and targeted improvements to model behavior [9, 10, 11]. Despite its importance, mechanistic understanding of AI remains difficult to obtain and hard to scale [10, 12, 13].
Discussion / Conclusion. We introduce Mechanist, a scientific instrument for the autonomous investigation of mechanisms underlying AI models. To support Mechanist, we design a large-scale knowledge graph of interpretability and a library of 32 foundational methods of mechanistic analysis. These resources enable Mechanist to formulate high-quality hypotheses, execute experiments reliably, establish robust causal evidence, and iteratively refine mechanistic explanations. Mechanist can uncover previously unrecognized risks in models deployed in scientific laboratory settings, reveal how AI models represent world knowledge, and improve model capabilities across computer science and other scientific domains. Then, we discuss related topics in the following section. AI for science.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Can brute-force automated research substitute for iterative depth and human research intuition?- How does automated mechanism discovery compare to human-led mechanistic research?
- Can bilevel autoresearch discover new search mechanisms for the inner research loop?
- How should research governance adapt to structural verification delays?
- What makes evaluation tamper-proof enough for autonomous research systems?
- Can a single dominant mechanism replace the combined effect of all five?
- Do interaction effects between research mechanisms depend on the task domain?
- What distinguishes research stages where the combined stack remains reliable?
- How does bilevel autoresearch balance outer loop cost against discovery improvements?
- How does multi-agent debate prevent degeneration from self-revision loops?
- Can autonomous teams sustain multiple competing hypotheses simultaneously?
- What five ecosystem conditions must coordination governance and evidence actually satisfy?
- Why does decentralization work better than central planning for open-ended research?
- What causes prolonged concentration on single approaches in decentralized research teams?
- What genuine cultural forms does AI homogeneity actually displace?
- Why does volume alone fail to explain the damage AI does to epistemic systems?