SYNTHESIS NOTE
Topics›this note

How does test-time scaling work for individual research agents?

This explores whether the scaling laws that apply to reasoning tokens also apply to search steps and information retrieval in agentic systems. Understanding this could reveal new ways to improve AI research capabilities through compute allocation.

Synthesis note

Core Insights


description: Navigation hub for deep research scaling, agentic system architectures, knowledge graph reasoning, and search-as-TTS — individual agent-level test-time compute type: topic-map created: 2026-02-24 topics: ["How do you navigate synthesis across fragmented research topics?"]

deep research and agentic systems topic map

How does test-time scaling work at the individual agent level? This sub-map covers deep research scaling (where TTS law generalizes from reasoning tokens to search steps), agentic system architectures (data efficiency, skill libraries, adaptation paradigms), and knowledge graph reasoning (externalized reasoning, graph-structured training data synthesis).

The key insight: search budget follows the same scaling curve as reasoning tokens, making deep research a TTS problem. At the agent level, data efficiency is extreme — 78 curated demonstrations outperform 10K samples for agency.

Parent map: How does test-time scaling work at the agent level?

Deep Research and Search Scaling

(From Arxiv/Deep Research)

Proactive Search Evaluation (2026-05-28 — VibeSearchBench)

The evaluation-experience gap in search: benchmarks reward what users never struggle with, and realistic evaluation needs vague intent, multi-turn dialogue, and open-ended structure.

Production Deployment

Agentic Systems

(From Arxiv/Agents — agent architectures, team optimization, data efficiency for agency, adaptation paradigms)

Knowledge Graph Reasoning and Training Data Synthesis

(From Arxiv/Knowledge Graphs — graph-structured reasoning, KG externalization, and synthetic training data from KGs)

Autonomous Science and Ideation — Batch #3 backlog (2026-06-03)

Three papers on AI doing research: two architectures for long-horizon autonomy, and one reframing of why LLM ideation underwhelms. A fourth (2609.26457, two notes here, excerpt-only) adds a premise the three do not state: automating the artifacts a research agent produces leaves the efficiency of the research process itself fixed, so the agent's own code becomes the object of optimization. It gives the section a different diagnosis of what limits a research loop: ASI-Evolve puts the gap in insight transfer across iterations and closes it with a cognition base and an analyzer, while this premise makes the loop's own efficiency the unmoved quantity and the remedy a rewrite of the agent. The premise rests on a trend the paper relays, diminishing returns to R&D spending, and the scaling-law entry in What actually constrains AI systems from learning misalignment? argues research progress becomes computation-scalable, so the two frame the same axis with different quantities and neither excerpt tests the other. The bridge note gives the premise a unit, research efficiency as score under a fixed evaluation budget, and that unit differs from the spending the premise cites. The loop itself, its open questions and a filed tension are in What actually constrains AI systems from learning misalignment?.

AI-for-AI and Research Venues — Batch #3 wave 2 (2026-06-03)

Related Areas

New — 2026-06-27

Inquiring lines that read this note 42

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can intelligent routing over smaller models outperform scaling a single large model? How should test-time compute scaling work in agentic systems? How do neural networks achieve compositional generalization at scale? Can parallel reasoning outperform sequential reasoning under fixed token budgets? How should retrieval systems handle complex multi-step reasoning? Can inference-time compute effectively substitute for model scale? How should inference compute be allocated based on problem difficulty? Do knowledge graphs offer advantages over embeddings for multi-hop retrieval? When do multi-agent systems outperform single frontier models? Why do persona simulations fail to predict authentic user behavior? How do capability benchmark scores systematically misrepresent true model abilities? Can brute-force automated research substitute for iterative depth and human research intuition? Can single-point security defenses protect multi-agent systems from multi-step attacks? How do agent-learned skills transfer and improve across different tasks? Why do standard benchmarks fail to predict agent deployment success? When do multi-agent systems provide sufficient quality returns on token investment? How can evolutionary algorithms maintain diversity during solution search? What should agent evaluation prioritize to reveal reliable behavior? How does harness optimization generalize across different model architectures and domains?

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

deep research and agentic systems topic map