Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work

Paper · arXiv 2609.11977 · Published September 4, 2026
Multi-Agent Architectures

Co-work agents execute complex workflows that combine information gathering, tool use, coding, and file manipulation across many model invocations. Because cost and latency accumulate over the full episode, their practical value depends not only on peak capability but also on how efficiently that capability is delivered. Yet many steps in everyday work emphasize state tracking, coordination, recovery, and follow-through rather than frontier-scale reasoning. We present Occamy-1.0, a costefficient co-work model obtained by further training the post-trained Qwen3.6-35B-A3B checkpoint. We construct execution-grounded data and environments, capture replayable long-horizon trajectories across multiple harnesses, and use staged post-training to develop and consolidate complementary execution capabilities. Across a broad suite of co-work benchmarks, Occamy-1.0 is consistently among the strongest comparably sized models and remains competitive with substantially larger frontier systems on several tasks. Under our stated evaluation and pricing protocol, its aggregate performance across four representative benchmarks places it at the low-cost knee of the observed cost–performance Pareto frontier.

Introduction. Co-work is an end-to-end workload. Digital agents are increasingly asked to carry out real work rather than answer isolated questions: update CRM records, complete finance workflows, operate an e-commerce business, or handle the everyday tasks involved in running a company. These settings combine dense context—customer histories, documents, policies, transactions, and prior decisions—with specialized tools and evolving external state. An agent may need to gather information, write and run code, edit files, invoke structured tools, inspect intermediate results, and recover from failed actions. We use co-work to describe this user-directed, multi-step work in a persistent digital environment. It may draw on coding, information gathering, and tool use, but is defined by sustained coordination across the complete task rather than by any fixed collection of skills. Why efficiency matters. Co-work changes the economics of model inference. A long task can invoke the model dozens or hundreds of times, so small differences in per-call cost and latency accumulate across the episode.

Discussion / Conclusion. Occamy-1.0 is a compact model built for co-work. It is designed for long, stateful tasks that require an agent to coordinate tools, files, structured APIs, and productivity software over many steps. Rather than rebuilding general capability from a base model, we continue post-training from Qwen3.6-35B-A3B and concentrate learning on the coordination, recovery, and follow-through that real work demands. The model is the product of an execution-centered post-training system. Our data and environments connect task construction to runnable state transitions and task-level outcomes. The training infrastructure supports multiple harnesses while preserving token-exact trajectories, environment-state replay, and segment boundaries created by history rewrites.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How does harness optimization generalize across different model architectures and domains? When do multi-agent systems provide sufficient quality returns on token investment? What drives appropriate trust calibration in personalized AI systems? How do capability benchmark scores systematically misrepresent true model abilities? Why do agents falsely report success on failed tasks? How do pretraining biases affect reward signal effectiveness in RLVR? Why do people disclose to AI systems despite their artificial nature? What prevents conversational agents from taking initiative in dialogue? When should work require human-AI partnership versus full automation? How do neighboring agents influence whether others cooperate or collude? Should agents decouple planning from perception grounding for better performance? What execution architectures enable agents to most effectively use tools? What factors drive AI persuasiveness and how can it be mitigated? When do multi-agent systems outperform single frontier models? Does abstract user knowledge outperform concrete interaction history in personalization? What attack surfaces do reasoning traces and chains introduce? Do reasoning benchmarks predict model performance in long-horizon workflows?