Can we measure how much risk open models actually add?
Whether current evidence adequately quantifies the marginal misuse risk of openly released foundation models compared to existing technology. This matters because policy decisions depend on knowing if open release meaningfully worsens real-world harm vectors.
The open-vs-closed release debate is heated and under-evidenced. This position paper clarifies it by defining open foundation models (broadly available weights — Llama 2, Stable Diffusion XL) via five distinctive properties (greater customizability, deeper inspectability, poor monitoring, etc.) that drive both their benefits (innovation, competition, distributed decision-making power, transparency) and risks. Its analytical contribution is a marginal-risk framework: assess misuse not in absolute terms but relative to pre-existing technology (search engines, prior models). Applying it across vectors (cyberattacks, bioweapons, disinformation), it finds current research insufficient to characterize the marginal risk — and shows that past disagreements stem from focusing on different parts of the framework under different assumptions.
The keeper is the marginal reframing: the policy question is not "could an open model help a bad actor?" but "how much does it help beyond what they could already do?" — and on that question the evidence is mostly missing, which is itself the finding.
This is a discourse/governance anchor for the vault. It complements the empirical risk register of Where do frontier AI models actually pose the greatest risk today? — both insist on measured marginal risk over speculation — and informs the open-weights side of the alignment-and-society conversation.
Inquiring lines that read this note 10
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What capability trade-offs arise from domain specialization through fine-tuning?- What distinctive properties make open foundation models different from closed ones?
- What benefits do open foundation models create that closed systems cannot?
- Do open model properties like customizability create net new misuse opportunities?
- How do we measure marginal risk instead of speculating about misuse scenarios?
- Can we empirically test whether open models lower barriers to harmful workflows?
- What information should governments disclose when issuing model suspension directives?
- Why do researchers disagree on open model risks despite same evidence?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Where do frontier AI models actually pose the greatest risk today?
Current AI safety discourse focuses on autonomous R&D and self-replication, but empirical risk assessment may reveal a different priority. Where should mitigation efforts concentrate?
both demand measured marginal risk over speculative misuse narratives
-
How soon do AI researchers expect artificial general intelligence?
A survey of 2,778 AI researchers reveals how expert timelines for human-level AI have shifted over the past year, and what factors drive disagreement among specialists on this critical timeline.
the broader risk-discourse context where open-model debates sit
-
Does measuring exploit capability help or harm defense?
Exploitation benchmarks can support defenders and attackers equally. How should we evaluate capabilities with unavoidable dual-use potential, and what safeguards make evaluation itself defensible?
a benchmark's own statement of a marginal-risk claim for the cyber vector ("lowering the barrier for offense"); the ExploitGym excerpt reports no results, so it is a measurement design, not evidence about uplift
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- On the Societal Impact of Open Foundation Models
- A Causal Model for Locating and Unlocking Sandbagging in Model Organisms
- Towards Faithfully Interpretable NLP Systems: How should we define and evaluate faithfulness?
- Foundation Priors
- What Has a Foundation Model Found? Using Inductive Bias to Probe for World Models
- Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report
- The Return of Pseudosciences in Artificial Intelligence: Have Machine Learning and Deep Learning Forgotten Lessons from Statistics and History?
- TrustLLM: Trustworthiness in Large Language Models
Original note title
open foundation models need a marginal-risk framework because current evidence cannot characterize their misuse risk relative to existing technology