Policymakers and biosecurity experts face a growing challenge: how to generate and interpret credible evidence about the biological capabilities and risks of AI systems that can autonomously perform multi-step scientific tasks. A new paper on arXiv, titled "Measuring Biological Capabilities and Risks of AI Agents," tackles this issue head-on by providing a framework for understanding and conducting evaluations of these so-called AI scientists.
The paper, authored by Paskov, Patricia, Lee, Jeffrey, Brady, Kyle, and Worland Alyssa, addresses a rapidly emerging policy challenge. As agentic AI systems enter real research workflows, decision-makers increasingly encounter evaluation results whose meaning depends on underlying design choices that are often implicit or under-documented. The authors synthesize current evidence on AI-enabled biological risks and introduce "biological agentic evaluations" as a promising but interpretation-sensitive tool for assessing these systems.
Central Contribution: Practical Considerations
The paper's central contribution is a set of practical, experience-grounded considerations drawn from the authors' own evaluations. These considerations show how choices around defining, designing, running, scoring, and documenting evaluations materially shape what results do and do not imply about risk.
- Defining evaluations: The scope of the evaluation—what biological capabilities or risks are being measured—must be clearly articulated. The choice of tasks, from benign research to potentially harmful outcomes, directly affects the interpretation of results.
- Designing evaluations: The design includes selecting appropriate benchmarks, controlling for confounding variables, and ensuring reproducibility. Without rigorous design, the evaluation may produce misleading implications.
- Running evaluations: Practical aspects such as compute constraints, access to models, and safety protocols during testing influence the reliability of outcomes.
- Scoring evaluations: The metrics used to score performance must align with the risk model being assessed. For instance, success on a narrow synthetic task may not translate to real-world biological capability.
- Documenting evaluations: Transparent documentation of all choices is crucial for others to interpret and reproduce results. The authors emphasize that evaluation documentation must include the rationale behind design decisions.
Target Audiences
The analysis is intended to help multiple stakeholders:
- Policymakers: To interpret biological evaluation outputs with appropriate caution and understand the limitations of the evidence.
- Funders: Public and private funders are guided toward high-leverage investments in AI-biology evaluation research.
- Biosecurity practitioners: Those assessing emerging AI systems can use the considerations to evaluate the credibility of claims.
- Researchers: A secondary audience includes researchers designing or conducting agentic evaluations within frontier AI labs, AI providers, scientific institutions, and third-party evaluation organizations.
Implications for Enterprise Technology Leaders
Although the paper focuses on biosecurity, the framework has broader relevance for any organization deploying autonomous AI systems in high-stakes domains. The core lesson—that evaluation design choices determine what results do and do not imply about risk—applies equally to supply chain AI agents, trade finance automation, and logistics decision systems. Enterprise technology decision-makers should ensure that any AI evaluation they rely on is transparent about its defining, design, execution, scoring, and documentation choices to avoid misinterpretation of capabilities and risks.