Capability benchmark · Biology · Computer science
Paper2Agent evaluates executable paper reuse across computational research repositories
A peer-reviewed evaluation of agents that expose paper code and data as tested executable tools for reuse.
Summary
The journal evaluation of Paper2Agent, first released as a preprint in 2025, reports successful conversion of 74 of 100 computational-biology papers into agents. Of 599 proposed tools, 593 passed automated validation. On 300 tutorial-derived queries, the system reported 91.2% accuracy versus 80.3% for Claude Code with repository access using the same Sonnet 4 model. The study also tests open-ended questions and cross-paper analyses.
AI role
LLM agents configured environments, wrapped paper methods as MCP tools, tested outputs and answered scientific queries through those tools.
Narrative role
Tests a concrete barrier between publishing a computational method and reusing it. The expanded journal evaluation is a research-workflow capability result, not evidence that all papers can become reliable autonomous scientists.
Caveat
Incomplete repositories, missing data and environment failures prevented many conversions. Reference-answer agreement primarily measures faithful execution, not validity of open-ended conclusions. Researchers still select hypotheses and assess evidence; no independent rerun here.