Capability benchmark · Computer science
PaperBench measures whether agents can replicate AI research papers
PaperBench measures whether agents can replicate AI research papers: capability signal for AI systems on research-adjacent tasks.
Summary
The arXiv paper introduces PaperBench, a benchmark in which agents attempt to replicate 20 ICML 2024 papers from scratch. The benchmark decomposes replication into thousands of graded tasks covering paper understanding, coding, experimentation and result reproduction.
AI role
AI systems are tested on research-adjacent capabilities relevant to computer science.
Narrative role
This is supporting evidence for whether AI systems can perform research-adjacent tasks needed before stronger discovery or acceleration claims.
Caveat
Benchmark, model, or tool performance is an upstream capability indicator, not proof of new scientific discovery.