Capability benchmark · Computer science

PaperBench measures whether agents can replicate AI research papers

PaperBench measures whether agents can replicate AI research papers: capability signal for AI systems on research-adjacent tasks.

Summary

The arXiv paper introduces PaperBench, a benchmark in which agents attempt to replicate 20 ICML 2024 papers from scratch. The benchmark decomposes replication into thousands of graded tasks covering paper understanding, coding, experimentation and result reproduction.

AI role

AI systems are tested on research-adjacent capabilities relevant to computer science.

Narrative role

This is supporting evidence for whether AI systems can perform research-adjacent tasks needed before stronger discovery or acceleration claims.

Caveat

Benchmark, model, or tool performance is an upstream capability indicator, not proof of new scientific discovery.