Capability benchmark · Computer science
RE-Bench compares AI agents with humans on AI R&D tasks
AI-agent and human performance compared on open-ended AI R&D environments.
Summary
The arXiv paper presents RE-Bench, seven open-ended machine-learning research-engineering environments with scoring functions and human baseline data. It compares frontier AI agents with human experts under time limits and reports mixed performance across tasks.
AI role
Attempted machine-learning research-engineering tasks under time limits and scored environments.
Narrative role
RE-Bench is important because it compares agents with humans on R&D-like work rather than only static tasks.
Caveat
The task suite is small and AI R&D-specific, so it should not be overgeneralized to science as a whole.