Capability benchmark · Computer science

RE-Bench compares AI agents with humans on AI R&D tasks

AI-agent and human performance compared on open-ended AI R&D environments.

Summary

The arXiv paper presents RE-Bench, seven open-ended machine-learning research-engineering environments with scoring functions and human baseline data. It compares frontier AI agents with human experts under time limits and reports mixed performance across tasks.

AI role

Attempted machine-learning research-engineering tasks under time limits and scored environments.

Narrative role

RE-Bench is important because it compares agents with humans on R&D-like work rather than only static tasks.

Caveat

The task suite is small and AI R&D-specific, so it should not be overgeneralized to science as a whole.