Capability benchmark · Computer science

FIRE-Bench measures full-cycle rediscovery of scientific insights

A benchmark where agents rediscover verified empirical findings from recent machine-learning research papers.

Summary

FIRE-Bench turns recent machine-learning empirical analysis papers into constrained rediscovery tasks. Agents are asked to independently design and run experiments from a high-level question, then their conclusions are compared with established findings.

AI role

Agents receive high-level research questions, design experiments, implement code, execute plans, and derive evidence-backed conclusions.

Narrative role

This fills a benchmark gap between paper replication and open-ended AI scientist claims by testing whether agents can reconstruct empirical insight without being handed the original methods and conclusions.

Caveat

The benchmark is limited to machine-learning research papers, so it measures one form of scientific rediscovery rather than wet-lab or field-science discovery.