Limitation or failure · General science · Computer science
EurekaBench finds research agents match prediction more closely than scientific insight
A 26-task, six-domain experiment-loop benchmark separates scientific constraints, held-out prediction and 306 expert-defined insights.
Summary
EurekaBench evaluates mechanisms on scientific validity, predictions and useful insights. Its best conditioned agent score is 29.8%, versus 63.7% for published human mechanisms. The two highest predictive scores are 47.4%, near the human 48.8%, while their insight scores are 42.4% and 29.4%, versus 69.7%. Most agents violate at least one scientific constraint on more than half the tasks.
AI role
Seven model-agent configurations conducted experiments and proposed executable mechanisms under four-hour, one-H100 budgets.
Narrative role
Directly tests the distinction between fitting observations and producing scientifically useful explanations, an important boundary for autonomous-discovery claims.
Caveat
Preprint with only 26 selected problems. Two LLM judges grade constraints and insight derivability. Human baselines are existing research outcomes with unmatched time and compute; blocked source retrieval does not eliminate pretraining contamination.