Limitation or failure · Computer science · General science

Shadow evaluations find agents miss the core of open-ended AI research

Two expert-graded shadow evaluations of frontier agents pursuing the central research questions of unpublished NeurIPS submissions.

Summary

Kirgis et al. introduce expert-graded shadow evaluations in which frontier agents pursued the core questions of two unpublished NeurIPS 2026 papers. The agents completed research-engineering work without human help, but the original authors judged neither to have made substantial progress; a second model and scaffold reproduced the reported failure pattern.

AI role

Frontier agents autonomously performed research engineering while attempting to formulate and execute open-ended machine-learning research projects.

Narrative role

This adds unusually direct, artifact-backed negative evidence about the open-ended research bottleneck: agents can execute engineering steps yet fail at the judgment, reframing, and recovery needed to advance a live research question.

Caveat

The evidence is a two-case preprint study in AI research, not a broad evaluation across scientific domains; expert assessment of unpublished work remains necessarily contextual.

Related events