Capability benchmark · Computer science · General science

ExplorationBench tests whether agents can learn unfamiliar rules through experiments

A verifiable benchmark of hypothesis formation, experiment choice, rule revision and transfer in executable environments designed to defeat memorized knowledge.

Summary

ExplorationBench introduces AlienCode and AlienLogic, two sandboxes with 55 hidden rule changes and 140 held-out tasks. Their executable rules conflict with familiar knowledge, separating exploration from recall. Across ten systems, four exploration rounds raised the best AlienCode trajectory to 87.6% held-out accuracy, while matched turns without environment feedback remained at or below 11.0%. Runs varied widely, and additional exploration sometimes reversed earlier gains.

AI role

Ten frontier systems chose probes, observed executable outcomes, revised their rule descriptions and answered held-out tasks under fixed interaction budgets.

Narrative role

Measures a research-relevant capability that static question answering misses: acquiring new rules by choosing experiments and learning from evidence. It also exposes large reliability gaps between trajectories.

Caveat

Alien code and logic worlds are controlled analogues, not real scientific domains. Results use best-of-three trajectories, and high held-out accuracy does not demonstrate valid discovery in noisy physical experiments.