Capability benchmark · Computer science · General science
CausalGame tests whether AI scientist agents can reason causally
An interactive benchmark for agent experiment design under selection bias, measurement error, and hidden confounders.
Summary
CausalGame evaluates LLM agents in interactive experimental settings where naive correlations can mislead. Across 30 agents, the paper reports weak causal understanding despite some successful outcomes, with best survival below analytical optima and low causal-reasoning rubric credit.
AI role
Agents design experimental protocols, collect observations, infer causal structure, and produce final decisions with explanations.
Narrative role
This benchmark clarifies a capability gap for AI scientist narratives: discovery agents may need causal experiment design, not just retrieval, coding, or trial-and-error optimization.
Caveat
The environment uses game abstractions rather than live scientific systems, and rubric-based explanation scoring adds evaluation-design assumptions.