Capability benchmark · Computer science · General science

CausalGame tests whether AI scientist agents can reason causally

An interactive benchmark for agent experiment design under selection bias, measurement error, and hidden confounders.

Summary

CausalGame evaluates LLM agents in interactive experimental settings where naive correlations can mislead. Across 30 agents, the paper reports weak causal understanding despite some successful outcomes, with best survival below analytical optima and low causal-reasoning rubric credit.

AI role

Agents design experimental protocols, collect observations, infer causal structure, and produce final decisions with explanations.

Narrative role

This benchmark clarifies a capability gap for AI scientist narratives: discovery agents may need causal experiment design, not just retrieval, coding, or trial-and-error optimization.

Caveat

The environment uses game abstractions rather than live scientific systems, and rubric-based explanation scoring adds evaluation-design assumptions.