Capability benchmark · Climate · Physics · Computer science
SciExam finds six agent-built ENSO models exceed a published reference on hidden tests
Executable low-order stochastic climate models evaluated on statistics, variable reconstruction and held-out ENSO forecasting.
Summary
Six of twelve final agent-produced models exceed the published reference’s composite score. The best scores 0.515 versus 0.378; the reference remains strongest on statistical fidelity. Agents receive 1980–2014 observations, while forecast grading uses 2015–2024. Hidden graders do not provide feedback during the run.
AI role
Agents process observations, freeze self-written diagnostics and develop coupled stochastic models within six eligible working hours.
Narrative role
This measures open-ended physical model construction against empirical observations rather than a known answer or language-model judge.
Caveat
Each system runs once, so scores characterize these outputs, not expected model performance. Systems differ in scaffold/settings. The reference was not optimized for the composite; mechanism compatibility does not settle the ENSO debate.