Capability benchmark · Climate · Physics · Computer science

SciExam finds six agent-built ENSO models exceed a published reference on hidden tests

Executable low-order stochastic climate models evaluated on statistics, variable reconstruction and held-out ENSO forecasting.

Summary

Six of twelve final agent-produced models exceed the published reference’s composite score. The best scores 0.515 versus 0.378; the reference remains strongest on statistical fidelity. Agents receive 1980–2014 observations, while forecast grading uses 2015–2024. Hidden graders do not provide feedback during the run.

AI role

Agents process observations, freeze self-written diagnostics and develop coupled stochastic models within six eligible working hours.

Narrative role

This measures open-ended physical model construction against empirical observations rather than a known answer or language-model judge.

Caveat

Each system runs once, so scores characterize these outputs, not expected model performance. Systems differ in scaffold/settings. The reference was not optimized for the composite; mechanism compatibility does not settle the ENSO debate.