Limitation or failure · Materials science · Computer science

CompMat-Bench finds scientific errors persist despite high guided materials-task success

Ninety-four reproduced computational-materials research steps with fixed-rule grading and an audit of scientific versus software failures.

Summary

CompMat-Bench grades steps drawn from published materials studies without rerunning expensive simulations. GPT-5.6 Sol completes 85/94 tasks in MatClaw and 86/94 in Codex CLI under full guidance. Most failures are scientific rather than software errors. In one CrN workflow, nine of fifteen failed trials continue after visibly anomalous Wannier outputs, producing exchange constants deviating by up to 116%.

AI role

Agents prepared simulation inputs and analyzed cached outputs under full or reduced methodological guidance, individually and in workflows.

Narrative role

Complements MatToolBench’s software-operation failures with a distinct measurement of scientific judgment: agents can execute a procedure while overlooking diagnostic evidence that makes its output unreliable.

Caveat

Selected research steps exclude new conceptual proposals and largely reuse precomputed simulations. Full-guidance single tasks are near saturation for stronger agents; workflow samples are small, fixed graders can reject valid answers, and about 1 TB of reproduction intermediates is unreleased.

Related events