Capability benchmark · Chemistry · Materials science · General science
PHREEQC-MCQ-200 diagnoses scientific simulator tool-use agents
A 200-question benchmark for agents operating deterministic aqueous-geochemistry simulations.
Summary
The preprint introduces PHREEQC-MCQ-200, a diagnostic benchmark for tool-augmented agents using the PHREEQC aqueous-geochemistry simulator. The benchmark tests whether agents can build simulator inputs, run deterministic scientific software, interpret structured outputs, and commit to final answers.
AI role
Agents construct PHREEQC inputs, execute the simulator, inspect structured outputs, and answer grounded computation questions.
Narrative role
This adds a physical-science computation benchmark where tool access is part of realistic research work, and where evaluation looks beyond aggregate accuracy to item retention, protocol sensitivity, and failure location.
Caveat
The task is multiple choice and centered on one simulator, so results should not be read as unconstrained geochemistry modeling ability.