Capability benchmark · Materials science · Chemistry · Computer science
AutoXRD benchmarks agents on executable powder-diffraction workflows
XRDBench combines 100 diagnostic tasks with 34 executable diffraction-analysis workflows.
Summary
Ten models completed 1,340 benchmark attempts. Mean scores fell from 61.9 on diagnostic questions to 53.7 on executable workflows; traces exposed inaccurate structural results even when software execution succeeded.
AI role
LLM agents inspected diffraction files, selected refinement actions, ran crystallographic software, and preserved evidence under physical and execution checks.
Narrative role
This adds a materials-characterization benchmark that separates operating scientific software from obtaining defensible scientific results.
Caveat
The author-developed preprint benchmark has only 34 end-to-end tasks and uses different tasks across tracks. Scores do not measure discovery acceleration or establish independent deployment reliability.