Capability benchmark · Materials science · Chemistry · Computer science

AutoXRD benchmarks agents on executable powder-diffraction workflows

XRDBench combines 100 diagnostic tasks with 34 executable diffraction-analysis workflows.

Summary

Ten models completed 1,340 benchmark attempts. Mean scores fell from 61.9 on diagnostic questions to 53.7 on executable workflows; traces exposed inaccurate structural results even when software execution succeeded.

AI role

LLM agents inspected diffraction files, selected refinement actions, ran crystallographic software, and preserved evidence under physical and execution checks.

Narrative role

This adds a materials-characterization benchmark that separates operating scientific software from obtaining defensible scientific results.

Caveat

The author-developed preprint benchmark has only 34 end-to-end tasks and uses different tasks across tracks. Scores do not measure discovery acceleration or establish independent deployment reliability.