Limitation or failure · Materials science · Computer science
MatToolBench exposes low agent success on professional materials software
A reproducible real-environment benchmark of multimodal agents completing 204 tasks across ten professional materials-science tools.
Summary
MatToolBench evaluates scientific-software use with authentic experimental files and fine-grained expert criteria. Even the best evaluated model completed only 25% of GUI tasks and 45% of code tasks. Models often made partial progress, but failures in domain-specific operations, cross-tool artifact handoffs and visually exposed state prevented full workflow completion despite stronger performance on general desktop benchmarks.
AI role
Seven frontier multimodal models operated GUI applications, wrote OriginPro scripts and queried materials databases inside standardized Windows virtual machines.
Narrative role
Adds a realistic workflow-level boundary to research-agent capability claims by measuring specialized tools and artifacts rather than question answering or isolated code generation.
Caveat
The benchmark covers a finite set of materials tools, fixed software versions and 50-step episodes. It measures autonomous task completion, not whether human scientists become faster when using agents interactively.