Capability benchmark · Biology · Physics · General science
ScienceBoard evaluates computer-using agents in realistic scientific workflows
A multi-domain environment and benchmark for visually rich scientific workflows using professional software.
Summary
ScienceBoard introduced a benchmark and environment for multimodal, computer-using agents in scientific workflows involving professional software. The benchmark spans domains such as biochemistry, astronomy, and geoinformatics, and reported low overall success despite promising partial results.
AI role
Computer-using agents interacted with scientific software environments and were graded on workflow progress and final task outcomes.
Narrative role
This backfills a 2025-2026 benchmark milestone for evaluating whether agents can operate within the software surfaces scientists actually use, not just answer questions or write isolated code.
Caveat
The event date uses the original arXiv submission; the ICLR 2026 camera-ready version appeared later, and benchmark success depends on the chosen agent interfaces and grading templates.