Capability benchmark · Biology · Physics · General science

ScienceBoard evaluates computer-using agents in realistic scientific workflows

A multi-domain environment and benchmark for visually rich scientific workflows using professional software.

Summary

ScienceBoard introduced a benchmark and environment for multimodal, computer-using agents in scientific workflows involving professional software. The benchmark spans domains such as biochemistry, astronomy, and geoinformatics, and reported low overall success despite promising partial results.

AI role

Computer-using agents interacted with scientific software environments and were graded on workflow progress and final task outcomes.

Narrative role

This backfills a 2025-2026 benchmark milestone for evaluating whether agents can operate within the software surfaces scientists actually use, not just answer questions or write isolated code.

Caveat

The event date uses the original arXiv submission; the ICLR 2026 camera-ready version appeared later, and benchmark success depends on the chosen agent interfaces and grading templates.