Capability benchmark · Physics · Chemistry · Computer science

PhySciBench tests deep-research agents in physics and chemistry workflows

A 200-question physical-science benchmark spanning physics and chemistry tasks such as reasoning, extraction, experimental design, and code generation.

Summary

PhySciBench is a June 2026 benchmark for deep-research agents in physical sciences, with 200 expert-curated questions balanced across physics and chemistry. Reported baseline performance remains low, while the accompanying DelveAgent framework improves accuracy and lowers inference cost relative to the strongest baseline.

AI role

Evaluated deep-research agents and models on multi-step physical-science tasks, then tested a specialized multi-agent framework against baselines.

Narrative role

This adds current capability evidence outside biology and ML engineering, showing how deep-research agents perform on realistic physical-science task mixes.

Caveat

The source is a newly submitted preprint, and some evaluation uses rubric-guided semantic judgment; it should be treated as benchmark evidence, not demonstrated scientific acceleration.