Capability benchmark · Physics · Chemistry · Biology
FrontierScience evaluates models on expert-level STEM research tasks
A science benchmark combining olympiad-level problems and PhD-level open-ended research subtasks.
Summary
FrontierScience proposes a benchmark for expert-level scientific reasoning across physics, chemistry, and biology. It includes both olympiad-style problems and PhD-level open-ended research subtasks, with rubric-based evaluation for the research track.
AI role
Language models attempt expert scientific reasoning tasks and are graded with rubrics for research-subtask performance.
Narrative role
This adds cross-domain STEM coverage for capability measurement, especially where existing benchmarks may be saturated or too dependent on published knowledge.
Caveat
It measures problem-solving and research subtasks, not autonomous end-to-end scientific discovery.