Capability benchmark · Physics · Chemistry · Biology

FrontierScience evaluates models on expert-level STEM research tasks

A science benchmark combining olympiad-level problems and PhD-level open-ended research subtasks.

Summary

FrontierScience proposes a benchmark for expert-level scientific reasoning across physics, chemistry, and biology. It includes both olympiad-style problems and PhD-level open-ended research subtasks, with rubric-based evaluation for the research track.

AI role

Language models attempt expert scientific reasoning tasks and are graded with rubrics for research-subtask performance.

Narrative role

This adds cross-domain STEM coverage for capability measurement, especially where existing benchmarks may be saturated or too dependent on published knowledge.

Caveat

It measures problem-solving and research subtasks, not autonomous end-to-end scientific discovery.