Capability benchmark · Computer science · Biology · Physics · Chemistry
GPQA benchmarks models on expert-written graduate science questions
Expert-written graduate-level science questions for testing advanced scientific reasoning.
Summary
The arXiv paper introduces GPQA, a set of 448 multiple-choice questions written by domain experts in biology, physics and chemistry. The questions are designed to be difficult for skilled non-experts using the web and for frontier AI systems, creating a benchmark for advanced scientific question answering.
AI role
Answered hard biology, physics, and chemistry questions designed to be difficult for non-experts using the web.
Narrative role
GPQA tracks whether models can reason over expert science content, a prerequisite for many research-assistant claims.
Caveat
Question answering is a capability indicator and does not directly measure discovery or experiment quality.