Capability benchmark · Computer science · Biology · Physics · Chemistry

GPQA benchmarks models on expert-written graduate science questions

Expert-written graduate-level science questions for testing advanced scientific reasoning.

Summary

The arXiv paper introduces GPQA, a set of 448 multiple-choice questions written by domain experts in biology, physics and chemistry. The questions are designed to be difficult for skilled non-experts using the web and for frontier AI systems, creating a benchmark for advanced scientific question answering.

AI role

Answered hard biology, physics, and chemistry questions designed to be difficult for non-experts using the web.

Narrative role

GPQA tracks whether models can reason over expert science content, a prerequisite for many research-assistant claims.

Caveat

Question answering is a capability indicator and does not directly measure discovery or experiment quality.