Capability benchmark · Computer science · General science

DiscoveryBench evaluates LLM agents on data-driven discovery tasks

LLM-agent performance on expert-curated data-driven discovery tasks.

Summary

The arXiv paper introduces DiscoveryBench, a benchmark for data-driven discovery with 264 expert-curated tasks across six domains and additional synthetic tasks. The authors evaluate LLM-based agents and report that the best systems solve only a minority of the benchmark.

AI role

Worked with data and generated testable findings instead of answering isolated questions.

Narrative role

DiscoveryBench expands the evidence ledger for whether agents can do realistic data-discovery work.

Caveat

Benchmark performance remains upstream of actual scientific discovery and downstream validation.