Capability benchmark · Computer science · General science
DiscoveryBench evaluates LLM agents on data-driven discovery tasks
LLM-agent performance on expert-curated data-driven discovery tasks.
Summary
The arXiv paper introduces DiscoveryBench, a benchmark for data-driven discovery with 264 expert-curated tasks across six domains and additional synthetic tasks. The authors evaluate LLM-based agents and report that the best systems solve only a minority of the benchmark.
AI role
Worked with data and generated testable findings instead of answering isolated questions.
Narrative role
DiscoveryBench expands the evidence ledger for whether agents can do realistic data-discovery work.
Caveat
Benchmark performance remains upstream of actual scientific discovery and downstream validation.