Capability benchmark · Computer science
DS-1000 benchmarks language models on data-science code
DS-1000 benchmarks language models on data-science code: capability signal for AI systems on research-adjacent tasks.
Summary
The arXiv paper presents DS-1000, a benchmark of 1,000 data-science programming problems spanning seven Python libraries and based on Stack Overflow-style questions. It evaluates generated code with executable tests and reports much lower performance on this setting than on simpler coding benchmarks.
AI role
AI systems are tested on research-adjacent capabilities relevant to computer science.
Narrative role
This is supporting evidence for whether AI systems can perform research-adjacent tasks needed before stronger discovery or acceleration claims.
Caveat
Benchmark, model, or tool performance is an upstream capability indicator, not proof of new scientific discovery.