Capability benchmark · Computer science
DSBench tests data-science agents on more realistic tasks
DSBench tests data-science agents on more realistic tasks: capability signal for AI systems on research-adjacent tasks.
Summary
The arXiv paper introduces DSBench, a benchmark with 466 data-analysis tasks and 74 data-modeling tasks sourced from Eloquence and Kaggle. It emphasizes long contexts, multimodal task backgrounds, large data files, multi-table structures and end-to-end modeling, and reports that tested agents struggled with most tasks.
AI role
AI systems are tested on research-adjacent capabilities relevant to computer science.
Narrative role
This is supporting evidence for whether AI systems can perform research-adjacent tasks needed before stronger discovery or acceleration claims.
Caveat
Benchmark, model, or tool performance is an upstream capability indicator, not proof of new scientific discovery.