Capability benchmark · Computer science
InfiAgent-DABench evaluates agents on executable data-analysis tasks
InfiAgent-DABench evaluates agents on executable data-analysis tasks: capability signal for AI systems on research-adjacent tasks.
Summary
The arXiv paper introduces InfiAgent-DABench, a benchmark and agent framework for data-analysis tasks that require interaction with an execution environment. It includes 257 questions derived from 52 CSV files and evaluates 34 language models.
AI role
AI systems are tested on research-adjacent capabilities relevant to computer science.
Narrative role
This is supporting evidence for whether AI systems can perform research-adjacent tasks needed before stronger discovery or acceleration claims.
Caveat
Benchmark, model, or tool performance is an upstream capability indicator, not proof of new scientific discovery.