Capability benchmark · Computer science

InfiAgent-DABench evaluates agents on executable data-analysis tasks

InfiAgent-DABench evaluates agents on executable data-analysis tasks: capability signal for AI systems on research-adjacent tasks.

Summary

The arXiv paper introduces InfiAgent-DABench, a benchmark and agent framework for data-analysis tasks that require interaction with an execution environment. It includes 257 questions derived from 52 CSV files and evaluates 34 language models.

AI role

AI systems are tested on research-adjacent capabilities relevant to computer science.

Narrative role

This is supporting evidence for whether AI systems can perform research-adjacent tasks needed before stronger discovery or acceleration claims.

Caveat

Benchmark, model, or tool performance is an upstream capability indicator, not proof of new scientific discovery.