Capability benchmark · Computer science

SWE-bench turns real GitHub issues into a coding-agent benchmark

SWE-bench turns real GitHub issues into a coding-agent benchmark: capability signal for AI systems on research-adjacent tasks.

Summary

The arXiv paper introduces SWE-bench, an evaluation framework with 2,294 software-engineering problems drawn from real GitHub issues and pull requests across 12 Python repositories. Initial evaluations found contemporary models solved only a small fraction of issues.

AI role

AI systems are tested on research-adjacent capabilities relevant to computer science.

Narrative role

This is supporting evidence for whether AI systems can perform research-adjacent tasks needed before stronger discovery or acceleration claims.

Caveat

Benchmark, model, or tool performance is an upstream capability indicator, not proof of new scientific discovery.