Capability benchmark · Computer science
TasteVal isolates experiment-planning efficiency and compares research agents with human experts
An eight-task AI R&D benchmark measuring serial experimental-compute efficiency while holding the implementation agent fixed.
Summary
TasteVal evaluates 20 models and 24 human experts on eight new AI R&D tasks, using the best expert attempt per task as baseline. Runs allow 40 H100 GPU-busy hours or 120 elapsed hours. Opus 5.5 achieves a 2.30x compute multiplier (95% CI 1.15–4.37), comparing the GPU time needed to reach a matched score. The study reports faster recent growth in this efficiency metric, while final normalized performance shows no trend break.
AI role
Evaluated models propose experiments and interpret results; a fixed Opus 4.8 Coder implements and runs the experiments for models and human experts.
Narrative role
Adds a controlled test of which experiments to run and how to learn from their outcomes, separating research judgment from coding ability and end-to-end agent performance.
Caveat
Eight private, single-GPU tasks with fast clean feedback; tasks are not released and results lack independent replication. The multiplier measures serial experimental compute, excluding planning inference and elapsed delays, rather than overall research speed. Problem selection is outside scope; observed gains primarily involve shallow/moderate optimization.