Capability benchmark · Biology · Medicine

SpatialBench-Long probes agents on end-to-end spatial biology reasoning

A long-horizon spatial biology benchmark requiring agents to recover biological claims from raw or near-raw measurement data.

Summary

SpatialBench-Long introduced a benchmark for long-horizon spatial biology analysis across cancer, organoid, lineage-tracing, aging, and intervention systems. The authors hardened candidate claims through reproduction, independent scientist review, and trajectory inspection, and reported only 11.1% success for the best tied model-harness pairs.

AI role

Agents analyzed spatial and single-cell biology datasets, selected methods, and produced controlled-vocabulary answers graded against hardened claims.

Narrative role

This adds a current biology benchmark that targets the harder step from running analyses to deriving accurate biological conclusions from complex measurements.

Caveat

The benchmark is small and preprint-stage, and the low success rates reflect particular harnesses, rubrics, and curated biological claims.