Capability benchmark · Biology · Medicine · Computer science
BioXArena benchmarks agents on biomedical ML pipeline tasks
A biomedical machine-learning benchmark with executable end-to-end modelling tasks across heterogeneous data domains.
Summary
BioXArena introduces a benchmark for LLM agents building biomedical machine-learning pipelines across multimodal datasets. Tasks require executable code, model training, and submissions against private test samples, making it closer to applied biomedical research work than static question answering.
AI role
Agents write code, train predictive models, and produce submissions for held-out biomedical machine-learning tasks.
Narrative role
This strengthens the benchmark layer for biomedical research automation by testing full modelling workflows rather than isolated biomedical knowledge questions.
Caveat
The work is a recent preprint and reported results depend on benchmark design, private graders, and a fixed two-hour single-GPU evaluation setting.