Capability benchmark · Biology · Medicine · Computer science

BioXArena benchmarks agents on biomedical ML pipeline tasks

A biomedical machine-learning benchmark with executable end-to-end modelling tasks across heterogeneous data domains.

Summary

BioXArena introduces a benchmark for LLM agents building biomedical machine-learning pipelines across multimodal datasets. Tasks require executable code, model training, and submissions against private test samples, making it closer to applied biomedical research work than static question answering.

AI role

Agents write code, train predictive models, and produce submissions for held-out biomedical machine-learning tasks.

Narrative role

This strengthens the benchmark layer for biomedical research automation by testing full modelling workflows rather than isolated biomedical knowledge questions.

Caveat

The work is a recent preprint and reported results depend on benchmark design, private graders, and a fixed two-hour single-GPU evaluation setting.