Capability benchmark · Chemistry · Medicine · Materials science
RetroChimera improves expert-rated retrosynthesis through complementary model ensembling
A peer-reviewed retrosynthesis model evaluated on rare reactions, distribution shifts and blind expert ratings of multistep routes.
Summary
RetroChimera combines two models with complementary failure modes and was tested across public chemistry data plus proprietary GSK and Novartis datasets. In a blind expert assessment of ten difficult targets, chemists accepted nine complete routes from RetroChimera, compared with five, four and two for its component or baseline systems. The implementation and public checkpoints were released.
AI role
A learned ensemble combines a sequence generator and a graph-template model to rank precursor reactions and construct routes from target molecules to purchasable inputs.
Narrative role
Adds a peer-reviewed, chemist-judged capability result at a real bottleneck in molecular research: proposing viable routes before laboratory synthesis. It measures planning quality rather than completed molecules or elapsed research time.
Caveat
Expert acceptance is not experimental execution. The multistep test used ten targets, proprietary distribution-shift datasets are not fully inspectable, lower-ranked generative predictions can hallucinate, and no causal design-make-test acceleration was measured.