Limitation or failure · Biology · Computer science
InquiTree finds research agents lose critical judgment in inquiry loops
A diagnostic benchmark of AI agents over paper-derived scientific inquiry trees
Summary
InquiTree models scientific work as an interactive inquiry loop over paper-derived research trees. The preprint reports that agents can develop cognitive tunneling during long-horizon interactions and perform worse on papers after model training cutoffs, suggesting apparent research competence can depend partly on memorized prior literature.
AI role
LLM-based agents proposed inquiry steps, interpreted study feedback, and updated beliefs across multi-turn research trees.
Narrative role
This adds a cautionary benchmark for research-agent reliability, shifting the evidence base from one-shot scientific QA toward whether agents preserve judgment, anomaly detection, and belief updating across a realistic inquiry process.
Caveat
The result is a preprint benchmark centered initially on neuroscience papers, so it measures diagnostic behavior rather than direct failure in deployed laboratories.