Limitation or failure · Biology · Computer science

InquiTree finds research agents lose critical judgment in inquiry loops

A diagnostic benchmark of AI agents over paper-derived scientific inquiry trees

Summary

InquiTree models scientific work as an interactive inquiry loop over paper-derived research trees. The preprint reports that agents can develop cognitive tunneling during long-horizon interactions and perform worse on papers after model training cutoffs, suggesting apparent research competence can depend partly on memorized prior literature.

AI role

LLM-based agents proposed inquiry steps, interpreted study feedback, and updated beliefs across multi-turn research trees.

Narrative role

This adds a cautionary benchmark for research-agent reliability, shifting the evidence base from one-shot scientific QA toward whether agents preserve judgment, anomaly detection, and belief updating across a realistic inquiry process.

Caveat

The result is a preprint benchmark centered initially on neuroscience papers, so it measures diagnostic behavior rather than direct failure in deployed laboratories.