Limitation or failure · Medicine · General science

AI literature tools fail to reliably flag retracted papers

A content-analysis evaluation of AI tools handling retracted scientific articles during literature-search tasks.

Summary

Labenbacher and coauthors test nine AI tools on literature tasks involving retracted papers and find that none consistently handles retraction status correctly. Research-focused tools did not produce a single fully correct response set, and several tools included retracted articles in topic overviews without adequate warning.

AI role

General-purpose and research-focused AI tools generated literature-search responses and were evaluated for correct retraction handling.

Narrative role

This is a practical caution signal for AI-assisted evidence synthesis, especially in medicine, where speed gains can propagate invalid evidence if retraction checks are weak.

Caveat

The evaluation used a fixed set of 15 retracted articles and free-access tool versions, so performance may change as products update.