Limitation or failure · Computer science · General science
Large-scale study finds AI research agents narrow scientific exploration
A comparative analysis of AI-generated and human research ideas across scientific fields and historical research landscapes.
Summary
Tang and Yang analyze 219,655 ideas produced by five AI research-agent frameworks and five language models. Across their experiments, AI-generated ideas were more concentrated than comparable human work, remained closer to their starting literature, aligned less with later human research, and occupied lower-impact parts of the historical research landscape.
AI role
Five agent frameworks and five language models generated research ideas that were measured for concentration, distance from seed literature, future alignment, and location in the historical landscape.
Narrative role
This shifts the limitations evidence beyond execution failures to a field-level exploration concern: scaling idea generation may reinforce local trajectories rather than expand the search space that produces high-impact research.
Caveat
This is a preprint using operational measures of novelty, future alignment, and historical impact; its conclusions may depend on the selected fields, seed literature, agent scaffolds, and bibliometric proxies.