Limitation or failure · Computer science · General science

Large-scale study finds AI research agents narrow scientific exploration

A comparative analysis of AI-generated and human research ideas across scientific fields and historical research landscapes.

Summary

Tang and Yang analyze 219,655 ideas produced by five AI research-agent frameworks and five language models. Across their experiments, AI-generated ideas were more concentrated than comparable human work, remained closer to their starting literature, aligned less with later human research, and occupied lower-impact parts of the historical research landscape.

AI role

Five agent frameworks and five language models generated research ideas that were measured for concentration, distance from seed literature, future alignment, and location in the historical landscape.

Narrative role

This shifts the limitations evidence beyond execution failures to a field-level exploration concern: scaling idea generation may reinforce local trajectories rather than expand the search space that produces high-impact research.

Caveat

This is a preprint using operational measures of novelty, future alignment, and historical impact; its conclusions may depend on the selected fields, seed literature, agent scaffolds, and bibliometric proxies.

Related events