Limitation or failure · Computer science · General science · Social science
Patterns paper frames data leakage as an ML-based science reproducibility crisis
A peer-reviewed survey and case study of data leakage failures in machine-learning-based scientific claims.
Summary
Kapoor and Narayanan survey reproducibility issues in machine-learning-based science and argue that data leakage can produce overoptimistic scientific claims. Their Patterns article reports leakage across 17 fields and 294 papers, proposes an eight-type taxonomy, and shows in a civil-war prediction case study that corrected analyses remove the claimed advantage of complex ML models.
AI role
Machine-learning models were used as evidence for scientific claims, but leakage in training, evaluation, or feature construction made some reported advantages overoptimistic.
Narrative role
This backfills a concrete methodological limitation from the last few years: AI and ML can accelerate publication and modeling while also amplifying hard-to-detect validity failures across scientific domains.
Caveat
The paper focuses on predictive ML used to support scientific claims, not all AI-for-science workflows, and the count reflects surveyed leakage evidence rather than a complete census of affected literature.