Source-linked AI summary
Troubling Trends in Machine Learning Scholarship
Zachary C. Lipton, Jacob Steinhardt
TL;DR
Machine learning scholarship often departs from ideals of rigorous, clear, reader-serving knowledge creation. This paper organizes examples of four troubling trends, discusses their consequences and possible remedies, and argues for greater rigor while acknowledging countervailing costs and limits to its perspective.
Problem
ML papers can depart from clear communication, foundational knowledge creation, and empirically rigorous understanding through recurring trends in explanation, empirical evaluation, mathematics, and language.
Method
The paper describes four trends, gives specific and positive examples, explains consequences, and discusses counterarguments, historical context, and possible community responses.
Results
Careful evaluation reveals that reported gains may stem from hyper-parameter tuning rather than architectural innovation, while challenge datasets and mathematical claims can have important limitations.
Takeaways & Limitations
Greater rigor in exposition, science, and theory is presented as essential for scientific progress, productive public discourse, and responsible ML deployment in critical domains.
Takeaways & Limitations
The authors do not claim to offer a full or balanced view of ML science and acknowledge that their own arguments are selective and fallible.
Abstract
from arXiv · showhide
Collectively, machine learning (ML) researchers are engaged in the creation and dissemination of knowledge about data-driven algorithms. In a given paper, researchers might aspire to any subset of the following goals, among others: to theoretically characterize what is learnable, to obtain understanding through empirically rigorous experiments, or to build a working system that has high predictive accuracy. While determining which knowledge warrants inquiry may be subjective, once the topic is fixed, papers are most valuable to the community when they act in service of the reader, creating foundational knowledge and communicating as clearly as possible. Recent progress in machine learning comes despite frequent departures from these ideals. In this paper, we focus on the following four patterns that appear to us to be trending in ML scholarship: (i) failure to distinguish between explanation and speculation; (ii) failure to identify the sources of empirical gains, e.g., emphasizing unnecessary modifications to neural architectures when gains actually stem from hyper-parameter tuning; (iii) mathiness: the use of mathematics that obfuscates or impresses rather than clarifies, e.g., by confusing technical and non-technical concepts; and (iv) misuse of language, e.g., by choosing terms of art with colloquial connotations or by overloading established technical terms. While the causes behind these patterns are uncertain, possibilities include the rapid expansion of the community, the consequent thinness of the reviewer pool, and the often-misaligned incentives between scholarship and short-term measures of success (e.g., bibliometrics, attention, and entrepreneurial opportunity). While each pattern offers a corresponding remedy (don't do it), we also discuss some speculative suggestions for how the community might combat these trends.
1 Introduction
The paper argues that ML scholarship should serve readers through clear, rigorous communication, but identifies four troubling trends that depart from these ideals. It attributes their causes tentatively to community growth, limited reviewer capacity, and misaligned short-term incentives.
- ML papers can pursue theoretical characterization, empirically rigorous understanding, or high predictive accuracy, while serving readers through foundational knowledge and clear communication.
- The paper identifies four trends: conflating explanation with speculation, obscuring empirical gains’ sources, using mathiness, and misusing language.
- Mathiness mixes formal and informal claims in ways that can obscure theoretical problems and strengthen weak prose through an appearance of technical depth.
- Misused language includes terms with misleading colloquial connotations or overloaded technical meanings, potentially confusing readers across ML’s widening audience.
- The trends’ causes are uncertain, with possible contributors including rapid community expansion, a thin reviewer pool, and incentives tied to bibliometrics, attention, and entrepreneurship.
- The authors aim to improve precision and clarity to accelerate research, reduce onboarding time, and support more constructive public discourse.
2 Disclaimers
The authors frame the paper as a discussion piece offering a critical, insider perspective on troubling trends in machine learning scholarship rather than a comprehensive assessment.
- The paper aims to instigate discussion for the ICML Machine Learning Debates workshop, not provide a full or balanced evaluation of ML science.
- The authors present their critique as an introspective view from insiders, while noting that the identified ills are not specific to any individual or institution.
- They acknowledge that they have exhibited these patterns themselves and may do so again, while stressing that such patterns do not make a paper bad or indict its authors.
3 Troubling Trends
The paper identifies four troubling trends in machine learning scholarship: speculation presented as explanation, empirical gains whose sources are obscured, mathiness, and misuse of language. These practices can mislead readers by weakening evidential clarity, overstating technical contributions, or creating terminological confusion.
- Failure to Distinguish Explanation from Speculation: Papers often present speculation as authoritative explanation, even when concepts lack crisp definitions or supporting experiments.The authors distinguish exploratory intuition from explanations that can withstand scientific scrutiny.
- Failure to Identify the Sources of Empirical Gains: Uncontrolled collections of model tweaks and absent ablations can obscure which change produced an empirical gain and falsely imply that all changes are necessary.The paper emphasizes that proper ablation studies are needed to identify the source of improvements.
- Failure to Identify the Sources of Empirical Gains: Published architectural improvements have sometimes been attributed to complex innovations when better hyper-parameter tuning actually produced the gains.Melis et al. found that vanilla LSTMs topped the leaderboard under equal evaluation conditions.
- Mathiness: Mathiness mixes formal and informal claims without clearly relating them, allowing vague definitions to conceal theoretical problems and technical form to bolster weak prose.The paper also criticizes unnecessary or irrelevant theorems that lend apparent authority to empirical work.
- Misuse of Language: The paper identifies suggestive definitions, overloaded terminology, and suitcase words as recurring forms of language misuse in machine learning.Examples include anthropomorphic terms, imprecise uses of “human-level,” and technical words whose meanings drift across subfields.
- Misuse of Language: Overloading established terms such as “deconvolution” and “generative model” creates confusion by collapsing distinct technical meanings and capabilities.The paper contrasts formal definitions with newer usage that refers more broadly to upconvolutions or realistic-looking structured outputs.
4 Speculation on Causes Behind the Trends
The authors speculate that troubling ML scholarship trends may be rising because progress, community growth, reviewing constraints, and external incentives reward strong results over careful argumentation.
- The authors suspect the trends are increasing and associate them with complacency, rapid community expansion, a thin reviewer pool, and misaligned short-term incentives.
- Strong results may license unsupported explanations, inadequate disentangling experiments, exaggerated terminology, and less careful mathematics.
- Single-round review may pressure reviewers to accept papers with strong quantitative findings despite unresolved flaws.
- Newer researchers may be especially susceptible to terminology misuse, although experienced researchers also fall into these patterns.
- Rapid growth increases submitted papers per reviewer and reduces the fraction of experienced reviewers, potentially enabling or incentivizing several trends.
- Media attention and startup investment create incentives for trends such as anthropomorphic descriptions of ML algorithms.
5 Suggestions
The authors propose community practices that improve empirical inquiry, mathematical and linguistic clarity, reviewing incentives, retrospective synthesis, and critical discourse.
- 5.1 Suggestions for Authors: Authors should ask “what worked?” and “why?”, using error analysis, ablations, and robustness checks rather than reporting headline numbers alone.
- 5.1 Suggestions for Authors: Careful empirical inquiry can produce insight without proposing a new algorithm, including findings about random-label fitting and dataset limitations.
- 5.1 Suggestions for Authors: Authors should test whether explanations support predictions or working systems, helping align mathematical concepts with their intended insight.
- 5.1 Suggestions for Authors: Clearly separating open from solved problems gives readers a clearer picture, encourages follow-up work, and guards against neglecting falsely presumed-resolved questions.
- 5.2 Suggestions for Publishers and Reviewers: Reviewers should distinguish genuine contribution from bundled untested changes, favoring simple ideas with negative results over equally effective combinations without ablations.
- 5.2 Suggestions for Publishers and Reviewers: Authoritative retrospective surveys could remove exaggerated claims, replace anthropomorphic names, and standardize notation, but the authors find too few strong examples.
- 5.2 Suggestions for Publishers and Reviewers: The authors advocate giving critical writing a voice at ML conferences because algorithms and experiments are insufficient for evaluating problems or inquiry methods themselves.
- 5.2 Suggestions for Publishers and Reviewers: Peer-review questions about open review and reviewer point systems remain topics for further discussion rather than settled recommendations.
6 Discussion
The authors acknowledge that stronger scholarship standards can slow research, delay publication, or require costly experiments, but argue that these standards are usually worthwhile and should guide rather than prohibit sharing. They frame the paper as a contribution to recurring debate through which the community can self-correct.
- Countervailing Considerations: Some major results may justify delaying ablation studies when experiments are computationally expensive.The ImageNet breakthrough combined significant results with experiments whose ablations were costly to complete.
- Countervailing Considerations: High standards might impede unusual, speculative ideas and consume resources through lengthy publication processes.The authors cite economics, where a single paper can take years and lengthy revisions consume resources that could support new work.
- Countervailing Considerations: Researchers generating conceptual ideas or systems need not be the same people who carefully collate and distill knowledge.The authors identify specialization as a possible reason to separate innovation from careful synthesis.
- Countervailing Considerations: The proposed standards are strong heuristics rather than unbreakable rules, and the authors prefer sharing an idea when strict adherence would otherwise prevent it.They describe the standards as exacting but often requiring only a few extra days of experiments and more careful writing.
- Historical Context: The issues discussed recur throughout academia, including earlier debates in AI and documented crises caused by undisciplined scholarship.The authors connect current concerns to historical discussions of empirical standards, reproducibility, and misleading scientific enthusiasm.
- Concluding Remarks: Although these problems may be self-correcting, the authors argue that self-correction occurs through recurring debate about reasonable scholarship standards.The paper aims to contribute constructively to that debate rather than offer a full or balanced account of ML science.