Source-linked AI summary
Machine Learning that Matters
Kiri Wagstaff
TL;DR
Machine learning research often focuses on isolated datasets and abstract metrics rather than concrete impact or sustained engagement with the problems it aims to address. This paper presents six Impact Challenges and discusses obstacles to reconnecting ML research with science and society, aiming to focus attention and inspire discussion.
Problem
Machine learning lacks a field-level objective tied to concrete impact, while datasets, metrics, and research practices often remain disconnected from real-world domains.
Method
The paper presents six Impact Challenges and discusses obstacles to conducting ML research that engages domain experts, relevant users, and originating scientific communities.
Results
The paper identifies six examples of Impact Challenges and several obstacles, including heavy reliance on benchmark datasets and abstract metrics and limited interpretation in domain context.
Takeaways & Limitations
Real impact requires connecting algorithmic advances with problem formulation, data and evaluation choices, domain interpretation, communication, user adoption, and deployment.
Takeaways & Limitations
UCI datasets lack both controlled generation and real-world context, limiting their utility as evidence for impactful ML research.
Abstract
from arXiv · showhide
Much of current machine learning (ML) research has lost its connection to problems of import to the larger world of science and society. From this perspective, there exist glaring limitations in the data sets we investigate, the metrics we employ for evaluation, and the degree to which results are communicated back to their originating domains. What changes are needed to how we conduct research to increase the impact that ML has? We present six Impact Challenges to explicitly focus the field?s energy and attention, and we discuss existing obstacles that must be addressed. We aim to inspire ongoing discussion and focus on ML that matters.
1. Introduction
The paper argues that much ML research has become disconnected from consequential problems in science and society. It calls for evaluating and communicating ML advances in terms of concrete impact, identifying the gap and proposing Impact Challenges and first steps to address it.
- Many published ML papers evaluate algorithms on isolated benchmark data sets, rarely communicate results back to their originating domains, or assess whether performance gains matter outside ML.
- The paper attributes this disconnection partly to limited emphasis in graduate training and peer review on connecting ML advances to the larger world.
- The authors ask whether ML should optimize performance on isolated data sets or characterize progress through the concrete impact of innovations.
- This position paper contains no algorithms, theorems, experiments, or results; instead, it identifies the connection problem, suggests first steps, issues Impact Challenges, and identifies obstacles.
- The paper aims to focus future research efforts on impact and stimulate thought and discussion, regardless of whether readers agree with every statement.
2. Machine Learning for Machine Learning’s Sake
The paper argues that ML research often prioritizes benchmark algorithms and abstract metrics while neglecting domain interpretation, real-world impact, and follow-through. Publishing incentives reinforce this imbalance by rewarding middle-stage technical work more than connecting research to meaningful applications.
- 2.1. Hyper-Focus on Benchmark Data Sets: ML research commonly evaluates new algorithms on isolated synthetic and benchmark data sets, with little interpretation in the originating domain.At ICML 2011, 34 of 148 papers used only UCI and/or synthetic data, while just 1 of 148 interpreted results in domain context.
- 2.1. Hyper-Focus on Benchmark Data Sets: Benchmark data sets enable comparisons but are undermined by inconsistent methodologies and limited connection to experts, users, and operational systems.The paper argues that UCI data are neither controlled synthetic data nor genuinely contextualized real-world data, and may overemphasize classification and regression.
- 2.2. Hyper-Focus on Abstract Metrics: Abstract metrics such as accuracy, RMSE, F-measure, and AUC obscure whether performance differences matter for a specific domain.The same percentage can have different practical meanings: 80% accuracy may suffice for iris classification, whereas mushroom edibility may require 99% or higher.
- 2.3. Lack of Follow-Through: ML publishing incentives disproportionately reward middle-stage algorithmic or theoretical contributions, leaving little incentive to connect advances with the outer world.The paper identifies this incentive structure as an obstacle to reconnecting active research with relevant real-world problems.
- 2.2. Hyper-Focus on Abstract Metrics: Statistical significance does not establish real-world significance, while aggregate benchmark averages and algorithm bake-offs reveal little about impact or domain-specific sources of performance.ROC and AUC reporting can also ignore relevant operating regimes and unequal costs of false positives and false negatives.
- 2.3. Lack of Follow-Through: Research seeking real impact must extend from problem identification and data collection through evaluation, interpretation, expert involvement, communication, adoption, and eventual difference-making.The paper presents these activities as necessary components of an impact-oriented research program, not merely optional application details.
3. Making Machine Learning Matter
The paper calls for ML research to measure and pursue impact in the originating domain, not merely improve isolated technical performance. It recommends impact-aware evaluation, domain collaboration, independent assessment, and problem selection tied to meaningful real-world gains.
- ML research should fundamentally change how projects are formulated, attacked, and evaluated rather than merely report isolated applications.
- 3.1. Meaningful Evaluation Methods: Impact-aware metrics can include dollars saved, lives preserved, time conserved, effort reduced, and quality of living increased.Publications can also report how accuracy improvements translate into impact for the originating domain.
- 3.1. Meaningful Evaluation Methods: Using domain-specific impact measures avoids treating equal percentage improvements as equivalent across unrelated domains.The paper contrasts profit improvements for an auto-tire business with avoided surgical interventions.
- 3.2. Domain Expertise: Domain experts can connect performance plots to scientific significance and identify systems that are numerically good but unreliable for adoption.
- 3.2. Domain Expertise: Independent domain researchers could assess an ML advance’s performance, utility, and impact while helping other communities understand its methods.
- 3.3. Eyes on the Prize: Research problems should be selected partly by the number of people, species, countries, or areas affected and by the meaningful improvement over the status quo.
- 3.3. Eyes on the Prize: A fetal-hypoxia system illustrates the proposed path from collaboration and measurable performance to clinical trials and potential deployment.It detected 50% of cases early enough for intervention with a 7.5% false positive rate.
4. Machine Learning Impact Challenges
The paper proposes six Impact Challenges that define machine learning success through concrete effects in law, finance, diplomacy, cybersecurity, medicine, and human development. Unlike technical contests or human-level benchmarks, they span domains and emphasize sufficient performance for real-world impact.
- 4. Machine Learning Impact Challenges: The six Impact Challenges include a law or legal decision relying on ML analysis, $100M saved, and a conflict averted through translation.
- 4. Machine Learning Impact Challenges: They also include a 50% reduction in cybersecurity break-ins, a life saved through ML-recommended care, and a 10% increase in one country’s HDI.
- 4. Machine Learning Impact Challenges: The challenges cover the full process of a successful ML endeavor, including performance, infusion, and impact.
- 4. Machine Learning Impact Challenges: Unlike domain-specific technical challenges, the Impact Challenges are not restricted to one problem domain or technical capability.
- 4. Machine Learning Impact Challenges: The list is explicitly non-comprehensive and is intended to inspire additional challenges benefiting the field.
- 4. Machine Learning Impact Challenges: Impact Challenges do not treat human-level performance as the gold standard; performance is sufficient when it can make an impact on the world.
5. Obstacles to ML Impact
The paper identifies communication barriers and poor scalability as obstacles that prevent ML from producing widespread impact. Specialized jargon blocks cross-domain communication, while deployment and maintenance can require scarce expert labor.
- Jargon: Specialized ML vocabulary creates conceptual barriers for domain experts, the public, and closely related fields such as statistics.The paper encourages expressing ideas in general or audience-familiar terms.
- Requiring a Ph.D. to deploy, maintain, and update ML systems does not scale to widespread impact.Simplifying, maturing, and robustifying algorithms and tools could permit wider independent use.
6. Conclusions
The paper argues that ML research often ends with publication to the ML community rather than usable communication back to the originating problem setting. It presents broad opportunities for impact and six challenges intended to stimulate discussion and adoption.
- Many ML researchers isolate themselves with a data set, optimize algorithmic performance, and stop the process at publication to the ML community.
- Successes are often not communicated back to the original problem setting or are not provided in a usable form.
- Law, finance, politics, medicine, and education could benefit from ML systems that analyze, adapt, and take or recommend action.
- The paper identifies six Impact Challenges and obstacles to inspire discussion about how ML can make a difference.It states that real impact is needed for the wider world to notice, value, and adopt ML solutions.