Source-linked AI summary

The Values Encoded in Machine Learning Research

Abeba Birhane, Pratyusha Kalluri, Dallas Card, William Agnew, Ravit Dotan, Michelle Bao

arXiv:2106.15590v2cs.LGcs.AIcs.CY

TL;DR

The paper examines whether machine-learning research is value-neutral and what values it advances. It introduces an annotation scheme and applies it to 100 papers, finding that performance and related technical values dominate while societal justification is uncommon.

  • Problem

    Machine-learning research is often cast as value-neutral, even though neutrality can insulate AI from critique and permit emphasis on its benefits.

  • Method

    The paper presents an open-source fine-grained annotation scheme and uses line-by-line inductive annotation to study values in documents, including randomly sampled research papers.

  • Results

    Performance appears in 96% of the 100 papers, followed by generalization at 89%, building on past work at 88%, quantitative evidence at 85%, efficiency at 84%, and novelty at 77%; 68% omit societal need or impact.

  • Takeaways & Limitations

    The findings support the conclusion that machine-learning research is not value-neutral and that its technical goals and notions of performance warrant scrutiny.

  • Takeaways & Limitations

    The paper notes that task choices can seem arbitrary, while authors often claim generalization without adequately justifying it.

Abstract

from arXiv · show

Machine learning currently exerts an outsized influence on the world, increasingly affecting institutional practices and impacted communities. It is therefore critical that we question vague conceptions of the field as value-neutral or universally beneficial, and investigate what specific values the field is advancing. In this paper, we first introduce a method and annotation scheme for studying the values encoded in documents such as research papers. Applying the scheme, we analyze 100 highly cited machine learning papers published at premier machine learning conferences, ICML and NeurIPS. We annotate key features of papers which reveal their values: their justification for their choice of project, which attributes of their project they uplift, their consideration of potential negative consequences, and their institutional affiliations and funding sources. We find that few of the papers justify how their project connects to a societal need (15\%) and far fewer discuss negative potential (1\%). Through line-by-line content analysis, we identify 59 values that are uplifted in ML research, and, of these, we find that the papers most frequently justify and assess themselves based on Performance, Generalization, Quantitative evidence, Efficiency, Building on past work, and Novelty. We present extensive textual evidence and identify key themes in the definitions and operationalization of these values. Notably, we find systematic textual evidence that these top values are being defined and applied with assumptions and implications generally supporting the centralization of power.Finally, we find increasingly close ties between these highly cited papers and tech companies and elite universities.

1 INTRODUCTION

The paper challenges portrayals of machine learning as value-neutral by examining which values the field prioritizes and how social forces shape research and its beneficiaries. It introduces an annotation scheme and applies it to 100 highly cited ICML and NeurIPS papers, finding dominant technical values alongside evidence of centralized power and growing corporate ties.

  • Machine learning research is shaped by social forces that influence what research gets done and who benefits.
  • The paper introduces a fine-grained annotation scheme for studying values in documents such as research papers.The authors describe it as the first scheme of its kind and make it openly available for further qualitative and quantitative analysis.
  • The most frequently emphasized values are Performance, Generalization, Quantitative evidence, Efficiency, Building on past work, and Novelty.The paper argues that these apparently technical values are socially and politically charged in their definitions and operationalization.
  • The paper finds textual evidence that these values are operationalized in ways that centralize power and neglect society’s least advantaged.
  • Corporate affiliation among top-cited papers rose from 24% in 2008/09 to 55% in 2018/19, while big-tech presence rose from 21% to 66%.

2 METHODOLOGY

The paper develops a qualitative annotation scheme and applies it to 100 highly cited ICML and NeurIPS papers, examining their justifications, uplifted values, negative-impact discussions, affiliations, and funding. Line-by-line annotation and consensus procedures produce a large repository and identify dominant values while preserving textual and contextual interpretation.

  • Corpus and sampling: The study analyzes 100 highly cited NeurIPS and ICML papers from 2008, 2009, 2018, and 2019, annotating over 3,500 sentences.The corpus was selected because highly cited papers reflect and shape disciplinary values.
  • Annotation scheme: The annotation scheme examines project justifications, uplifted values, potential negative impacts, author affiliations, and stated funding sources.Researchers reviewed abstracts, introductions, discussions, and conclusions, then quantified and contextualized the annotations.
  • Outputs and interpretation: The researchers provide the annotation scheme, complete annotations, quantitative patterns, sampled quotations, and themes showing how values become socially loaded.They also compare prominent values with alternatives that might have been valued instead.
  • Annotation process: A hybrid inductive-deductive content analysis began with prior ethical principles, added emergent values through discussion, and then applied constant comparative annotation across the corpus.The process combined predefined categories with discovery of values grounded in the texts.
  • Scope and limitations: The paper treats qualitative analysis as both a methodological contribution and a limitation, because nuanced values can be misinterpreted in quantitative-leaning contexts.The authors argue that automated labeling misses emergent and context-dependent values, supporting the need for qualitative methodology.

3 QUANTITATIVE SUMMARY

Across 100 annotated papers, technical performance-oriented values dominate, while societal justification and discussion of negative impacts are uncommon. The corpus also shows increasingly close institutional ties to corporations and big tech, alongside substantial university connections.

  • Dominant values: Performance appears in 96% of papers, followed by generalization (89%), building on past work (88%), quantitative evidence (85%), efficiency (84%), and novelty (77%).These are the values most frequently used to justify and assess machine-learning research.
  • User rights and ethics: None of the papers mentioned autonomy, justice, or respect for persons, values associated with user rights and ethical principles.The finding contrasts with the prevalence of technical and research-community-oriented values.
  • Project justification: 68% of papers made no mention of societal need or impact, while only 4% rigorously connected their research to societal needs.Most papers justified their internal technical goals rather than broader societal relevance.
  • Negative impacts: 98% of papers contained no reference to potential negative impacts; one discussed them and a second merely mentioned their possibility.The negative-potential annotations therefore indicate very limited engagement with possible harms.
  • Institutional ties: Between the earlier and later publication periods, corporate ties nearly doubled to 79%, big-tech ties more than tripled to 66%, and university ties declined to 81%.Corporate presence consequently approached university presence in the later papers.

4 TEXTUAL ANALYSIS

The papers primarily justify projects through ML community needs and technical progress, while rarely connecting them rigorously to societal needs or discussing negative potential. Their dominant values—especially performance, generalization, efficiency, and novelty—are operationalized through assumptions that can centralize power.

  • Justification and negative potential: 15% of papers justify how their project connects to a societal need, while most motivate projects through ML research-community needs.Examples include understanding model behavior, improving efficiency, and creating benchmarks.
  • Justification and negative potential: Societal applicability is usually mentioned in introductions but rarely justified, revisited, or connected to the paper’s specific contribution.Papers often cite applications such as object detection or text classification without explaining why those applications are worth advancing.
  • Justification and negative potential: Negative potential is extremely rarely discussed, including in papers advancing facial surveillance, DeepFakes, and misleading videos.Among the two papers that mention negative potential, discussions are mostly abstract or hypothetical rather than assessments of the proposed models.
  • Dominant values: The dominant values are Performance, Generalization, Building on past work, Quantitative evidence, Efficiency, and Novelty.The paper argues that these values are not purely technical because their definitions and operationalizations encode taken-for-granted assumptions.
  • Performance: Performance is commonly operationalized as equally weighted average correctness, which can deprioritize underrepresented people, data, and inclusion needs.The annotated papers contained no discussion of alternative weighting approaches from fairness-related research.
  • Generalization: Large preestablished datasets and benchmark conventions privilege institutions with resources, while generalization often promotes one model across diverse settings.The paper links these assumptions to representational harms and reduced attention to socially distinct contexts.

4.6 Efficiency

Efficiency is commonly framed as enabling larger-scale computation rather than conserving resources or improving accessibility. This framing favors scaling by powerful actors and can make models less accessible to resource-constrained communities.

  • Efficiency: Efficiency is commonly defined through resource use, including data, energy, labels, memory, cost, speed, and training time.The paper finds that defining efficiency specifies which resources matter and what purpose efficiency serves.
  • Efficiency: Efficiency is frequently presented as enabling scale-up, reflecting the Jevon’s paradox that greater resource efficiency can increase total resource utilization.The operational emphasis is on doing more computation rather than reducing aggregate consumption.
  • Efficiency: 84% of papers mention valuing efficiency, but only 15% of those value requiring few resources.More efficient inference is often used to run larger models or datasets with the same or greater resources.
  • Efficiency: No papers present evidence that efficiency facilitates low-resource communities or reduces hardware needs, resource extraction, or carbon emissions.Instead, scaling computation can make models less accessible to people without comparable resources and reduce their ability to compete.
  • Efficiency: An accessibility-oriented definition of efficiency could prioritize more equitable conditions instead of scalability.The paper presents accessibility as an alternative usage of the value rather than the dominant operationalization.

5 CORPORATE AFFILIATIONS AND FUNDING

Corporate and Big Tech ties to highly cited ML papers increased substantially from 2008/09 to 2018/19, alongside strong representation from elite universities. The authors connect this institutional concentration with values that may reinforce centralized power.

  • Corporate presence: Big Tech author affiliations increased nearly fourfold, from 13% in 2008/09 to 47% in 2018/19.The named examples of large technology firms include Google and Microsoft.
  • Corporate presence: Corporate ties among annotated papers increased from 45% in 2008/09 to 79% in 2018/19.The measure includes corporate-affiliated authors or corporate funding.
  • Institutional concentration: 80% of papers with university affiliations came from elite universities, defined as the top 50 universities by QS World University Rankings.The analysis also reports pronounced corporate presence among the most-cited ICML and NeurIPS papers.
  • Values and power: The paper identifies a connection between concentrated institutional influence and ML values such as performance, generalization, and efficiency.It states that these values may facilitate Big Tech objectives while suppressing beneficence, justice, and inclusion.
  • Values and power: Large datasets and priorities such as accuracy, efficiency, and scale can make user safety, informed consent, and participation appear costly or time-consuming.The authors frame this as evading social needs in a context of centralized power.

6 DISCUSSION AND RELATED WORK

The paper extends interdisciplinary critiques of value-laden technology to machine learning, situating its analysis within STS, critical theory, philosophy, and justice-oriented scholarship. It presents a characterization of the field’s current values as a resource for understanding and transforming them.

  • Related work: STS, critical theory, and philosophy treat science and technology as value-laden rather than value-neutral.
  • Related work: Prior scholarship examines how technology can encode political, colonial, racist, sexist, and other marginalizing values.
  • Contribution: The paper extends these critiques to machine learning as part of a broader interdisciplinary body of work.
  • Context: Growing institutional and grassroots attention to ML’s societal impacts includes new organizations, workshops, FAccT, and NeurIPS broader-impact statements.
  • Contribution: The characterization is intended to help people understand, shape, dismantle, or transform current field values and articulate alternatives.

7 CONCLUSION

The analysis finds robust qualitative and quantitative evidence that influential machine learning research is not value-neutral. These papers often neglect societal needs and harms while prioritizing values and institutional ties associated with concentrated power.

  • The analysis finds robust evidence against treating machine learning as value-neutral, describing it as socially and politically loaded.
  • Influential papers favor research-community and large-firm needs over broader social needs and generally fail to acknowledge critiques or alternatives.
  • Project selection, limited attention to negative impacts, and operationalized values such as performance, generalization, efficiency, and novelty disfavor societal needs.
  • Big tech and elite universities have an overwhelming and increasing presence in the highly cited papers, consistent with power-centralizing commitments.
  • The papers’ operationalization of these values is described as supporting concentration of resources, tools, knowledge, and power among already powerful actors.

A.1 Data Sources

The study selected highly cited NeurIPS and ICML papers using Semantic Scholar citation data and examined their affiliations and funding sources. It also documents database and corpus boundaries relevant to interpreting the sample.

  • Paper selection: Semantic Scholar supplied bibliographic information and citation counts for determining which papers were most cited.
  • Paper selection: The sample consists of the most cited papers from NeurIPS and ICML in 2008, 2009, 2018, and 2019.
  • Scope and limitations: Semantic Scholar is imperfect: the selection included one 2010 paper and one retracted paper, while citation counts were static and may differ across sources.
  • Scope and limitations: The annotations reproduce sentences from published papers, which may contain personally identifying information or offensive content present in those papers.
  • Affiliations and funding: The authors quantify affiliations and funding using definitions of elite universities based on QS computer-science rankings and big tech based on a specified company list.

C COMBINING VALUES

The appendix consolidates overlapping fine-grained annotations into broader value categories for the main analysis. It lists the pre-combination values separately and reports the resulting sets.

  • Combining values: Overlapping values such as Performance and Accuracy were identified as closely related during annotation.
  • Combining values: Fine-grained categories such as Data Efficiency and Label Efficiency were combined under the broader value Efficiency.
  • Combining values: The authors combined related values for the main analysis following identified best practices.
  • Combining values: Figure C.1 lists all annotated values before combining, while Table 7 lists the combinations used in the main paper.
  • Combining values: The main value sets include Performance, Building on past work, Generalization, and Efficiency with their associated subvalues.

D REFLEXIVITY STATEMENT

The authors explicitly challenge objectivity claims while reflecting on how their identities, affiliations, and institutional histories shape the research. They acknowledge that their interpretation may omit values important to marginalized communities and reproduce educational hierarchies.

  • The authors challenge the idea that machine learning research is insulated from values and historical context.
  • The research team spans technical, critical, organizing, artistic, and philosophical backgrounds, while remaining multi-racial and multi-gender.
  • Affiliations with well-resourced Western universities made the research and publication process easier than it may have been for peers in the global South.
  • The authors acknowledge that different researchers might identify different values and that their analysis may overlook what matters to communities at the margins.
  • They caution that treating some universities as “elite” imposes a potentially unjust hierarchy and that broader ecosystem assumptions may remain unexamined.

E EXPERIMENTS WITH USING TEXT CLASSIFICATION TO IDENTIFY VALUES

The paper tests simple text classifiers as a way to estimate value prevalence beyond the manually annotated sample. Classifier performance is generally poor, so the broader prevalence and temporal trends require substantial caution.

  • Classifier method: The authors train separate regularized logistic-regression classifiers for values with at least 20 relevant sentences, using unigram features and held-out evaluation.
  • Classifier results: Most classifiers achieve F1 scores around 0.5 or lower, with Unifying Ideas reaching an F1 score of 0.
  • Classifier results: Performance, Accuracy, State-of-the-art, Effectiveness, and Facilitating Use are exceptions, with F1 scores greater than 0.75.
  • Corpus expansion: The authors apply the classifiers to ICML and NeurIPS papers from 2008–2020, estimating the proportion containing at least one sentence predicted to express each value.
  • Corpus expansion: The estimated relative prevalence broadly resembles the annotated sample, but predicted frequencies are often higher and should be treated skeptically because classifier performance is poor.
  • Temporal analysis: Performance-related values appear to have gradually become more common in NeurIPS over time, although the authors emphasize that further investigation is needed.

F CODE AND REPRODUCIBILITY

The authors release code and annotations and describe computational experiments as modest in resource use. They also discuss ethical risks, attribution choices, and concerns that the paper may be perceived as unrepresentative or harmful to conference norms.

  • Reproducibility: The paper’s code and annotations are publicly available under a CC BY-NC-SA license.
  • Reproducibility: The text-classification experiments ran on a 2019 MacBook Air, while the computational work was conducted locally with resource usage comparable to everyday computer use.
  • Ethics: The authors report minimal concern about risks to living beings, human rights, livelihoods, and annotator participation, since all annotators were co-authors.
  • Attribution: They omit author attributions from randomly selected quoted examples to reduce punitive effects while retaining a full cited-paper list for transparency and reproducibility.
  • Scope and reception: The authors acknowledge that some readers may view the paper as unrepresentative or detrimental to an ML conference, while arguing that these conversations belong in prominent venues.
  • Scope and reception: They frame prevailing norms as potentially contingent and open to being transformed, dismantled, or reenvisioned.

H RANDOM EXAMPLES

The random examples illustrate how ML papers encode values through claims about performance, efficiency, novelty, robustness, formal analysis, practical use, and understanding. They also show that papers sometimes foreground limitations, failures, or unresolved theoretical questions.

  • Values in examples: The examples commonly associate contributions with Performance, Accuracy, Quantitative evidence, Generalization, Efficiency, Novelty, and Robustness.
  • Limitations: Some examples explicitly identify limitations, including failure to address test loss, adversarial examples, non-convergence, inaccurate evaluation methods, and open theoretical problems.
  • Practical relevance: Other examples connect value to practical relevance through real-world data, practice, popular applications, user influence, and diverse outputs.
  • Efficiency and novelty: Several papers present new algorithms or formulations as valuable because they are fast, scalable, simple, or applicable to large datasets.
  • Theory and guarantees: The examples treat theoretical guarantees and formal analysis as important evidence, including results about convergence, sparse recovery, rank minimization, and generalization.
  • Limitations: The examples also show that empirical claims can qualify benefits: robustness evaluations may poorly represent real-world performance, and methods may fail to outperform simpler combinations.
Loading 2106.15590v2…