Source-linked AI summary

Lessons from Archives: Strategies for Collecting Sociocultural Data in Machine Learning

Eun Seo Jo, Timnit Gebru

arXiv:1912.10389v1cs.LGcs.AIcs.CY

TL;DR

ML fairness, accountability, transparency, and ethics problems are rooted partly in overlooked data collection and annotation practices. The paper compares sociocultural ML collection with archival methods and finds that archives provide institutional and procedural structures for addressing these concerns, while noting scaling and motivation differences.

  • Problem

    Data collection remains an overlooked and insufficiently systematic part of ML despite shaping outcomes and raising concerns about consent, inclusivity, power, transparency, and ethics & privacy.

  • Method

    The paper draws parallels between sociocultural ML data collection and archival practices, including interventionist procedures and community-controlled approaches.

  • Results

    Archives have institutional and procedural structures regulating data collection, annotation, and preservation that ML can draw from.

  • Takeaways & Limitations

    ML research should become more cognizant and systematic in data collection and draw on interdisciplinary expertise from archives and libraries.

  • Takeaways & Limitations

    Applying archival methods to ML datasets may require substantial overhead, time, financial cost, coordination, and community effort at scale.

Abstract

from arXiv · show

A growing body of work shows that many problems in fairness, accountability, transparency, and ethics in machine learning systems are rooted in decisions surrounding the data collection and annotation process. In spite of its fundamental nature however, data collection remains an overlooked part of the machine learning (ML) pipeline. In this paper, we argue that a new specialization should be formed within ML that is focused on methodologies for data collection and annotation: efforts that require institutional frameworks and procedures. Specifically for sociocultural data, parallels can be drawn from archives and libraries. Archives are the longest standing communal effort to gather human information and archive scholars have already developed the language and procedures to address and discuss many challenges pertaining to data collection such as consent, power, inclusivity, transparency, and ethics & privacy. We discuss these five key approaches in document collection practices in archives that can inform data collection in sociocultural ML. By showing data collection practices from another field, we encourage ML research to be more cognizant and systematic in data collection and draw from interdisciplinary expertise.

1 INTRODUCTION

Data collection and annotation shape ML outcomes, yet sociocultural data practices remain insufficiently systematic. The paper draws on archival institutions and procedures to address consent, inclusivity, power, transparency, and ethics & privacy in ML.

  • Motivation: Haphazardly categorizing people can harm vulnerable groups and propagate societal biases through ML systems.Phenotypic traits have historically been used to target groups, while face recognition creates related risks.
  • Motivation: Diverse training data and disaggregated testing can be important for accurately diagnosing all demographic groups.Melanoma detection illustrates the need for demographic diversity, while such testing can require sensitive information and categorization.
  • Research gap: ML data collection lacks a rigorous, systematic process and an industry-wide standard for fairness-aware collection.Researchers have characterized dataset generation as the “wild west.”
  • Approach: The paper examines archives as an interdisciplinary source of language and institutional approaches for collecting and annotating sociocultural data.It organizes archival lessons around consent, inclusivity, power, transparency, and ethics & privacy.
  • Contribution: Archives have institutional and procedural structures regulating data collection, annotation, and preservation that ML can draw from.The paper presents these structures as lessons for a more cognizant and systematic ML data pipeline.

2 WHAT ARE ARCHIVES?

Archives are collective systems for systematically storing human materials for scholarly, heritage, and future uses. Archival studies have developed concepts and procedures relevant to sociocultural data collection in ML.

  • Archival purpose: Archives systematically store historical and current materials as collective human records for academic, scholarly, heritage, and legacy purposes.They predate digital materials and have existed for thousands of years.
  • Archival purpose: Institutional, governmental, foundational, and research-oriented archives share the objective of collecting human materials for future uses.Many modern archives also include digital components.
  • Relevance to ML: Archival studies have developed literature and procedures addressing data labeling, private information, dataset sharing, diversity, inclusivity, appraisal, and selection.These concerns echo recent fairness initiatives in ML.
  • Supervision scale: The paper frames data collection practices along a supervision scale that includes laissez-faire collection, community archives, and curatorial archives.The figure presents example categories rather than a universally preferred position.

3 DIFFERENCES BETWEEN ARCHIVAL AND ML DATASETS

Archival and ML datasets differ in supervision, objectives, and collection practices. The paper argues that examining these differences helps ML researchers communicate strategies and intervene before inherited biases are embedded in datasets.

  • Implications: Identifying collection differences gives ML researchers vocabulary for communicating strategies and encourages more cognizant, active data collection.The paper presents archival approaches as parallels rather than a single universally preferred method.
  • Supervision: Current ML collection often emphasizes dataset size and efficiency, producing minimally supervised and potentially indiscriminate data collection.The paper contrasts this with curatorial archives’ more interventionist supervision.
  • Supervision: Curatorial archives use mission statements, trained archivists, and filtering layers to decide which sources enter a collection.Some archives deliberately target minority-group collections to diversify holdings.
  • Objectives: ML datasets commonly target task accuracy, whereas curatorial archives prioritize heritage, memory, education, authenticity, privacy, inclusivity, and rarity.The paper argues that archival concerns should also inform ML objectives and supervision levels.
  • Bias: Datasets lacking adequate intervention can replicate historical and representational biases that arise before sampling, weighting, and balancing.The paper therefore places intervention before these later techniques.
  • Bias: Internet-crawled datasets require an interventionist layer because their materials reflect demographic and sociocultural inequities.News sources also retain political and topical biases despite editorial screening.

5 LESSONS FROM ARCHIVES

The paper organizes archival lessons around fairness, accountability, transparency, and ethics, showing how institutional and interventionist practices can inform sociocultural ML data collection.

  • Inclusivity: ML data collection has often been driven by task requirements, convenience, or digital availability rather than fair representation and diversity.This can narrow demographic scope and make models reflect the specific source from which data were gathered.
  • Inclusivity: Archives begin collection with explicit objectives and mission statements, rather than simply using datasets available through existing channels.Public missions can guide source searches, filtering decisions, and continuing contributions as sociocultural norms change.
  • Consent: Community archives give represented groups ownership and agency to contribute cultural materials and define their own categorization and access protocols.Mukurtu lets Indigenous communities upload materials, flag sensitive content, and specify preferences for access, use, and circulation.
  • Consent: ML could adapt participatory archival structures by enabling open-ended community input and participant control over access, circulation, and sensitivity.The paper presents these structures as starting templates for community-centered ML data collection.
  • Ethics & privacy: Ethical data collection requires labor, expertise, infrastructure, and resources, which can disadvantage institutions with limited resources.Disaggregated testing and privacy-preserving handling add annotation and infrastructure costs.
  • Power: Data consortia can reduce technological overhead and redundant collections while increasing access to unique holdings, but may also create bureaucracy and power inequities.In ML, proprietary datasets held by large technology organizations may be unavailable for consortium sharing.

1. Mission Statement

Mission statements establish the highest-level agenda for the topics and concepts an archive intends to collect.

  • Mission statements formulate the highest-level agenda by determining which topics or concepts are of concern.

2. Collection Development Policy

Collection development policies translate a mission statement into specific rules for what to collect and how to locate relevant sources.

  • Collection development policies specify what is collected, what is excluded, and where and how to search for sources.

3. Appraisal

Appraisal evaluates whether potential sources fit the mission and are worth preserving for future use.

  • Appraisal evaluates sources for mission fit, rarity, provenance authenticity, and value for future generations.

4. Processing/Indexing (Micro-Appraisal)

Archives use layered review, processing, documentation, and ethical oversight to govern what sociocultural data is retained and how it is handled. These institutional procedures offer fair ML models despite adding development cost and delay.

  • Sources can be processed individually or at the folder/document level, indexed, and reflected in updated finding aids.
  • Committee-based review and delegated professional curators or processors distribute appraisal decisions across multiple levels of supervision.
  • Multiple review and record-keeping levels are largely absent from ML data collection, although adopting them increases development cost and lag.
  • Archival ethics systems address dilemmas involving retention, access to sensitive content, intellectual property, and professional conduct.
  • Professional membership, ethics panels, and case-by-case evaluation provide incentives and mechanisms for enforcing archival standards in data collection.
  • Cross-institutional organizations can help ethical principles withstand pressure from individual employers to reduce safeguards.

6 TWO LEVELS OF ACTIONABLES

The paper proposes coordinated action at macro and micro levels to improve ML data collection and annotation. Community-wide institutions and project-level practices reinforce one another, especially when accountability extends beyond individual projects.

  • Macro level: Macro-level actions include data consortia, professional organizations enforcing ethical guidelines, community archives, and a dedicated data-collection subfield.
  • Micro level: Micro-level actions include revising mission statements, hiring professionally accountable data-collection staff, supporting public datasets, and adopting rigorous documentation.
  • Micro level: Micro-level policies also call for stronger collection-development rules and informed committee decisions about discretionary data.
  • Macro and micro actions reinforce one another because central organizations outside individual projects can hold data collectors accountable.

7 CASE STUDY: GPT-2 AND REDDIT

The WebText case study applies archival concepts to GPT-2’s Reddit-derived corpus, examining its motivations, sociocultural constraints, and possible improvements. The authors propose clearer mission statements, more inclusive collection strategies, ethical supervision, and documentation.

  • WebText composition and motivation: WebText contains approximately 40G of text from 8 million online documents outlinked from Reddit pages with at least three net upvotes.Its stated objective was to collect large, diverse, high-quality language data for transfer across multiple NLP tasks.
  • WebText composition and motivation: WebText’s collection approach was cost- and time-effective, contemporary, and aligned with web-scraped validation data, but was not optimized for sociocultural inclusivity.The authors infer that model-performance goals, rather than cultural diversity, shaped the dataset’s design.
  • Sociocultural constraints: WebText’s hard constraint is reliance on publicly accessible online documents, while its softer limitation is inheritance of Reddit’s demographic and platform biases.The dataset therefore emphasizes written online language and topical relevance to Reddit’s predominantly young, male, urban U.S. user base.
  • Proposed improvements: A mission statement can clarify WebText’s origin, intervention level, intended applications, scope, and opportunities for participation or contribution.The paper suggests a complementary dataset targeting groups whose language is sparsely digitized or underrepresented in Reddit-centric internet English.
  • Proposed improvements: More inclusive collection could combine community archives with participatory schemes in rural, multi-ethnic, and low-internet-use communities.These approaches would collect materials such as oral histories and interviews from demographics underrepresented in existing online sources.
  • Proposed improvements: Ethical supervision requires full-time data-collection and management staff, an ethics code, screening for private data, and documentation of selection decisions.The paper also recommends documenting subreddit selection and scraping conditions to improve transparency.

8 LIMITATIONS

The archival strategies proposed for ML face scalability, incentive, and governance challenges. Their transferability depends on substantial coordination, resources, and adaptation to ML’s institutional context.

  • Scalability and cost: Scaling archival guidelines to ML datasets may require substantial time, financial resources, staffing, documentation, and coordination.Community data consortia could reduce costs through economies of scale and resource sharing, but require sustained collective effort.
  • Institutional incentives: ML datasets and archives have different motivations: commercial and corporate incentives often favor internet users and proprietary data, whereas archives preserve cultural heritage and diversity.The paper identifies incentive-model analysis as outside its scope but relevant to adapting archival recommendations.
  • Governance risks: Interventionist collection can concentrate agenda-setting power in archivists and expose collection priorities to undue social or political influence.Multi-person systems and governance bodies can diffuse power and provide accountability, but such safeguards are not currently in place across ML.

9 CONCLUSION

The paper argues that archival and library practices offer institutional and procedural strategies for addressing ethics, representation, power, transparency, and consent in sociocultural ML data collection. It also points to other social sciences as complementary sources of expertise.

  • Archival lessons for ML: Archives and libraries have developed institutional and procedural responses to ethics, representation, power, transparency, and consent in human data collection.Applying these strategies in ML requires allocated funds and collective institutional effort.
  • Interdisciplinary expertise: Sociology, psychology, history, and anthropology can complement archival lessons on human subjects, privacy, representation, historical context, and cultural sensitivities.The paper encourages ML researchers to learn from older fields’ successes and failures on comparable matters.
Loading 1912.10389v1…