Source-linked AI summary
Do Datasets Have Politics? Disciplinary Values in Computer Vision Dataset Development
Morgan Klaus Scheuerman, Emily Denton, Alex Hanna
TL;DR
Computer-vision datasets shape algorithm development and consequential applications, yet existing documentation guidance does not fully articulate the values embedded in dataset creation. The paper analyzes dataset documentation through structured and thematic content analysis, finding that the field privileges efficiency, universality, impartiality, and model work over care, contextuality, positionality, and data work, and recommends attending to these silenced values.
Problem
Existing dataset documentation frameworks improve transparency but generally do not articulate the values shaping computer-vision dataset development.
Method
The authors analyze computer-vision dataset documentation using structured and qualitative content analysis across collection, curation, annotation, and release practices.
Results
The analysis finds that computer-vision dataset authors value efficiency, universality, impartiality, and model work over care, contextuality, positionality, and data work.
Takeaways & Limitations
The authors recommend incorporating the silenced values into dataset creation and curation to support more trustworthy, ethical, and human-centered practices.
Takeaways & Limitations
The authors’ perspectives shape their interpretation, and computer-vision researchers may characterize the field’s values differently.
Abstract
from arXiv · showhide
Data is a crucial component of machine learning. The field is reliant on data to train, validate, and test models. With increased technical capabilities, machine learning research has boomed in both academic and industry settings, and one major focus has been on computer vision. Computer vision is a popular domain of machine learning increasingly pertinent to real-world applications, from facial recognition in policing to object detection for autonomous vehicles. Given computer vision's propensity to shape machine learning research and impact human life, we seek to understand disciplinary practices around dataset documentation - how data is collected, curated, annotated, and packaged into datasets for computer vision researchers and practitioners to use for model tuning and development. Specifically, we examine what dataset documentation communicates about the underlying values of vision data and the larger practices and goals of computer vision as a field. To conduct this study, we collected a corpus of about 500 computer vision datasets, from which we sampled 114 dataset publications across different vision tasks. Through both a structured and thematic content analysis, we document a number of values around accepted data practices, what makes desirable data, and the treatment of humans in the dataset construction process. We discuss how computer vision datasets authors value efficiency at the expense of care; universality at the expense of contextuality; impartiality at the expense of positionality; and model work at the expense of data work. Many of the silenced values we identify sit in opposition with social computing practices. We conclude with suggestions on how to better incorporate silenced values into the dataset creation and curation process.
1 Introduction
This section frames computer vision datasets as consequential components of machine-learning systems and examines how their documentation communicates disciplinary values. The study analyzes dataset documentation to identify both emphasized and silenced values in curation practices.
- Computer vision datasets train, evaluate, and standardize algorithms used in consequential domains, making dataset practices important to scrutinize.Applications include facial recognition, autonomous vehicles, robotics, military identification, refugee resettlement, and humanitarian aid.
- Existing reporting frameworks seek greater transparency but generally do not explain how specific values shape dataset curation.
- Undocumented choices in task formulation, data collection, annotation, and related processes can render value-laden decisions invisible.
- The study analyzes documentation from 113 datasets and 114 publications across face-based, body-based, and non-corporeal tasks using structured and thematic content analysis.The codebook covers stages from collection and annotation to dissemination.
- The authors identify efficiency over care, universality over contextuality, impartiality over positionality, and model work over data work.
- They argue that incorporating the silenced values could make computer-vision data curation more trustworthy, ethical, and human-centered.
2 Related Work
Related work treats data and dataset development as value-laden scientific and technical practices. This paper extends critiques of computer-vision datasets by systematically studying documentation and the disciplinary values it reveals.
- Scientific and technological production involves judgments about what becomes visible or invisible and documented or undocumented.
- Data collection, classification, and analysis reflect disciplinary and researcher values, making data organization politically consequential.
- Social-computing design frameworks have explicitly incorporated politics and values into design practices.
- Prior audits identify representational problems, offensive categories, privacy concerns, consent issues, and licensing problems in computer-vision datasets.
- Dataset development is undervalued relative to algorithmic work and is often marginalized by publication and review practices.
- This paper provides the first large-scale systematic study of computer-vision dataset documentation and shared disciplinary values.
- Unlike prospective documentation frameworks, the study retrospectively examines current practices and the values structuring dataset development.
3 Methods
The methods combine a broad manuscript corpus, task-based dataset identification, structured coding, thematic analysis, and explicit reflection on researcher positionality. The authors acknowledge that their interpretations may not match computer-vision researchers’ own characterizations.
- The authors state that their perspectives shape interpretation and that computer-vision researchers may characterize the field’s values differently.
- The manuscript corpus contains 50,694 computer-vision papers identified through IEEE proceedings and complementary searches in IEEE Xplore.
- The analysis examined documentation across dataset collection, curation, annotation, and release using structured and qualitative content analysis.
- Task categories were developed from standardized and author keywords, then clustered into 21 conceptually related topics.
- Researchers sampled papers within task categories, identified referenced datasets, and expanded the corpus through snowballing.
3.4 Sampling Datasets for Analysis
The dataset sample was drawn from a corpus of 487 datasets using popularity-based and stratified sampling. The resulting sample proportionally represented broad categories and was judged broadly representative of the population.
- The full corpus contained 487 datasets, from which the authors sampled a more manageable analysis set.
- Datasets with over 4,000 Google Scholar citations were selected to capture popular datasets.
- The popularity-based sample included 13 datasets spanning face-based, body-based, and non-corporeal categories.
- An additional 100 datasets were sampled proportionally across the face, body, and non-corporeal categories.
- The sample and population had the same rounded mean publication year, 2011, while their medians were 2012 and 2011, respectively.
- The authors judged the sample a good representation of dataset types common in the population corpus.
- The codebook captured dataset-creation stages from motivations and annotations through data availability.
3.6 Analysis
The study combined structured and thematic analyses to examine what computer vision dataset documentation communicates and omits about dataset development. It used a codebook spanning curation stages and coded value statements into higher-level themes.
- Coding covered dataset curation from data collection through annotation and dissemination across 114 coded datasets.The corpus included 114 dataset publications, with coding responsibilities divided among the authors.
- The analysis combined structured content analysis with qualitative thematic analysis of computer vision dataset documentation.The structured analysis examined explicitly documented curation practices, while thematic coding examined language communicating underlying values and silences.
- The analysis treated both communicated and uncommunicated dataset information as signals of what dataset authors and the computer vision community value.This approach supported abductive reasoning about dimensions, motivations, and silences around dataset creation.
- Researchers split value statements into open codes and then grouped them into higher-level themes and categorical relationships.Examples included grouping codes such as “unbiased data” under broader themes.
- The researchers released their codebook, dataset corpus, coding documentation, and Python analysis code in an open-access repository.The repository also listed publication, venue, year, and citation-count information for each dataset.
4 Findings
The qualitative analysis organized findings into three broad themes: dataset authors’ disciplinary practices, data properties, and human actors involved as annotators and data subjects.
- The fourteen focused qualitative codes were grouped into three themes: dataset authors, data properties, and human actors.The codes were not mutually exclusive, and some were highly related.
4.1 Dataset Authors and their Disciplinary Practices
Dataset authors presented data as essential for scientific progress, standardization, reproducibility, and open research, but documentation and maintenance practices were uneven. Their papers emphasized methodological and technical work more consistently than data collection and annotation work.
- 4.1.1 Data as Essential for Scientific Progress: 18 of 114 datasets (15.8%) described data as crucial for advancing their subfields beyond the state of the art.Datasets were presented as enabling new challenges, tasks, or shared research agendas.
- 4.1.1 Data as Essential for Scientific Progress: Dataset authors often described insufficient relevant data as limiting research and framed new datasets as barriers-to-entry reductions and standards for underresearched tasks.Leeds Sports Pose introduced 2,000 publicly available annotated consumer images to address limited training data.
- 4.1.1 Data as Essential for Scientific Progress: The dataset was the main contribution in 59 of 114 papers (51.8%), while 97 of 114 papers (85.1%) also introduced a new algorithm.This pattern indicates that dataset releases were commonly paired with methodological innovation.
- 4.1.1 Data as Essential for Scientific Progress: Dataset documentation showed a bimodal distribution, with papers either largely describing the dataset or providing scant dataset information while emphasizing methodological innovations.The proportion dedicated to dataset description had a mean of 0.41 and standard deviation of 0.33; 105 of 114 papers (92%) were archival publications.
- 4.1.2 Standardization for Evaluation and Reproducibility: Authors created standardized benchmarks to compare methods using common data, annotations, and evaluation protocols.The 300-W dataset was motivated by developing a standardized benchmark for facial landmark localization.
- 4.1.2 Standardization for Evaluation and Reproducibility: Standardization was associated with quantitative evaluation viewed as objective, reliable, and reproducible, while qualitative assessment was treated as less rigorous.HumanEva authors described heuristic and qualitative evaluation as making state-of-the-art comparisons difficult.
- 4.1.3 Open Source Data: 60.5% of datasets provided a URL in the paper, and 85% had a site discoverable through the paper or dataset name.However, only 3 of 114 datasets (2.6%) had a DOI, and only one (0.9%) was posted on an institutional repository.
- 4.1.3 Open Source Data: Only 46 of 69 datasets (66.7%) with a paper URL remained available, while 59 of 80 openly accessible datasets (73.8%) remained downloadable.The findings distinguish nominal openness from continued availability and stable maintenance.
4.2 Data and its Desirable Properties
Dataset authors describe desirable data through diversity, realism, lack of bias, quality, comprehensiveness, and scale, often linking these properties to model performance and generalization. These priorities can conflict: challenging, realistic variation may reduce image-resolution quality, while large-scale collection makes annotation labor costly.
- Diversity and variety: Diversity is valued as variation in scenes, capture conditions, poses, backgrounds, and appearances that makes datasets more realistic and challenging.Authors connect diversity to deep learning effectiveness and argue that insufficient variation limits new approaches.
- Unbiased data: Unbiased data is treated as desirable because dataset bias is associated with failures to generalize across datasets or from datasets to the real world.Bias discussions cover collection processes and image properties, including selection, photographer, recency, pose, and illumination biases.
- High-quality data: High-quality data is defined mainly through clear, high-resolution images and accurate or consistent annotations.Annotation quality is assessed against gold-standard labels, annotator agreement, or model performance; one benchmark reports 1024 × 440 px versus 360 × 288 px resolution.
- Realistic data: Realistic data reflects uncontrolled conditions and is valued for estimating performance and developing models that generalize to real-world settings.BiosecurID explicitly contrasts laboratory performance with practical implementations and calls for realistic multimodal biometric data.
4.3 Human Actors as Annotators and Data Subjects
Dataset documentation presents annotators and data subjects as necessary to data construction while often treating human subjectivity, variation, privacy, and labor costs as problems to control or minimize. Human diversity is discussed mainly as a technical property of data instances, with limited attention to annotator demographics or ethics.
- Annotation labor and time costs: 63 of 114 datasets (55%) used human annotation, but only 5 of 63 papers reported annotator demographics and 4 of 63 reported compensation.Annotator identities were reported in 40 of 63 cases, most often identifying third-party workers, authors, or students.
- Annotation labor and time costs: Authors valued manual annotation for accuracy while seeking to minimize the time and money spent on human labor.Crowdworking platforms were valued for supplying substantial labor at low cost, and annotation costs were described as barriers to large-scale datasets.
- Humans as data subjects: Human characteristics were often framed as difficult-to-control sources of variation that could complicate data collection, task specification, or model accuracy.Examples include variation in appearance, pose, clothing, imaging conditions, walking style, and other behaviors.
- Humans as diverse data: The term diversity usually referred to variation in visual instances rather than human conditions such as race, gender, or ability.When human diversity was discussed, authors most often reported age, sometimes race or ethnicity, and rarely gender beyond a binary framing.
- Humans as diverse data: Human diversity was generally justified through technical accuracy, although PPB also noted that underrepresented groups may be frequently targeted despite lower benchmark representation.This links representational gaps to consequences beyond model performance without treating diversity solely as a technical property.
5 Discussion
The discussion interprets computer vision dataset development as a site where disciplinary values become embedded in technical artifacts. It identifies trade-offs between efficiency and care, universality and contextuality, impartiality and positionality, and model work and data work, while calling for broader institutional support for change.
- Discussion: Computer vision datasets are a key point in the model pipeline for examining how values become embedded in artifacts used to train, test, and validate models.The paper connects this importance to the increasing application of computer vision technologies in public life.
- Discussion: The study uses structured and qualitative content analysis of dataset collection, curation, annotation, and release practices across human and non-human tasks.The analysis centers on recurring themes including annotation, availability, categories, and collection processes.
- Discussion: The paper characterizes four value trade-offs: efficiency versus care, universality versus contextuality, impartiality versus positionality, and model work versus data work.The authors argue that explicit dataset values often ignore or implicitly critique values associated with human-centered computing.
- Discussion: Institutional incentive structures may prevent individual dataset authors from implementing changes effectively, so conferences, journals, and departments are also encouraged to act.The authors note that the highlighted values and recommendations may extend to other machine learning domains while still advocating domain-specific analyses.
5.1 Efficiency over Care
Computer vision dataset creators prioritize efficient data collection and annotation, often treating care for people, rights, and ethical practice as costs. The paper identifies practices that could make dataset creation more reflexive and attentive to participants and labor.
- Efficiency over Care: Dataset creators value data that is quickly and cheaply available, classifiable, and accurately annotatable.Efficiency includes time, monetary, and computational costs.
- Efficiency over Care: Debiasing and quality-control investments are justified by potential algorithmic gains but framed as costs to minimize.
- Efficiency over Care: The field pays limited attention to the perspectives, labor, and rights of human actors involved in dataset curation.The authors connect this omission to dehumanizing accounts of human subjects in machine learning.
- Efficiency over Care: Efficiency receives priority over thoughtful decision-making, ethical collection, fair compensation, and consideration of harmful social implications.Compensation for data and annotation labor was generally not reported.
- Efficiency over Care: CAFE provides a counterexample by obtaining parental consent, undergoing IRB review, and addressing harms and benefits for child participants.
- Recommendations for Incorporating the Value of Care: The authors recommend privacy-conscious data collection, explicit permissions, and compensation suited to data subjects and annotators.They note that compensation should reflect context and what annotators value.
5.2 Universality over Contextuality
Dataset creators favor large-scale, diverse, realistic data and universal classifications, associating them with comprehensive world representation, generalization, standardization, and reproducibility. This emphasis often leaves temporal, cultural, geographic, and use-specific context underexamined.
- Universality over Contextuality: Dataset creators value large-scale, diverse, realistic data because they assume it better approximates the world and supports models that generalize.
- Universality over Contextuality: Universal classifications support standardization and reproducibility but imply that the world can be organized through objective labels.
- Universality over Contextuality: Dataset documentation rarely addresses image geography, classification language, identity markers, or other circumstances shaping represented diversity.
- Universality over Contextuality: The authors argue for contextuality by considering intended use, inclusion criteria, culture, language, and location when curating data.
- Universality over Contextuality: Context can concern where a dataset will be used or which world it is intended to capture, including particular workplaces, communities, and cultures.
- Recommendations for Incorporating Contextuality: The authors recommend context-specific datasets and empirical studies of intended use through surveys, interviews, or ethnography.
5.3 Impartiality over Positionality
Dataset authors seek impartial, unbiased data by minimizing human subjectivity, while often omitting how interpretation, social position, and annotator expertise shape dataset construction. The paper recommends reflexive practices that make these influences explicit.
- Impartiality over Positionality: Dataset authors pursue impartiality by treating selection and observer bias as threats to trustworthy data.
- Impartiality over Positionality: Documentation rarely acknowledges that interpretation is unavoidable or explains how researchers’ social and professional positions shape resources and knowledge.
- Impartiality over Positionality: Claims of false objectivity can reduce data trust and utility by refusing to describe relevant human stakeholders.
- Impartiality over Positionality: Only five publications reported demographic information about annotators, despite annotators’ critical role in determining dataset contents.
- Impartiality over Positionality: PPB reports selection rationales, identity-based limitations, and ground-truth labeling by a board-certified surgical dermatologist.The authors present professional training as a possible basis for trust in the data process.
- Recommendations for Incorporating Positionality: The authors recommend reflexivity, including positionality statements describing how researchers’ perspectives and values shape the work.
5.4 Model Work over Data Work
Computer vision publishing privileges algorithmic and model contributions over documentation, maintenance, availability, and reproducibility of datasets. The authors therefore frame persistent data stewardship as necessary data work.
- Model Work over Data Work: Dataset documentation is often missing, and many datasets are unavailable despite data being presented as crucial to machine learning.Some archival papers preserve URLs that no longer work.
- Model Work over Data Work: Publications typically require algorithmic improvement, while dataset-only contributions are rarely published archivally.
- Model Work over Data Work: The field treats data as auxiliary to models and undervalues dataset maintenance and stability after publication.
- Model Work over Data Work: Dataset creators often express interest in sharing but remain silent about infrastructure for persistence and upkeep.
- Model Work over Data Work: Investments in data sharing affect long-term scientific policies and practices, while data dependencies contribute to machine-learning technical debt.
- Model Work over Data Work: CAFE exemplifies data work by using an institutional repository and receiving a DOI that enables persistent access.
- Recommendations for Incorporating Data Work: The authors recommend stable DOIs and institutional repositories to improve dataset transparency, usability, and reproducibility.
6 Conclusion
The study examines what computer vision dataset documentation reveals about disciplinary values and what developers leave unsaid. It finds that dataset practices broadly value efficiency, while foregrounding some aspects of development and omitting others.
- The study situates dataset practices within broader scholarship on the politics and values embedded in technical artifacts and data practices.
- Computer vision dataset documentation reveals how authors value different aspects of dataset development and reporting.
- The analysis also examines what dataset publications leave unsaid to better understand what dataset developers valued.