Source-linked AI summary

Large image datasets: A pyrrhic win for computer vision?

Vinay Uday Prabhu, Abeba Birhane

arXiv:2006.16923v2cs.CYstat.APstat.ML

TL;DR

The paper examines ethical problems in large-scale vision datasets, focusing on consent, justice, privacy, stereotypes, and harmful content. Using ImageNet-ILSVRC-2012, it combines a 57-metric quantitative census with human-in-the-loop curation to identify transgressions and motivate justice-centered dataset governance.

  • Problem

    Large-scale vision datasets can contain consent, privacy, justice, stereotyping, and harmful-content problems, while their societal consequences remain insufficiently scrutinized.

  • Method

    The paper audits ImageNet-ILSVRC-2012 through cross-categorical model annotations, 57 metrics, human-in-the-loop analysis, and hand-curation of flagged subsets.

  • Results

    The census helped discover five dozen plus images across four consent-violating or pornographic categories in five ImageNet classes.

  • Takeaways & Limitations

    The paper advocates justice-centered practice and motivates mandatory Institutional Review Boards for large-scale dataset curation.

  • Takeaways & Limitations

    The paper cautions that social and ethical challenges are deeply rooted and that quick fixes, including complete bias removal, may create false reassurance.

Abstract

from arXiv · show

In this paper we investigate problematic practices and consequences of large scale vision datasets. We examine broad issues such as the question of consent and justice as well as specific concerns such as the inclusion of verifiably pornographic images in datasets. Taking the ImageNet-ILSVRC-2012 dataset as an example, we perform a cross-sectional model-based quantitative census covering factors such as age, gender, NSFW content scoring, class-wise accuracy, human-cardinality-analysis, and the semanticity of the image class information in order to statistically investigate the extent and subtleties of ethical transgressions. We then use the census to help hand-curate a look-up-table of images in the ImageNet-ILSVRC-2012 dataset that fall into the categories of verifiably pornographic: shot in a non-consensual setting (up-skirt), beach voyeuristic, and exposed private parts. We survey the landscape of harm and threats both society broadly and individuals face due to uncritical and ill-considered dataset curation practices. We then propose possible courses of correction and critique the pros and cons of these. We have duly open-sourced all of the code and the census meta-datasets generated in this endeavor for the computer vision community to build on. By unveiling the severity of the threats, our hope is to motivate the constitution of mandatory Institutional Review Boards (IRB) for large scale dataset curation processes.

1 Introduction

ImageNet transformed computer vision through unprecedented scale and influence, while motivating continuing scrutiny of consent, privacy, and societal implications in large image datasets.

  • 1 Introduction: Informed consent developed to protect human dignity, agency, identity, and control over disseminated personal information, including photographic data.Broad consent provides a less stringent framework that retains basic identity safeguards in large-scale databases.
  • 1 Introduction: ImageNet expanded image classification from smaller datasets to over 14 million images across 21,841 synsets.It also included 1,034,908 bounding-box annotations.
  • 1 Introduction: ImageNet remains influential more than a decade after creation, making continued critical dialogue relevant despite its age.Its societal impact and influence on later large-scale datasets motivate ongoing auditing.
  • 1 Introduction: The paper surveys ethical dimensions and threats, proposes possible solutions, and presents quantitative auditing using ILSVRC2012 as an example.It also describes curated data assets for the computer-vision community.

2 Background and related work

Related work identifies ethical problems in taxonomies, labels, dataset inheritance, representation, and annotation labor, including offensive categories retained across large image datasets.

  • 2 Background and related work: Taxonomies can make some identities visible while rendering others invisible, and classification systems may encode dominant ideologies and exclude marginalized groups.Essentialist gender categories, for example, can exclude non-binary and transgender people.
  • 2 Background and related work: Image-search and recognition systems have shown under-representation and unequal accuracy across gender and demographic groups, with darker-skin females suffering the most misclassification in one cited system.The cited literature also reports stereotype exaggeration and higher pedestrian-recognition error rates for darker skin tones.
  • 2 Background and related work: Image datasets inherit problematic assumptions from taxonomy sources, with ImageNet specifically built on WordNet’s structure.The ImageNet team acknowledged limitations in WordNet’s stagnant concept vocabulary.
  • 2 Background and related work: Crowdsourced labeling can exploit underpaid workers, while single-label procedures and restrictive proposals introduce annotation shortcomings for real-world images containing multiple objects.These concerns extend ethical scrutiny from dataset content to curation labor and validation practices.

3 The threat landscape

The paper describes harms from insufficiently scrutinized image datasets, including privacy threats, inherited stereotypes, opaque scaling, consent violations, and downstream burdens in models trained on tainted data.

  • 3 The threat landscape: Face-search services can help identify ImageNet individuals in the real world, creating risks for vulnerable people and marginalized populations.The paper emphasizes that large datasets built without societal consideration threaten individual welfare and well-being.
  • 3 The threat landscape: ImageNet helped normalize appropriating real people’s images as raw material, encouraging larger opaque datasets and raising concerns about privacy.The paper connects this culture to secretive systems such as Clearview AI.
  • 3 The threat landscape: Creative Commons licensing addresses copyright but not privacy rights, consent for training, research ethics, or regulation of online surveillance tools.The paper cites deleted datasets as further support for this distinction.
  • 3 The threat landscape: Models trained on tainted datasets can carry ethical residues into downstream applications, including neural generative art and face-recognition systems.The paper also describes privacy leakage through accurate extraction of subsets of training facial images.
  • 3 The threat landscape: Classification and labeling can decide what counts as legitimate, normal, or correct, thereby perpetuating historical and cultural prejudices.The paper states that AI systems trained on such data amplify and normalize these patterns.

4 Candidate solutions: The path ahead

The paper presents multiple corrective paths for ethical problems in large-scale vision datasets, while emphasizing that social and ethical challenges resist a single solution. Proposed measures include removal or replacement, privacy protection, synthetic data, curation-stage checks, and dataset audit cards.

  • Social and ethical challenges in large-scale vision datasets have no single straightforward solution, and quick fixes may conceal problems.Bias, discrimination, and injustice vary with context, history, and place.
  • Remove, replace, and open strategy: The authors advocate removing offensive labels and verifiably pornographic or non-consensual images, while considering consensually produced, compensated replacements.They distinguish concern about non-consensual imagery from concern about the category or its content.
  • Remove, replace, and open strategy: Dataset auditors can use reverse-image-search abuse portals to request removal of indexed images containing identifiable individuals.The authors frame this as a way to mitigate some immediate harms.
  • Privacy and synthetic alternatives: Privacy-preserving alternatives include DP-Blur and synthetic images, though synthetic-data approaches remain a nascent field.Examples include sketches, GAN-generated images, and dataset distillation.
  • Curation-stage safeguards: Explicit ethics instructions and interface checks during crowd validation could help filter ethical transgressions at the source.The authors propose integrating ethics checks into future humans-in-the-loop curation interfaces.
  • Dataset audit cards: Dataset audit cards should publish goals, curation procedures, known shortcomings, and caveats to provide context for evaluating dataset ethics.The paper illustrates this proposal with an ImageNet audit card based on its quantitative analyses.

5 Quantitative dataset auditing: ImageNet as a template

The audit uses ImageNet as a template for cross-categorical quantitative analysis, combining model-based metrics with curated meta-datasets to investigate ethical transgressions.

  • The census covers 57 metrics spanning person count, age, gender, NSFW scoring, class-label semanticity, and classification accuracy.It includes both image-level and class-level analysis.
  • The audit curated separate datasets for CAG statistics, NSFW scores, class-wise accuracy, and semanticity for community reuse.The assets include CSV files documenting these analysis dimensions.
  • Five classes were selected for further NSFW investigation: bikini, maillot, tank suit, miniskirt, and brassiere.
  • The analysis identified more than five dozen images across beach-voyeuristic, exposed-private-parts, verifiably pornographic, and upskirt categories.The images were found by combining gender, age, class semanticity, and NSFW information.

6 Conclusion and discussion

The authors argue that large image datasets remain ethically troublesome despite ImageNet’s achievements and call for justice-centered practices and IRBs for dataset curation.

  • Large image datasets can disproportionately harm vulnerable and marginalized communities through direct and indirect societal impacts.
  • The authors urge the machine learning community to center justice and the welfare of disproportionately impacted communities.
  • The paper hopes to motivate Institutional Review Boards for large-scale dataset curation processes.

Appendix A Risk of privacy loss via reverse search engines

The appendix describes privacy risks from increasingly efficient reverse image search engines that can identify people depicted in ImageNet.

  • Reverse image search services with face-search capabilities can uncover the real-world identities of ImageNet subjects for a small fee.
  • The quantitative audit includes NSFW scoring, semanticity, and classification accuracy among its analysis dimensions.

B.1 Count, Age and Gender

The count, age, and gender analysis combines InsightFace and DEX outputs to estimate person prevalence and examine agreement, gender skew, and demographic reliability.

  • InsightFace identified 101,070 persons across 83,436 images, estimating a 7.6% prevalence of people included without explicit consent.
  • DEX estimated person prevalence at 10.3%, higher than InsightFace’s 7.6% estimate because DEX has a higher identification false-positive rate.
  • The detected population showed relatively older male presence, with 73,746 males averaging 33.24 years versus 26,840 females averaging 25.58 years.
  • The authors caution that pretrained models can be error-prone across ethnicities and invite re-auditing with more ethical tools.
  • Pearson r = 0.973(0.0) indicates strong agreement between DEX and InsightFace on class-wise person counts.
  • Pearson r = 0.723(0.0) for gender skewness and Pearson r = 0.567(0.0) for age estimates indicate weaker cross-model agreement.

B.2 NSFW scoring aided misogynistic imagery hand-labeling

The authors use NSFW model scores to narrow ImageNet classes for human review of misogynistic and potentially pornographic imagery, while also examining semanticity, gender, and child-image prevalence.

  • Four previously identified categories—beach-voyeur photography, upskirt images, verifiably pornographic images, and exposed private parts—anchor the hand-labeling effort.
  • The NSFW-Mobilenet-v2 score sums the softmax values for hentai, porn, and sexy, then averages scores across each image class.
  • The authors compare class-wise mean NSFW scores with DEX-derived mean-gender scores and identify five natural clusters using Affinity Propagation.
  • Class semanticity is represented with 300-dimensional GloVe embeddings and reduced to two or three dimensions using UMAP for scatter-plot analysis.
  • The audit also reports infant and child images across 30 ImageNet classes, with particularly high infant-image density in bassinet, cradle, crib, and bib.

B.5 Blood diamond effect in models trained on this dataset

The paper argues that problematic ImageNet seed images can propagate into neural artworks, creating a moral analogue to blood diamonds in which harmful origins are obscured by the final product.

  • The authors frame neural art made from problematic dataset images as analogous to blood-diamond jewelry, because ill-considered seed images can affect the resulting artwork.
  • Their concern centers on datasets containing pornographic, non-consensual, voyeuristic, and underage-nudity images used as model-training material.
  • In a GanBreeder example, volunteers did not recognize the problematic seed classes from the generated artwork’s appearance.

B.6 Error analysis

The error analysis tests whether class-wise recognition accuracy varies with asymmetric human co-occurrence between training and validation images, finding a significant decrease for the most imbalanced classes.

  • The authors sort 1,000 classes by human-delta, defined from differences in the proportions of images containing people between validation and training sets.
  • The comparison uses ResNet50 and NasNet inference results against the general population of ImageNet classes.
  • Top-5 accuracy falls significantly for the top-25 human-delta classes, with T-test values between −3.87 and −3.06.
  • The authors present these resources as motivation for further investigation rather than claiming human-delta is the primary cause of cultural change.
  • They acknowledge methodological and dataset-context critiques, including reliance on pretrained gender classifiers and risks of deanonymization.

C.2 Bluewashing of AI ethics and revisiting the enterprise of Big data

The paper questions whether Big Data practices adequately protect marginalized communities and criticizes symbolic ethics responses that leave harmful dataset practices insufficiently examined.

  • The authors question whether Big Data can serve marginalized communities disproportionately affected by algorithmic injustice, especially where people have little agency or recourse.
  • They argue that collective silence and practices such as ethics shopping, bluewashing, lobbying, dumping, and shirking cause harm.
  • The paper calls for informed consent to be treated rigorously rather than bypassed through the creative-commons loophole.
  • It highlights non-consensual images and hidden problems in categorizing people as societal and ethical implications requiring continued machine-learning discussion.
Loading 2006.16923v2…