Source-linked AI summary

The HAM10000 dataset, a large collection of multi-source dermatoscopic images of common pigmented skin lesions

Philipp Tschandl, Cliff Rosendahl, Harald Kittler

arXiv:1803.10417v3cs.CV

TL;DR

Automated diagnosis research lacks large, diverse dermatoscopic image collections and reliable multi-class data. The paper releases HAM10000 as a benchmark dataset with unified diagnostic classes for machine-learning research and human–machine comparisons, while noting that curated images may differ from real-world submissions.

  • Problem

    High-quality dermatoscopic images with reliable diagnoses are limited or restricted to a few disease classes, hindering neural-network training.

  • Method

    The dataset unifies diagnoses into seven generic classes, avoids ambiguous classifications, and incorporates images acquired with different cameras and modalities.

  • Results

    The released dataset covers more than 95% of pigmented lesions examined in daily clinical practice at the two study sites.

  • Takeaways & Limitations

    HAM10000 is intended as a benchmark for automated diagnosis and comparisons between human experts and machines.

  • Takeaways & Limitations

    Dataset images may differ from those end-users, especially lay persons, would provide in real-world scenarios.

Abstract

from arXiv · show

Training of neural networks for automated diagnosis of pigmented skin lesions is hampered by the small size and lack of diversity of available datasets of dermatoscopic images. We tackle this problem by releasing the HAM10000 ("Human Against Machine with 10000 training images") dataset. We collected dermatoscopic images from different populations acquired and stored by different modalities. Given this diversity we had to apply different acquisition and cleaning methods and developed semi-automatic workflows utilizing specifically trained neural networks. The final dataset consists of 10015 dermatoscopic images which are released as a training set for academic machine learning purposes and are publicly available through the ISIC archive. This benchmark dataset can be used for machine learning and for comparisons with human experts. Cases include a representative collection of all important diagnostic categories in the realm of pigmented lesions. More than 50% of lesions have been confirmed by pathology, while the ground truth for the rest of the cases was either follow-up, expert consensus, or confirmation by in-vivo confocal microscopy.

Background & Summary

Automated diagnosis research needs large, reliably annotated dermatoscopic datasets, but existing resources are small, diagnostically narrow, biased toward melanocytic lesions, or difficult to access. HAM10000 was released to broaden automated diagnosis research and support future human–machine comparisons.

  • Neural-network diagnostic algorithms require many annotated images, while high-quality images with reliable diagnoses remain limited or restricted to few disease classes.
  • Earlier public resources included PH2’s 200 images, comprising 160 nevi and 40 melanomas, with pathology confirmation for melanomas but not most nevi.
  • The commercially available Interactive Atlas dataset was diagnostically diverse but probably limited in use by constrained accessibility.
  • The ISIC archive was large, permissively licensed, and accessible, but 12893 of 13786 images were nevi or melanomas.
  • Past research therefore focused on melanoma-versus-nevus classification, while non-melanocytic pigmented lesions and reliable multiclass dermatoscopic prediction were underrepresented.
  • HAM10000 was released to boost automated dermatoscopic diagnosis research and provide a benchmark for comparing human experts with machines.

Methods

The HAM10000 training set combines images collected over 20 years from dermatology sites in Vienna, Austria, and Queensland, Australia.

  • 10015 dermatoscopic images were collected over 20 years from two sites in Austria and Australia.
  • The two sites used different storage formats and acquisition histories, including pre-digital image archives and heterogeneous metadata.

Extraction of images and meta-data from PowerPoint files

The dataset workflow extracted images and metadata from heterogeneous sources, filtered and reviewed records, standardized images, and unified diagnoses into seven generic classes.

  • Extraction of images and meta-data from PowerPoint files: Automated extraction processed PowerPoint files containing monthly clinical and dermatoscopic images, lesion identifiers, and associated text fields.
  • Digitization of diapositives: Scanned diapositives were digitized, cropped with centered lesions, and manually corrected for visual contrast and color reproduction.
  • Image standardization: MoleMax HD images were extracted from SQL tables, cropped to 800x600px, centered when necessary, and converted to quadratic pixels.
  • Image extraction and categorization: A screening method categorized more than 30000 images to separate dermatoscopic images from clinical close-ups and overviews.
  • Unifying pathologic diagnoses: Cases with uncertain or colliding histopathologic diagnoses were excluded, except melanomas associated with a nevus.
  • Unifying pathologic diagnoses: The workflow unified diagnoses into seven generic classes covering more than 95% of pigmented lesions examined at the two sites.
  • Final validation: Final manual validation removed identifiable, out-of-focus, artifact-obstructed, non-pigmented, ocular, subungual, and mucosal cases; remaining images received color or luminance correction when necessary.
  • Reproducibility: Custom code for the described processing methods was made available through the HAM10000_dataset GitHub repository.

Data Records

HAM10000 records were deposited in the Harvard Dataverse and made accessible through the ISIC archive, while combining data from distinct clinical sources and acquisition modalities.

  • Data Records: HAM10000 data records were deposited at the Harvard Dataverse, and Table 1 compares its diagnosis counts with existing databases.
  • Data access: Images and metadata were also made accessible through the public ISIC archive gallery and standardized API calls.
  • Rosendahl image set (Australia): The Australian collection originated from Cliff Rosendahl’s skin cancer practice and was acquired with DermLite Fluid or DermLite DL3 devices.
  • Dataset comparison: HAM10000 was presented alongside publicly available datasets in Table 1.
  • ViDIR image set (Austria): The dataset incorporated Austrian records from different times, including analog diapositives photographed with the Heine Dermaphot system.
  • ViDIR image set (Austria): Digital Austrian records included triplets of the same lesion at different magnifications to show local features and general patterns.
  • Ground truth and selection: Vienna follow-up data supported inclusion of nevi with more than 1.5 years of digital dermatoscopic follow-up, alongside selected benign and excised lesions.

Technical Validation

HAM10000 ground truth combines pathology, reflectance confocal microscopy, longitudinal follow-up, and expert consensus, with manual image corrections applied when needed for visual inspection.

  • Image validation: Manual histogram correction targeted underexposed images and visible yellow or green hues, shifting estimated illuminants toward blue and red.Corrections were evaluated using grey-world illuminant estimates in Lab-color space before and after adjustment.
  • Ground truth: Histopathologic diagnoses were reviewed for plausibility, with ambiguous or mismatched cases excluded when necessary.Specialized dermatopathologists supplied diagnoses, and available slides were scanned for later review.
  • Ground truth: Reflectance confocal microscopy provided near-cellular-resolution confirmation for some facial benign keratoses.Most such cases came from a prospective confocal study that included one year of follow-up.
  • Ground truth: Nevi unchanged across three follow-up visits or 1.5 years were accepted as biologically benign.This follow-up ground truth was applied to nevi but not to other benign diagnoses.
  • Ground truth: Typical benign lesions without pathology or follow-up received an expert-consensus label only when both authors independently agreed on an unequivocal diagnosis.These lesions were generally photographed for educational purposes and did not require biopsy or further follow-up.

Usage Notes

The dataset spans clinically distinct populations and diagnostic categories, while repeated lesion views provide augmentation but do not represent unique lesions and may differ from real-world user images.

  • Dataset populations: Austrian cases come from a tertiary referral center serving high-risk patients, whereas Australian cases come from primary care in a high-incidence, chronically sun-damaged population.The two source populations therefore reflect different clinical settings and patient characteristics.
  • Diagnostic categories: HAM10000 covers pigmented actinic keratoses and Bowen’s disease, including variants that commonly show surface scaling and limited pigment.Both lesion types are UV-associated, although some Bowen’s disease cases arise from human papilloma virus infection.
  • Diagnostic categories: Basal cell carcinoma is represented across flat, nodular, pigmented, and cystic morphologic variants and can grow destructively if untreated.It rarely metastasizes but remains clinically important because of destructive growth.
  • Diagnostic categories: Benign keratosis groups seborrheic keratoses, solar lentigines, and lichen-planus-like keratoses because they are biologically similar and often share histopathologic terminology.Lichen-planus-like keratoses can mimic melanoma dermatoscopically and are especially challenging.
  • Diagnostic categories: The dataset includes melanocytic nevi, multiple melanoma variants except non-pigmented, subungual, ocular, or mucosal forms, vascular lesions, and dermatofibromas.Vascular lesions range from cherry angiomas to angiokeratomas and pyogenic granulomas, with hemorrhage also included.
  • Usage considerations: Image counts exceed unique-lesion counts because lesions may be photographed at different magnifications, angles, or with different cameras.These repeated views are intended as natural data augmentation exposing models to general and local features.
  • Usage considerations: Dataset images may differ from images supplied by end users, while insufficiently magnified or out-of-focus images were removed.This limits direct equivalence between benchmark inputs and some real-world user-provided images.
  • Usage considerations: Automated screening, manual reviews, and EXIF-data removal were used to support anonymization to the authors’ best knowledge.The data collection received ethics approval from the Medical University of Vienna and the University of Queensland.

Additional Information

The article is openly documented and licensed, with citation and publisher information provided alongside the dataset description.

  • Declarations: The authors declare no competing interests.This statement appears in the article’s competing-interests disclosure.
  • Publication information: The paper is cited as Tschandl et al., “The HAM10000 dataset,” published in Scientific Data 5:180161 in 2018.The citation includes DOI 10.1038/sdata.2018.161.
  • Publication information: The publisher states neutrality regarding jurisdictional claims in published maps and institutional affiliations.This is a standard publisher’s note accompanying the article.
  • Access and licensing: The article is distributed under a Creative Commons Attribution 4.0 International License, subject to attribution and change-notice requirements.Third-party material may have separate licensing unless otherwise indicated.
  • Access and licensing: Metadata files are made available under the Creative Commons Public Domain Dedication waiver.The waiver is identified as CC0 1.0.
Loading 1803.10417v3…