Source-linked AI summary

BIMCV COVID-19+: a large annotated dataset of RX and CT images from COVID-19 patients

Maria de la Iglesia Vayá, Jose Manuel Saborit, Joaquim Angel Montell, Antonio Pertusa, Aurelia Bustos, Miguel Cazorla, Joaquin Galant, Xavier Barber, Domingo Orozco-Beltrán, Francisco García-García, Marisa Caparrós, Germán González, Jose María Salinas

arXiv:2006.01174v3eess.IVcs.CVcs.LG

TL;DR

Public COVID-19 imaging studies were limited, often being proprietary, so the paper presents BIMCV COVID-19+ as an open dataset. It combines multimodal images with clinical metadata, radiological findings, and annotations, with a first iteration spanning 1,311 patients and multiple imaging types.

  • Problem

    Few publicly available COVID-19 image studies limited access to data for developing diagnostic and prognostic AI algorithms.

  • Method

    The paper constructs an open, multi-institutional dataset combining chest X-ray and CT images with reports, DICOM metadata, UMLS labels, diagnostic tests, and radiologist annotations.

  • Results

    1,380 CR, 885 DX, and 163 CT full-resolution images from 1,311 patients comprise the first iteration, including 2,429 image studies and 5,530 image series.

  • Takeaways & Limitations

    BIMCV COVID-19+ provides an open dataset with radiological findings and ROI annotations for studying COVID-19 imaging and clinical evolution.

  • Takeaways & Limitations

    The database is an initial iteration that will grow as images and annotations become available.

Abstract

from arXiv · show

This paper describes BIMCV COVID-19+, a large dataset from the Valencian Region Medical ImageBank (BIMCV) containing chest X-ray images CXR (CR, DX) and computed tomography (CT) imaging of COVID-19+ patients along with their radiological findings and locations, pathologies, radiological reports (in Spanish), DICOM metadata, Polymerase chain reaction (PCR), Immunoglobulin G (IgG) and Immunoglobulin M (IgM) diagnostic antibody tests. The findings have been mapped onto standard Unified Medical Language System (UMLS) terminology and cover a wide spectrum of thoracic entities, unlike the considerably more reduced number of entities annotated in previous datasets. Images are stored in high resolution and entities are localized with anatomical labels and stored in a Medical Imaging Data Structure (MIDS) format. In addition, 10 images were annotated by a team of radiologists to include semantic segmentation of radiological findings. This first iteration of the database includes 1,380 CX, 885 DX and 163 CT studies from 1,311 COVID-19+ patients. This is, to the best of our knowledge, the largest COVID-19+ dataset of images available in an open format. The dataset can be downloaded from http://bimcv.cipf.es/bimcv-projects/bimcv-covid19.

Background & Summary

BIMCV COVID-19+ addresses the limited public availability of COVID-19 imaging data by providing an open, multi-institutional dataset with radiological findings, reports, metadata, and annotations. Its first iteration combines multiple imaging modalities and repeated studies across patients.

  • Motivation: BIMCV COVID-19+ is an open, multi-institutional dataset intended to support research on COVID-19 detection and evolution.The authors emphasize that worldwide access may maximize the usefulness of the data.
  • Dataset scope: 1,380 CR, 885 DX, and 163 CT full-resolution images from 1,311 patients comprise the first iteration.The dataset is intended to grow incrementally as new images and annotations become available.
  • Dataset scope: Each study includes images, anonymized DICOM metadata, Spanish radiological reports, and UMLS concept identifiers associated with images.The identifiers are organized in semantic trees and include entities such as infiltrates and pleural effusion.
  • Annotations: Radiologists manually annotated COVID-19-related regions of interest in 10 images for pixel-level localization.These annotations support training semantic segmentation models and extracting lesion extent and exact location.
  • Related datasets: Earlier public datasets included limited patient counts, extracted CT images with suboptimal quality, or few COVID-19-positive images, while other datasets were private.The comparison is summarized through characteristics including image and patient counts, modality, reports, and public availability.
  • Novelty: The dataset is presented as the first COVID-19 dataset with radiological findings and, to the authors’ knowledge, the largest by images and patients.It also includes multiple samples per patient for analyzing clinical evolution.

Methods

The study describes data acquisition, anonymization, and labeling within an ethics-approved, HIPAA-compliant retrospective cohort framework. Publication of the open database required institutional and data-protection approvals.

  • Methods: The methodology covers data production, ethics, data anonymization, and labeling.
  • Ethics statement: The study received approval from the Miguel Hernandez University Institutional Review Board and a local institutional ethics committee.The retrospective cohort study was described as HIPAA-compliant.
  • Ethics statement: Healthcare authorities authorized publication after data-protection, confidentiality, impact-assessment, and security documentation was provided.

Data acquisition

Data were acquired through an open-source clinical-data infrastructure and identified from regional laboratory records. The dataset included all medical images available for subjects meeting SARS-CoV-2 test-positivity criteria during the specified period.

  • Acquisition infrastructure: R & D Cloud CEIB provided search, clinical-trial management, anonymization, and biomedical knowledge-engine functions using open-source technology.
  • Cohort identification: All consecutive subjects with at least one positive PCR or IgM, IgG, or IgA test between February 26 and April 18, 2020 were identified through regional laboratory records.
  • Cohort identification: All medical images acquired for eligible subjects during that period were included in BIMCV COVID-19+.The passage begins describing chest X-ray modalities included in the dataset.

Data Geo-positioning

The geographic analysis maps the number of CR, DX, and CT tests across Valencian health departments. In the first iteration, most represented departments were located in Alicante and Castellón, with future map completion planned.

  • Health departments: Valencian health departments are geographic healthcare demarcations that organize regional healthcare resources and services.
  • Geographic distribution: Figure 1 maps the number of CR, DX, and CT tests performed by health department.
  • Geographic distribution: 11 represented health departments were located mainly in Alicante and Castellón in the first dataset iteration.The authors planned to complete the map in future iterations.

Data anonymization

BIMCV COVID-19+ applies report, DICOM, and image-level anonymization procedures to protect patient confidentiality while preserving dataset utility.

  • Data anonymization: DICOM anonymization follows DICOM PS3.15 Annex E using a Clinical Trial Processor server.The Basic Profile removes patient, personnel, organization, identifying, demographic, temporal, and private attributes.
  • Data anonymization: Anonymization removes identifying information from Spanish radiological reports using the DisMed named-entity-recognition methodology.The method extracts and locates seven predefined entity categories in radiological reports.
  • Data anonymization: The anonymization workflow covers names, addresses, geographic locations, and other information that could reveal identities in reports or metadata.Selected entities are associated with Protected Health Information categories for systematic removal.
  • Data anonymization: RX, DX, and CT images required no deformation because patient data could not be reconstructed from their image information.Some chest X-rays were nevertheless visually inspected to remove or crop burnt-in personal information.

Image preprocessing

The dataset preserves image resolution while standardizing pixel representation and estimating radiographic projection with a neural network.

  • Image preprocessing: Raw DICOM pixels were rescaled using available window width and center values and stored as 16-bit PNG images without resizing.Avoiding resizing was intended to prevent loss of resolution.
  • Image preprocessing: An EfficientNet-based neural network estimated image projection as frontal or lateral.Frontal combines antero-posterior and postero-anterior views; the system does not distinguish erect position or left versus right lateral views.

Labeling

BIMCV COVID-19+ labels radiological findings and anatomical locations from reports, maps them to UMLS identifiers, and adds pixel-level ROI annotations for selected lesions.

  • Labeling: Radiological findings such as ground-glass opacity and consolidation are clinically relevant because their location and progression support diagnosis, prognosis, and decisions.These lesions commonly affect both lungs, especially lower and posterior peripheral or subpleural regions.
  • Labeling: Clinical reports, rather than images, supplied radiological findings, differential diagnoses, and localization labels.This separates report-derived labeling from subsequent image-based ROI annotation.
  • Labeling: A retrained bidirectional LSTM with attention automatically extracts 336 labels, including COVID-19 and COVID-19 uncertain, from preprocessed reports.The training corpus contains 23,439 manually annotated sentences, including 724 newly annotated COVID-19 sentences.
  • Labeling: Diagnoses, findings, and anatomical localizations are mapped to UMLS controlled-vocabulary CUIs and organized into semantic concept trees.Although reports remain in Spanish, CUI mapping makes the labels usable regardless of language.
  • Labeling: Ten images received radiologist-defined ROIs for frequent COVID-19-associated lesions, including infiltrates, ground-glass patterns, and consolidation.Eight radiologists annotated the ROIs at pixel level with XNAT OHIF Viewer, producing XML paths useful for semantic-segmentation training.

Code availability

The paper provides code for report-based annotation and a notebook for dataset statistics, alongside documentation of ROI findings.

  • Code availability: The report annotation pipeline, including preprocessing and the attention-based bidirectional LSTM classifier, is publicly available on GitHub.The code covers the automated extraction of labels from medical reports.
  • Code availability: A Python Jupyter notebook for acquiring dataset statistics is also available on GitHub.The notebook is hosted in the BIMCV-COVID-19 repository.
  • Code availability: Table 4 documents the findings annotated with ROIs using XNAT OHIF Viewer.These ROI annotations complement the released analysis and labeling resources.

Data Records

BIMCV COVID-19+ organizes imaging, diagnostic, report, terminology, and metadata records in the MIDS hierarchy. Each subject can have multiple studies and series, with radiological concepts and tests stored in linked files.

  • Data Records: MIDS extends BIDS to organize medical imaging and associated information in a simple hierarchical folder structure.It adapts DICOM organization using folder-based TSV and JSON files, with imaging stored in NIfTI or PNG formats.
  • Data Records: Diagnostic records include PCR, IgG, or IgM tests with positive, negative, or indeterminate outcomes, including repeated tests over time.The test list is stored in sil_reg_covid.csv.
  • Data Records: Figure 4 presents the dataset’s MIDS organization through a conceptual schema, a general template, and an example folder structure.The three views connect the abstract data model with its reusable organizational template and concrete storage layout.
  • Data Records: Radiological reports are anonymized and stored with automatically extracted UMLS concept identifiers for diseases and anatomical locations.The report labels and extracted concepts share the labels_COVID-19_posi.tsv file, while the terminology hierarchy is stored separately.
  • Data Records: Each subject may have multiple image studies, each containing multiple image series stored in the MIDS directory structure.Image data are extracted from DICOM and stored as .nii.gz files, while relevant series metadata are stored in same-named JSON files.

Technical Validation

Technical validation characterizes the dataset’s scale, modalities, acquisition variability, diagnostic-test timing, image quality, and automated report-labeling performance. The reported checks support both dataset description and validation of extracted COVID-19 labels.

  • Technical Validation: 1,311 subjects contributed 2,429 image studies and 5,530 image series, with 602 female patients and a mean age of 63.11 ± 16.75 years.The age distribution was reported as coherent with COVID-19+ demographics in Spain.
  • Technical Validation: 1.9 image studies per subject were recorded on average, supporting repeated observations across subjects.The distribution of studies per subject is shown in Figure 5.
  • Technical Validation: 1,380 CX, 885 DX, and 163 CT studies span multiple acquisition systems, producing modality and vendor variability.The dataset includes 2,427 chest X-rays, with 2,425 in monochrome 2 and 751 in monochrome 1 photometric interpretation.
  • Technical Validation: 135 studies were deemed suboptimal, while COVID-19, densities, pneumonia, consolidations, and infiltrates were the most frequent reported findings.The paper notes that these findings are closely related to COVID-19.
  • Technical Validation: F1-micro reached 0.922 for COVID-19 label validation using 2,343 manually labeled sentences, with further evaluation on an independent test sample.The independent sample included negative, uncertain, and affirmative COVID-19 mentions in a COVID-19-negative partition.
  • Technical Validation: Ten images received radiologist-marked regions of interest distinguishing ground-glass opacities from consolidations.Figure 6 uses green for ground-glass opacities and purple for consolidations.
  • Technical Validation: F1-weighted=0.9320, F1-micro=0.9378, and accuracy=0.8281 were obtained when evaluating all included entities on the independent test set.This evaluation covered entities related and unrelated to COVID-19 pneumonia.

Usage Notes

The dataset is intended for research and clinical-AI development, including radiologist training and decision-support methods. It is openly distributed through specified repositories and a web-based request process under licensing conditions.

  • Usage Notes: The dataset can support radiologist training and deep-learning methods using images and labels to assist radiologists’ decision-making.The stated uses include understanding lesions associated with this pathology and developing assistance tools.
  • Usage Notes: The data are freely available for research and may also be used commercially under certain conditions after accepting an End-User License Agreement.Access is provided through the BIMCV webpage.
  • Usage Notes: The dataset is stored in OSF and EUDAT repositories, including an OSF DOI and EUDAT infrastructure support.The repositories are identified as OSF in Germany and EUDAT through TransBioNet at Barcelona Supercomputing Center.
  • Usage Notes: A longitudinal series contains 15 image studies for the same subject between March 14 and March 30, 2020.This illustrates the availability of repeated studies for longitudinal analysis.
Loading 2006.01174v3…