Source-linked AI summary
Deep learning-based transformation of the H&E stain into special stains
Kevin de Haan, Yijie Zhang, Jonathan E. Zuckerman, Tairan Liu, Anthony E. Sisk, Miguel F. P. Diaz, Kuang-Yu Jen, Alexander Nobori, Sofia Liou, Sarah Zhang, Rana Riahi, Yair Rivenson, W. Dean Wallace, Aydogan Ozcan
TL;DR
Additional special stains can improve evaluation of non-neoplastic kidney disease but require extra histochemical processing and waiting. The paper uses supervised deep learning to transform H&E images into PAS, Masson’s Trichrome, and Jones stains using registered training pairs and evaluates the approach in 58 kidney cases. The authors report improved preliminary diagnosis, stain quality statistically equivalent to standard histochemical staining, and transformation within one minute or less per core slide.
Problem
Special stains support evaluation of non-neoplastic kidney diseases, but traditional preparation requires additional tissue processing, time, cost, and laboratory resources.
Method
A supervised stain-transformation network converts H&E images into PAS, Masson’s Trichrome, and Jones stains using perfectly registered virtual-staining pairs with style-transfer augmentation.
Results
The approach improved preliminary diagnosis in several non-neoplastic kidney diseases, while generated special-stain quality was statistically equivalent to standard histochemical staining.
Takeaways & Limitations
Computational special stains may reduce slide-preparation time and laboratory costs while preserving tissue and supporting preliminary diagnosis when additional stains are needed.
Takeaways & Limitations
The proof-of-concept was trained on H&E stains from a few institutions and one scanner model, so additional data are needed for broader generalization and diagnostic validation.
Abstract
from arXiv · showhide
Pathology is practiced by visual inspection of histochemically stained slides. Most commonly, the hematoxylin and eosin (H&E) stain is used in the diagnostic workflow and it is the gold standard for cancer diagnosis. However, in many cases, especially for non-neoplastic diseases, additional "special stains" are used to provide different levels of contrast and color to tissue components and allow pathologists to get a clearer diagnostic picture. In this study, we demonstrate the utility of supervised learning-based computational stain transformation from H&E to different special stains (Masson's Trichrome, periodic acid-Schiff and Jones silver stain) using tissue sections from kidney needle core biopsies. Based on evaluation by three renal pathologists, followed by adjudication by a fourth renal pathologist, we show that the generation of virtual special stains from existing H&E images improves the diagnosis in several non-neoplastic kidney diseases sampled from 58 unique subjects. A second study performed by three pathologists found that the quality of the special stains generated by the stain transformation network was statistically equivalent to those generated through standard histochemical staining. As the transformation of H&E images into special stains can be achieved within 1 min or less per patient core specimen slide, this stain-to-stain transformation framework can improve the quality of the preliminary diagnosis when additional special stains are needed, along with significant savings in time and cost, reducing the burden on healthcare system and patients.
Introduction
Special stains provide complementary tissue contrast but add time, cost, and repeated sectioning to pathology workflows. This study presents supervised deep learning to transform existing H&E whole-slide images into special stains for kidney disease evaluation.
- Special stains complement H&E by highlighting different tissue constituents, including connective tissue, but are standard of care for certain non-neoplastic diseases.
- Traditional workflows require separate sectioning and chemical staining procedures for each additional stain, increasing preparation time, cost, and diagnostic delay.
- Virtual staining can reduce physical processing and enable multiple computational stains from a single tissue section, but stain transformation instead operates on an already stained H&E whole-slide image.
- The proposed framework uses supervised learning with spatially registered image pairs, avoiding unpaired data and distribution-matching losses for stain-to-stain transformation.
- Three special stains—PAS, Masson’s Trichrome, and Jones methenamine silver—were computationally generated from H&E sections and evaluated in kidney tissues from 58 unique patients.
- The network was augmented with eight style-transfer networks to address staining and scanner variability, then tested on previously unseen digitized H&E slides at approximately 1.5 mm^2/s.A typical needle-core kidney biopsy slide required approximately 0.5–1 minute for transformation.
Evaluation of stain transformation networks for kidney disease diagnoses
Across 58 kidney cases, computationally generated special stains improved preliminary diagnoses over H&E alone, with performance comparable to histochemically stained special stains. The examples show both diagnostic gains from enhanced structural contrast and discordances caused by interpretation or virtual-stain misrepresentation.
- Study design: 58 unique H&E tissue sections from non-neoplastic kidney diseases were reviewed by three pathologists before and after adding three stain-transformed special stains.The study used a blinded design with a washout period between evaluations.
- Virtual-stain evaluation: 13 improved diagnoses (22.4%), 38.3 concordant diagnoses (66.1%) and 6.7 discordant diagnoses (11.5%) occurred on average across 58 cases with virtual special stains.Ten cases improved for at least two pathologists, while three cases were discordant for more than one pathologist.
- Histochemical comparison: 15 improved diagnoses (25.8%), 38.6 concordant diagnoses (66.6%) and 4.3 discordant diagnoses (7.4%) occurred with H&E plus histochemically stained serial-section special stains.Twelve cases improved for at least two pathologists, while two cases were discordant for more than one pathologist.
- Statistical comparison: Virtual special stains improved diagnoses over H&E alone (P=0.0095), while histochemically stained special stains also improved diagnoses (P=0.0003).Differences between the virtual-stain and histochemical comparisons were not statistically significant for the three pathologists, with P values of 0.60, 0.34, and 0.92.
- Diagnostic examples: Enhanced basement-membrane contrast helped pathologists localize inflammatory cells and characterize rejection more precisely in an example where H&E alone provided insufficient contrast.A generated JMS also revealed basement-membrane changes characteristic of membranous nephropathy that were appreciated only after reviewing the transformed stain.
- Discordances: Discordances reflected either pathologist interpretation error or likely virtual-stain misrepresentation, including unusually pale fibrin thrombi on transformed PAS and overly dark amyloid on transformed JMS.In both cited misrepresentation cases, two of three pathologists made concordant diagnoses.
Evaluation of the quality of stain-transformed special stain images
A separate reader study compared computationally generated and histochemically stained special-stain images across multiple quality dimensions. The measured quality differences were smaller than rating variability, supporting equivalent stain quality in this evaluation.
- Evaluation design: Three pathologists rated 16 unique rectangular fields of view from validation slides for computationally generated and histochemically stained special stains.The fields ranged from approximately 150 μm ×175 μm to 375 μm ×500 μm.
- Scoring criteria: 2304 unique assessments rated stain quality from 1 to 4, where 4 was perfect and 1 was not acceptable.Masson’s trichrome included overall, nuclear, cytoplasmic, and extracellular-fibrosis quality; PAS and Jones Silver included overall, nuclear, cytoplasmic, and basement-membrane detail.
- Results: For all measured stain aspects, the quality difference was significantly smaller than the standard error between ratings.The authors interpret this result as indicating equivalent quality between stain-transformed and histochemically stained tissue.
Discussion
The stain-transformation framework uses perfectly registered virtual-stain pairs to reduce chemical processing while preserving structural fidelity. It shows practical benefits for speed, cost, tissue preservation, and stain consistency, but remains limited by scanner and staining-domain scope and requires larger validation studies.
- Method and advantages: Perfectly registered virtual-stain pairs provide a structural fidelity constraint and avoid stain-to-stain misalignments during supervised training.The shared autofluorescence source enables precisely matched images across stains, supporting transformation reliability and accuracy.
- Method and advantages: The framework reduces chemical processing by avoiding de-staining and re-staining, while preserving tissue for subsequent analysis.These benefits accompany decreased slide preparation time and laboratory costs.
- Method and advantages: CycleGAN-based style augmentation helped the transformation networks generalize across slides with varying H&E distributions, while virtual PAS outputs showed little variation.The augmentation networks can also be expanded using existing H&E image databases.
- Method and advantages: The stain-transformation approach avoids the hallucinations observed when CycleGAN architectures are directly applied to stain transformation, particularly for PAS and Jones silver stains.Those hallucinations incorrectly labeled tubular basement membranes and brush borders.
- Limitations: The current network is trained on H&E stains from a few institutions and microscopes of one vendor and remains a proof-of-concept requiring larger training and test datasets.Generalization to different microscope specifications, vendors, or substantially different staining procedures requires additional data.
- Limitations: The study excluded immunofluorescence and electron microscopy while isolating standard light microscopy, although these modalities could provide further confirmation and safety in clinical use.Their application is described as supporting the final pathological diagnosis rather than replacing the stain-transformation technique.
- Practical implications: H&E-to-special-stain transformation was prioritized because H&E covers approximately 80% of human tissue staining procedures and can run at 1.5 mm^2/s on a consumer-grade desktop computer with two GPUs.The method also saves labor, time, and chemicals and could be extended to other stain-to-stain transformations.
Training of stain transformation network
The stain transformation networks use GANs with a generator for image transformation and a discriminator for matching generated images to ground-truth stain distributions. Training combines spatial-color accuracy, noise regularization, adversarial learning, and extensive patch-based augmentation.
- Network and objective: GAN training pairs a generator that transforms input images with a discriminator that evaluates whether outputs match ground-truth stained-image distributions.The generator uses x_input to produce transformed images, while the discriminator distinguishes generated from ground-truth images.
- Network and objective: The generator objective combines L1 accuracy, total-variation regularization, and discriminator loss.L1 preserves spatial and color accuracy, whereas total variation reduces noise introduced by adversarial training.
- Architecture: The generator uses a modified U-net, while the discriminator uses a VGG-style network with five convolutional blocks and a sigmoid real-image probability output.The U-net contains four up-blocks and four down-blocks with skip connections; the discriminator progressively increases channels and reduces its output to one value.
- Optimization: Training used Adam with discriminator and generator learning rates of 1×10^-5 and 1×10^-4, respectively, initially training the generator seven times per discriminator iteration.The generator-to-discriminator training ratio decreases over time to a minimum of one discriminator iteration per three generator iterations.
- Training data: The networks learned from approximately 7,836 randomly cropped 256×256-pixel patches derived from 1,013 images across 10 tissue sections, with 76 validation images.The approach also used eight data-augmentation networks, random rotations, and flipping.
Image data acquisition
The study used stained kidney needle-core biopsy tissue imaged by microscopy and whole-slide scanning. Histochemical staining was performed at multiple pathology facilities for H&E, Masson's Trichrome, PAS, and Jones silver stain.
- Tissue and imaging: All neural networks were trained on microscopic images of thin tissue sections from needle-core kidney biopsies.Unlabeled tissue sections came from existing UCLA Translational Pathology Core Laboratory specimens under UCLA IRB 18-001029.
- Histochemical staining: H&E, Masson's Trichrome, and PAS staining were performed at UC San Diego, while Jones silver staining was performed at Cedars-Sinai.The stained slides were digitized for use in the study.
- Tissue and imaging: Stained kidney needle-core biopsy whole-slide images were acquired using Aperio AT2 slide-scanning microscopes.
Image co-registration
Autofluorescence and stained brightfield images were co-registered through progressively refined matching to achieve subpixel accuracy. The autofluorescence images were then normalized using tissue-area mean and standard deviation.
- Registration: Co-registration began with cross-correlation matching and was progressively refined until subpixel-level accuracy was achieved.The process extracted the most similar portions of the autofluorescence and stained images before further refinement.
- Normalization: Autofluorescence images were normalized by subtracting the tissue-area mean and dividing by the tissue-area standard deviation.
Class conditional virtual staining of label-free tissue
A class-conditional GAN generated virtual H&E and special-stain images from a shared information source, enabling perfectly matched training pairs. A digital staining matrix selected the target stain during testing.
- Class-conditional generation: A class-conditional GAN generated both input and ground-truth images for training the stain transformation networks.Using one network ensured that virtual H&E and special-stain images were automatically registered to the same source information.
- Stain conditioning: The digital staining matrix was concatenated to the generator and discriminator inputs to define stain coordinates within each image field.
- Stain conditioning: During testing, one-hot encoding allowed the network to generate separate H&E and corresponding special-stain images for each field of view.
- Architecture: Network channels were doubled relative to the stain transformation architecture to support the larger dataset and two distinct stain transformations.
- Training data: Four adjacent tissue sections trained the virtual staining networks, using H&E images from 10 patients and stain-specific datasets from approximately 10–11 patients.The passage reports 1,058 H&E images, 946 PAS images, 816 Jones images, and 966 Masson's Trichrome images, each sized 1424×1424 pixels.
Style transfer for H&E image data augmentation
CycleGAN style transfer augments H&E training data to support stain transformation across images from different laboratories or hospitals.
- Style transfer for H&E image data augmentation: CycleGAN style transfer maps virtually stained H&E images between domains representing different laboratory or hospital image styles.The model learns mappings G: X→Y and F: Y→X, using samples x and y from the two domains.
- Style transfer for H&E image data augmentation: Adversarial losses match generated images to the target stain style, while cycle consistency losses constrain the learned mappings.The generator loss combines adversarial and cycle consistency terms.
- Style transfer for H&E image data augmentation: The augmentation networks use U-net generators with three down-blocks and three up-blocks, plus four-block discriminators.The generators share block designs with the stain transformation network, while the discriminators use four rather than five blocks.
- Style transfer for H&E image data augmentation: Training used Adam with learning rates of 2×10^-5 for both generators and discriminators, one generator iteration per discriminator step, and batch size 6.These settings were used during CycleGAN training.
Training of single-stain virtual staining networks
Separate single-stain networks were trained alongside a multi-stain neural network. These networks generated rough virtual stains to enable elastic co-registration.
- Separate networks were trained to generate each individual virtual stain.
- The single-stain networks supported rough virtual staining for elastic co-registration.
- Their training followed procedures outlined in Rivenson et al.7.
Implementation details
The study implemented the networks in a specified software and hardware environment and evaluated diagnostic changes across staged pathologist studies, including 58 cases.
- Implementation details: The networks were implemented in Python 3.6.2 with TensorFlow 1.8.0 on a Windows 10 computer with two Nvidia GeForce GTX 1080 Ti GPUs.The system also had 64GB of RAM and an Intel I9-7900X CPU.
- Implementation details: An initial feasibility study compared diagnoses from H&E alone with diagnoses from H&E plus stain-transformed special stains across 16 non-neoplastic kidney cases.Three board-certified renal pathologists reviewed the cases.
- Implementation details: After a greater-than-three-week washout period, pathologists repeated diagnostic assessments using H&E together with computationally generated special stains.The review used the same clinical history and H&E images, with additional transformed stains in the second round.
- Implementation details: An adjudicator classified changes between diagnostic rounds as concordance, discordance, or improvement.The adjudicator was not among the three diagnosticians.
- Implementation details: The expanded study repeated the procedure across 58 cases using an online viewer that allowed pathologists to switch among cases and stains.Patient history, whole-slide images, and applicable stain options were presented through the custom server.
- Implementation details: A second washout was followed by review of histochemically stained H&E and special-stain slides from serial tissue sections.Two preliminary-study cases were excluded from the final analysis because complete special-stain whole-slide images were unavailable.
- Implementation details: Pathologist 2 was replaced in the expanded study because of time availability, with the initial diagnoses documented separately.Pathologists’ diagnoses and comments were provided in the Supplementary Data.
Statistical analysis
The analysis was powered for the expanded sample and tested whether stain-transformed or histochemical special stains improved diagnoses relative to H&E alone.
- Statistical analysis: 41 samples were calculated as necessary for statistical significance, so the study was expanded to 58 patients for adequate power.The calculation used power 0.8, alpha 0.05, and a one-tailed t-test.
- Statistical analysis: Across 58 cases, average diagnostic improvement with additional stain-transformed or histochemical stains was statistically significant.Scores were +1 for improvement, -1 for discordance, and 0 for concordance; significance corresponded to an average score greater than zero.
- Statistical analysis: A chi-squared test with two degrees of freedom compared proportions of improvements, concordances, and discordances between tested methods.Comparisons were also performed separately for each pathologist.
- Statistical analysis: P values of 0.05 or less were considered statistically significant.