Source-linked AI summary
Benchmarking Single Image Dehazing and Beyond
Boyi Li, Wenqi Ren, Dengpan Fu, Dacheng Tao, Dan Feng, Wenjun Zeng, Zhangyang Wang
TL;DR
Single-image dehazing lacks large-scale benchmarks and evaluation criteria that reflect both human perception and machine-vision effectiveness. The paper introduces RESIDE and RESIDE-β, combining diverse synthetic and real-world data with full-reference, no-reference, subjective, and task-driven evaluations. Experiments reveal that no single dehazing model is best across all criteria and motivate evaluation beyond PSNR and SSIM.
Problem
Single-image dehazing lacks large-scale benchmarking, while PSNR and SSIM inadequately capture human perceptual quality and machine-vision effectiveness.
Method
The paper constructs RESIDE and RESIDE-β and systematically evaluates nine state-of-the-art algorithms using objective, subjective, perceptual, and task-driven criteria.
Results
No single dehazing model is best across all criteria: different methods lead on reference metrics, no-reference metrics, perceptual loss, subjective quality, detection, or efficiency.
Takeaways & Limitations
Dehazing should be evaluated and optimized using dedicated perceptual and high-level task criteria rather than solely PSNR and SSIM.
Abstract
from arXiv · showhide
We present a comprehensive study and evaluation of existing single image dehazing algorithms, using a new large-scale benchmark consisting of both synthetic and real-world hazy images, called REalistic Single Image DEhazing (RESIDE). RESIDE highlights diverse data sources and image contents, and is divided into five subsets, each serving different training or evaluation purposes. We further provide a rich variety of criteria for dehazing algorithm evaluation, ranging from full-reference metrics, to no-reference metrics, to subjective evaluation and the novel task-driven evaluation. Experiments on RESIDE shed light on the comparisons and limitations of state-of-the-art dehazing algorithms, and suggest promising future directions.
I. INTRODUCTION
Single-image dehazing addresses visibility degradation caused by haze, using the atmospheric scattering model to recover haze-free scene radiance from one hazy image. Existing methods estimate atmospheric light and transmission through physical priors or data-driven models.
- Haze reduces visibility, contrast, surface clarity, and color fidelity, making single-image dehazing a challenging restoration problem.
- Single-image dehazing is practically important because haze can obstruct cameras and hinder vision systems such as autonomous driving.
- The atmospheric scattering model represents a hazy image using clean radiance, transmission, and atmospheric light.
- Transmission depends on atmospheric scattering and object-camera distance, while the clean image is recovered by re-writing the model.
- Most contemporary methods estimate atmospheric light and transmission using physically grounded priors or data-driven approaches, with deep learning improving top-method performance.
B. Existing Methodology: An Overview
Existing dehazing methods estimate physical parameters with priors, CNNs, or unified formulations, but the field lacks a sufficiently broad benchmark and evaluation framework. RESIDE addresses these gaps with diverse data and multiple evaluation criteria.
- Existing methodology: Most dehazing pipelines estimate transmission, atmospheric light, and clean radiance in three stages, with primary attention usually given to transmission estimation.
- Existing methodology: Prior-based methods use image statistics and structural constraints, but their assumptions can fail when scene objects resemble atmospheric light.
- Existing methodology: CNN-based methods learn transmission directly from data, while end-to-end models can generate the clean image without intermediate parameter estimation.
- Evaluation gap: The field lacks large-scale benchmarking, and PSNR and SSIM inadequately characterize human perceptual quality or machine-vision effectiveness.
- Proposed benchmark: RESIDE combines large-scale data, diverse evaluation strategies, and systematic experiments comparing nine state-of-the-art algorithms.
II. DATASET AND EVALUATION: STATUS QUO
Before RESIDE, dehazing benchmarks were limited by small scale, synthetic data, missing real-world references, and metrics poorly aligned with practical quality. These limitations motivated broader objective and subjective evaluation.
- Benchmark status quo: Large-scale dehazing benchmarks were missing because realistic hazy images with clean ground truth are difficult to collect or create.
- Training data: Existing training datasets were often small or synthetic, with prior examples ranging from 12 to 25,000 images.
- Testing data: Testing sets were mostly synthetic with known ground truth, despite some visual evaluation on real hazy images.
- Evaluation criteria: PSNR and SSIM rely on clean references and may have limited practical relevance when synthetic and real hazy-image content diverges.
- Evaluation criteria: Objective metrics often align poorly with human visual quality, motivating subjective user studies because algorithmic differences can be visually subtle.
- Benchmark status quo: Existing datasets were generally too small or lacked sufficient real-world images and annotations for diverse evaluations.
III. A NEW LARGE-SCALE DATASET: RESIDE
RESIDE is a large-scale benchmark designed to compare single-image dehazing algorithms fairly across diverse data and evaluation viewpoints. Its training and testing components combine synthetic generation with objective and subjective assessment.
- Dataset and evaluation: RESIDE provides a large-scale benchmark with full-reference, no-reference, subjective, and task-driven evaluation options.
- Training data: The training set contains 13,990 synthetic hazy images generated from 1,399 clear indoor images.
- Training data: Each clear image generates multiple hazy counterparts under varied atmospheric-light and scattering-coefficient parameters, preserving clean-hazy pairs.
- Testing data: The testing set includes SOTS and HSTS, which are designed to provide different objective and subjective evaluation viewpoints.
B. Evaluation Strategies
The evaluation strategy combines full-reference, no-reference, and subjective measures to address the practical and perceptual limits of conventional dehazing evaluation.
- PSNR and SSIM are limited because clean ground truth is usually unavailable in practice and their scores may poorly align with human perception.
- The study applies PSNR, SSIM, SSEQ, and BLIINDS-II on SOTS and HSTS, comparing objective rankings with subjective ratings.
- The subjective study uses pairwise comparisons and separates perceptual quality into Clearness and Authenticity.Clearness measures haze removal, while Authenticity measures how realistic the dehazed image looks.
- Bradley-Terry modeling converts pairwise survey outcomes into subjective scores for ranking dehazing algorithms.
- Subjective evaluation is not automatically scalable to new results, motivating analysis of its correlation with objective metrics and periodic leaderboard reviews.
A. Objective Comparison on SOTS
Objective and subjective comparisons reveal that dehazing performance depends on the metric, haze density, image realism, and evaluation dimension rather than one universal ranking.
- Objective Comparison on SOTS: Learning-based methods outperform earlier prior-based methods in most SOTS cases for PSNR and SSIM.DehazeNet achieves the highest PSNR, while GRM achieves the highest SSIM with AOD-Net and DehazeNet close in SSIM.
- Objective Comparison on SOTS: No-reference rankings are less consistent: AOD-Net leads BLIINDS-II indoors, FVR leads SSEQ, and NLD remains competitive on both metrics.
- Objective Comparison on SOTS: DehazeNet is best for light and medium haze, whereas GRM achieves the highest PSNR and SSIM for thick haze.
- Subjective Comparison on HSTS: On synthetic HSTS images, DCP leads Clearness and DehazeNet leads Authenticity; on real images, CNN-based methods occupy the top three, with MSCNN best on both.
- Subjective Comparison on HSTS: Clearness and Authenticity are hardly correlated on synthetic images but correlate better on real images, reflecting multifaceted subjective quality.
- Subjective Comparison on HSTS: Subjective and objective evaluations diverge: MSCNN is the best subjective performer despite low synthetic-indoor PSNR/SSIM and moderate SSEQ/BLIINDS-II results.
C. Running Time
The running-time comparison evaluates per-image processing on synthetic indoor SOTS images and identifies AOD-Net as the most efficient method.
- Running time is measured per image on 620 × 460 synthetic indoor SOTS images using a 3.6 GHz CPU and 16G RAM.Implementations use MATLAB except AOD-Net, which uses Pycaffe.
- AOD-Net has a clear efficiency advantage over the other methods because of its lightweight feed-forward structure.
V. WHAT ARE BEYOND: FROM RESIDE TO RESIDE-β
RESIDE-β extends RESIDE toward practical outdoor dehazing by addressing mismatched training content and evaluation criteria. It introduces realistic outdoor training data and task-oriented directions for machine vision.
- RESIDE-β is an exploratory supplement addressing hurdles in training-data content and evaluation criteria.The authors characterize it as a beta-stage effort intended to inspire follow-up work.
- Indoor versus Outdoor Training Data: Synthetic training data often comes from indoor scenes, although dehazing is applied to outdoor environments.This creates a mismatch between training content and real application subjects.
- Indoor versus Outdoor Training Data: Depth estimation produces more visually plausible outdoor hazy images than Make3D-based synthesis in the authors’ comparison.The comparison is shown using examples in Figure 5.
- Indoor versus Outdoor Training Data: OTS contains 2,061 clean outdoor images and 72,135 paired synthetic hazy images generated across seven β values and five A values.The set uses estimated depth and is intended for training; visual inspection found most generated images free of noticeable artifacts.
- Indoor versus Outdoor Training Data: Including OTS generally preserves SOTS PSNR/SSIM performance while improving visual-quality generalization to real-world images.The authors report this pattern from preliminary experiments.
B. Restoration versus High-Level Vision
Dehazing for downstream vision should be evaluated by semantic-task utility rather than only restoration or perceptual quality. The paper compares perceptual loss and finds partial alignment with some quality measures but weak alignment with clearness and several other metrics.
- Task-driven dehazing should optimize utility for the downstream semantic task, not merely pixel-level or perceptual-level image quality.The motivation concerns images subsequently used for recognition and detection.
- Perceptual loss compares clean and dehazed images using Euclidean distances between features at relu2_2, relu3_3, relu4_3, and relu5_3.The metric is full-reference and is computed on SOTS.
- DehazeNet and CAP consistently achieve the lowest perceptual-loss differences on SOTS and HSTS.Their results generally align with PSNR, but not SSIM or the two no-reference metrics.
- On HSTS, perceptual loss correlates somewhat with authenticity but hardly with clearness.The authors suggest realistic visual appearance may better preserve semantic similarity than thorough haze removal.
2) No-Reference Task-driven Comparison on RTTS:
RTTS enables no-reference, task-driven comparison of dehazing methods on real-world hazy images. Detection rankings vary across models, but MSCNN, BCCR, and DCP show the strongest overall performance and correlate weakly with no-reference scores.
- The task-driven evaluation applies pretrained FRCNN, YOLO-V2, SSD-300, and SSD-512 models and ranks dehazing algorithms by mAP.The detection models operate on dehazed real-world images.
- RTTS contains 4,322 annotated real-world hazy images, 41,203 bounding boxes, and five traffic-related object categories.It primarily covers traffic and driving scenarios and is organized like VOC2007.
- MSCNN, BCCR, and DCP are the top three choices favored by detection tasks on RTTS overall.The tendency is not perfectly consistent across the four detection models.
- Detection mAP rankings show only weak correlation with no-reference results on RTTS.For example, BCCR has the highest BLIINDS-II value, while MSCNN has lower SSEQ and BLIINDS-II scores.
- RTTS fixes FRCNN for fair comparison, whereas earlier joint dehazing-detection work trained on annotated synthetic hazy images.The paper distinguishes these evaluation scopes while identifying joint optimization as a future direction.
VI. CONCLUSIONS AND FUTURE WORK
The evaluation finds no single best dehazing model across restoration, perceptual, subjective, efficiency, and detection criteria. The authors therefore emphasize real-world generalization and task-specific evaluation beyond PSNR and SSIM.
- No single dehazing model performs best across all criteria: strengths differ among PSNR/SSIM, no-reference metrics, perceptual loss, subjective quality, detection, and efficiency.AOD-Net and DehazeNet lead PSNR/SSIM, while MSCNN leads subjective quality and real-image detection.
- Deep learning methods are advantageous under PSNR and SSIM but may not generalize well to real-world hazy images.The authors also caution that PSNR and SSIM do not necessarily reflect human perceptual quality.
- Classical prior-based methods seem more favored by human perception, while typical MSE training can over-smooth visual details.The authors relate this difference to priors emphasizing illumination, contrast, or edge sharpness.
- MSCNN is endorsed by RTTS detection results, consistent with the use of multi-scale features in object detection.The paper presents this as an empirical alignment rather than a universal guarantee.
- Future dehazing research should evaluate and optimize subjective visual quality and high-level task performance rather than relying solely on PSNR and SSIM.The authors report that these traditional metrics align poorly with other evaluation criteria.