Source-linked AI summary

Real-world Underwater Enhancement: Challenges, Benchmarks, and Solutions

Risheng Liu, Xin Fan, Ming Zhu, Minjun Hou, Zhongxuan Luo

arXiv:1901.05320v2cs.CV

TL;DR

Underwater enhancement research lacks sufficiently diverse real-world benchmarks and task-specific evaluation across visibility, color correction, and downstream recognition. This paper constructs and evaluates the RUIE benchmark across those dimensions, finding task-dependent algorithm performance and no strong correlation between image-quality assessment and detection accuracy.

  • Problem

    Existing underwater enhancement evaluations rely on varied datasets and metrics, many synthetic, with limited real-world coverage across visibility, color casts, and higher-level tasks.

  • Method

    The paper constructs RUIE from over 4,000 real-world sea images, divides it into three task-oriented subsets, and evaluates enhancement using quality metrics and object detection accuracy.

  • Results

    No single UIE algorithm performs best across all tasks and criteria, and image-quality assessment has no strong correlation with detection accuracy.

  • Takeaways & Limitations

    Underwater enhancement evaluation should consider visibility, color correction, and downstream detection together, motivating joint and task-specific approaches.

Abstract

from arXiv · show

Underwater image enhancement is such an important low-level vision task with many applications that numerous algorithms have been proposed in recent years. These algorithms developed upon various assumptions demonstrate successes from various aspects using different data sets and different metrics. In this work, we setup an undersea image capturing system, and construct a large-scale Real-world Underwater Image Enhancement (RUIE) data set divided into three subsets. The three subsets target at three challenging aspects for enhancement, i.e., image visibility quality, color casts, and higher-level detection/classification, respectively. We conduct extensive and systematic experiments on RUIE to evaluate the effectiveness and limitations of various algorithms to enhance visibility and correct color casts on images with hierarchical categories of degradation. Moreover, underwater image enhancement in practice usually serves as a preprocessing step for mid-level and high-level vision tasks. We thus exploit the object detection performance on enhanced images as a brand new task-specific evaluation criterion. The findings from these evaluations not only confirm what is commonly believed, but also suggest promising solutions and new directions for visibility enhancement, color correction, and object detection on real-world underwater images.

I. INTRODUCTION

Underwater image degradation limits computer-vision applications, while existing benchmarks and metrics inadequately cover real-world visibility, color casts, and task-specific utility. The paper addresses these gaps with the large-scale RUIE benchmark and systematic evaluations across enhancement objectives.

  • Degradation challenges: Underwater scattering causes low contrast and haze-like effects, while wavelength-dependent attenuation produces bluish or greenish color casts.Suspended particles absorb and scatter reflected light, and red light is absorbed more strongly than green and blue light.
  • Evaluation objectives: Enhancement algorithms target visibility, color-cast correction, and the accuracy of downstream detection or classification tasks.The third objective applies when enhancement serves as preprocessing for higher-level vision.
  • Benchmark gap: Existing evaluations rely on varied datasets and quality metrics, many using synthetic images, while real-world benchmark coverage remains insufficient.The paper motivates enriching real-world data for algorithm evaluation and training data-driven networks.
  • Benchmark gap: Current underwater datasets have limited suitability for visibility evaluation and insufficient diversity in scenes and tones.These limitations are especially relevant for evaluating model-driven methods and varied enhancement conditions.
  • Contributions: The authors construct RUIE from over 4,000 real-world sea images, dividing it into three diverse subsets targeting distinct enhancement objectives.The subsets address visibility quality, color casts, and higher-level detection or classification.
  • Contributions: Systematic experiments on RUIE evaluate algorithms across multiple degradation degrees and color-cast types, revealing both advantages and limitations.The study aims to confirm established observations and identify directions for underwater enhancement research.
  • Contributions: Object detection accuracy on enhanced images provides a task-specific evaluation criterion, but current quality metrics show no strong correlation with classification accuracy.This motivates considering enhancement and higher-level vision jointly rather than as independent cascaded processes.

II. UNDERWATER IMAGE ENHANCEMENT ALGORITHMS

Typical underwater image enhancement algorithms aim to improve human-perceived image quality from a single degraded input by increasing visibility or reducing color casts. They are categorized according to how they model the underwater imaging process.

  • Overview: Typical UIE algorithms enhance a single degraded input by increasing visibility or alleviating color casts caused during underwater image capture.These methods target image quality favored by human observers.
  • Overview: The paper categorizes UIE algorithms according to whether and how they model the underwater imaging process.This modeling distinction organizes the subsequent discussion of enhancement methods.

A. Model-free Methods

Model-free methods adjust observed pixel values without explicitly modeling image formation, using spatial- or transform-domain operations. They can improve visual quality but remain vulnerable to artifacts and distortions.

  • Method definition: Model-free methods adjust pixel values without explicitly modeling the image formation process, in either spatial or transform domains.Examples include histogram equalization, Gray World, CLAHE, multiscale retinex, white balance, and color-constancy methods.
  • Limitations: Spatial-domain methods can improve visual quality but may accentuate noise, introduce artifacts, and cause color distortions.These limitations arise from relying on observed image information rather than an explicit degradation model.
  • Limitations: Transform-domain methods can suppress noise but may suffer from low contrast, detail loss, and color deviations.The paper contrasts their noise behavior with their other quality degradations.

B. Model-based Methods

Model-based methods estimate physical imaging parameters from degraded observations and prior assumptions, then invert the degradation process to recover the clear scene. Their effectiveness depends on the validity of the assumed priors and environmental conditions.

  • Method definition: Model-based methods characterize underwater image formation, estimate its parameters, and invert the degradation process to restore the clear scene.The approach derives from the Jaffe-McGlamery model and explicitly represents physical imaging effects.
  • Imaging model: The restoration model uses observed image Ic(x), scene radiance Jc(x), atmospheric light Ac, and transmission tc(x) for each color channel.Atmospheric light and transmission are identified as the two critical restoration parameters.
  • Imaging model: Transmission represents the portion of scene radiance reaching the camera, while recent methods estimate atmospheric light and transmission to improve visibility and correct color casts.These estimates may be combined with model-free techniques such as color balance or histogram equalization.
  • Prior-based methods: Prior-based methods adapt dehazing assumptions to underwater attenuation, including modified dark channel priors and physical priors based on channel discrepancies or attenuation.The paper surveys several underwater-specific prior constructions.
  • Limitations: A common limitation is that physical priors become invalid under specific scenery configurations or severe color casts, including white objects or regions for DCP.Thus, model-based restoration depends on the environmental validity of its assumptions.

C. Data-driven Enhancement Neural Networks

Data-driven underwater enhancement networks draw on dehazing architectures, but underwater-specific factors motivate more complex designs and losses. A comprehensive benchmark is still needed to assess these methods across visibility, color, and higher-level vision objectives.

  • Dehazing CNN architectures motivate data-driven underwater enhancement because underwater and hazy images share similar imaging models.
  • Dynamic water flow, color deviations, and low illumination require more complex network structures and/or carefully designed loss functions.
  • Existing evaluations use varied data sets and metrics, often on synthetic images or isolated quality indices such as contrast, saturation, and luminance.
  • A comprehensive study is needed to determine whether UIE algorithms improve visibility, correct color casts, and support higher-level vision tasks.

III. THE PROPOSED DATA SET

The proposed RUIE benchmark addresses visibility, color-cast correction, and higher-level task evaluation with diverse real-world underwater images. It is built from multi-view sea-floor video capture and organized into task-specific subsets.

  • Dataset acquisition: RUIE uses a multi-view system with twenty-two waterproof cameras mounted along a 10 meters by 10 meters square frame.
  • Dataset acquisition: The cameras captured videos across varying scene depths, lighting periods, water depths, and natural marine conditions near Zhangzi island.
  • Dataset acquisition: More than 250 hours of video yielded about four thousand manually selected images spanning illumination, depth of field, blur, and color-cast diversity.
  • RUIE subsets: UIQS evaluates visibility using UCIQE-ranked images divided into five equally sized quality subsets, A through E, in descending score order.
  • RUIE subsets: UCCS contains 300 images divided among bluish, greenish, and blue-green tones to evaluate color-cast correction.
  • RUIE subsets: UHTS contains 300 sea-life images with labeled bounding boxes and three classes—scallop, sea cucumbers, and sea urchins—for detection and classification evaluation.

IV. EVALUATION RESULTS AND DISCUSSIONS

The study applies the RUIE benchmark to quantitative and qualitative evaluation of representative underwater image enhancement algorithms.

  • Eleven representative underwater image enhancement algorithms were evaluated quantitatively and qualitatively on the RUIE benchmark.

A. Comprehensive Image Quality Comparisons on UIQS

On UIQS, the study compares eleven enhancement methods using qualitative visibility judgments and non-reference quality metrics, revealing method-specific strengths and metric–perception discrepancies.

  • Qualitative comparisons: The comparison evaluates visibility enhancement across UIQS inputs spanning quality levels A through E.The five input quality levels are ordered A, B, C, D, and E; the study compares eleven methods qualitatively.
  • Qualitative comparisons: Most methods improve images in quality levels A–C, where underwater scattering is subtle.MSRCR, CLAHE, and Fusion show distinct trade-offs in color tone, saturation, contrast, brightness, and haze.
  • Qualitative comparisons: For severely degraded D–E images, DPATN, HLPcb, and BCCRcb remove haze-like effects, with BCCRcb best improving visibility and contrast.HLPcb improves brightness and recovers more image details, while BCCRcb performs best for visibility and contrast.
  • Quantitative comparisons: UIQM and UCIQE provide non-reference quality scores because real-world sea images lack ground-truth reference scenes.UIQM combines colourfulness, sharpness, and contrast; UCIQE combines chroma, saturation, and contrast.
  • Quantitative comparisons: DPATN, DCPcb, BCCRcb, MSRCR, CLAHE, and Fusion stably improve both comprehensive metrics across all five quality categories.The gains over other methods are more evident for lower-quality categories C, D, and E.
  • Summary: The prior-based BP algorithm suits less-degraded images, whereas Fusion, CLAHE, and prior-aggregated DPATN better handle severe degradation.This summary reflects the authors’ overall recommendation for different degradation levels.
  • Discussion: Quantitative scores can disagree with visual quality because UIQM and UCIQE emphasize low-level intensities while ignoring semantic perception and reasonable global intensity ranges.NOM can receive high UIQM scores despite severe reddish shifts, while UCIQE can favor unnatural excessive contrast.

B. Color Correction Comparisons on UCCS

UCCS evaluates color-cast correction across bluish, blue-green, and greenish images. Results show method-specific strengths, with DCPcb, BCCRcb, and DPATN better suited to bluish images, while model-free methods better handle greener scenes.

  • UCCS contains diverse color tones and evaluates UIE algorithms’ ability to correct color casts.The set includes bluish, greenish, and blue-green subsets, while Avga and Avgb quantify green-red and blue-yellow components.
  • MSRCR corrects greenish and bluish tones well, but often shifts Avga and Avgb positive, producing reddish results.
  • DCPcb, BCCRcb, and DPATN perform more satisfactorily on bluish images, whereas model-free methods are more suitable for images with stronger green components.
  • Model-based transmission-prior methods can struggle with color casts; BP performs poorly on greenish images, while color-balance cascades produce more appealing results.
  • Combining model-based visibility enhancement with color balance outperforms either module alone on UIQS, although gains vary by dehazing prior and degradation level.DCP and BCCR gain more from color balance, while CAPcb and HLPcb show only small advantages on severely degraded subsets.

C. Higher-level Task-driven Comparison on UHTS

The UHTS experiments evaluate UIE as preprocessing for marine object detection, revealing that improved image quality does not reliably translate into better detection. Detection gains vary substantially across algorithms and metrics.

  • The study applies a marine object classifier to images enhanced by eleven UIE algorithms and measures mean Average Precision (mAP) and detection number (Num).
  • Higher-level task-specific evaluation remains important because existing underwater benchmarks lacked marine-object labels for evaluating UIE algorithms.
  • Most UIE algorithms increase detection numbers but provide little significant mAP improvement; BCCRcb and HLPcb improve Num while stably increasing mAP across UHTS subsets.
  • CLAHE improves or maintains mAP while increasing Num, whereas Fusion increases Num but negatively affects mAP; NOM performs poorly, especially on subsets B and C.
  • The study also considers synthetic outdoor-object images generated with WaterGan and evaluates the original YOLO-V3 detector after enhancement.
  • mAP correlates weakly with no-reference quality metrics: CLAHE and Fusion can raise UCIQE/UIQM while lowering mAP, whereas UHP can improve mAP despite the worst quality scores.

V. CONCLUSIONS AND FUTURE WORK

The paper introduces RUIE as a three-subset real-world benchmark for visibility, color-cast, and higher-level detection/classification evaluation. Its conclusions emphasize weak alignment between image-quality assessment and detection accuracy, motivating improved assessment and task-specific models.

  • RUIE comprises UIQS, UCCS, and UHTS, targeting visibility degradation, color cast, and higher-level detection/classification, respectively.
  • No strong correlation exists between image-quality assessment and detection accuracy measured by mAP.
  • The authors call for better quality assessment, joint visibility-and-color-correction paradigms, and more accurate underwater object detection/classification networks.
Loading 1901.05320v2…