Source-linked AI summary

Subjective and Objective Quality Assessment of Image: A Survey

Pedram Mohammadi, Abbas Ebrahimi-Moghadam, Shahram Shirani

arXiv:1406.7799v1cs.MMcs.CV

TL;DR

Image quality assessment is needed because processing stages introduce distortions and subjective evaluation, although reliable, is costly and difficult to deploy. This survey organizes subjective and objective IQA, emphasizes nine FR-IQA measures, and reviews HDR and 3-D assessment. It reports evaluations across subjective datasets and identifies scope boundaries for 3-D IQA and the surveyed methods.

  • Problem

    Subjective assessment is reliable but expensive, time consuming, and impractical for real-world or optimization use, while HDR assessment challenges conventional FR-IQA assumptions.

  • Method

    The paper surveys subjective and objective IQA classifications, datasets, performance measures, nine FR-IQA methods, HDR metrics, and 3-D IQA issues.

  • Results

    The survey evaluates the prediction performance and computation time of nine FR-IQA methods on four subjective datasets.

  • Takeaways & Limitations

    Objective IQA methods are presented as tools for automatically evaluating image quality, while 3-D methods must account for display, content, viewer, and depth perception.

  • Takeaways & Limitations

    Only a small number of the rapidly growing IQA methods are discussed in detail, and 3-D IQA remains dependent on display, content, viewer, and depth-perception factors.

Abstract

from arXiv · show

With the increasing demand for image-based applications, the efficient and reliable evaluation of image quality has increased in importance. Measuring the image quality is of fundamental importance for numerous image processing applications, where the goal of image quality assessment (IQA) methods is to automatically evaluate the quality of images in agreement with human quality judgments. Numerous IQA methods have been proposed over the past years to fulfill this goal. In this paper, a survey of the quality assessment methods for conventional image signals, as well as the newly emerged ones, which includes the high dynamic range (HDR) and 3-D images, is presented. A comprehensive explanation of the subjective and objective IQA and their classification is provided. Six widely used subjective quality datasets, and performance measures are reviewed. Emphasis is given to the full-reference image quality assessment (FR-IQA) methods, and 9 often-used quality measures (including mean squared error (MSE), structural similarity index (SSIM), multi-scale structural similarity index (MS-SSIM), visual information fidelity (VIF), most apparent distortion (MAD), feature similarity measure (FSIM), feature similarity measure for color images (FSIMC), dynamic range independent measure (DRIM), and tone-mapped images quality index (TMQI)) are carefully described, and their performance and computation time on four subjective quality datasets are evaluated. Furthermore, a brief introduction to 3-D IQA is provided and the issues related to this area of research are reviewed.

1 Introduction

Image distortions arise throughout acquisition, compression, transmission, and other processing stages, making reliable quality assessment essential. The survey introduces subjective and objective IQA, reviews datasets and measures, and emphasizes FR-IQA alongside color, HDR, and 3-D assessment.

  • Motivation: Processing stages such as acquisition, compression, and transmission introduce distortions that degrade image quality.Lossy compression can cause blurring and ringing, while limited transmission bandwidth can cause data loss.
  • Motivation: IQA supports image communication, management, acquisition, processing, segmentation, printing, display, fusion, and biomedical imaging.
  • Subjective and objective IQA: Subjective evaluation is accurate and reliable but expensive, time consuming, and sensitive to viewing conditions, displays, lighting, vision, and mood.
  • Subjective and objective IQA: Objective IQA seeks mathematical models that automatically predict image quality in agreement with average human observers.
  • Survey scope: The survey reviews subjective and objective IQA, six quality datasets, performance measures, and the three objective-IQA categories.
  • Survey scope: It emphasizes nine FR-IQA measures, evaluates their performance and computation time on four subjective datasets, and reviews color, HDR, and 3-D IQA.

2 Subjective image quality assessment

Subjective IQA asks human observers to rate or compare displayed images under standardized procedures, producing reliable measurements but requiring careful scoring and costly experiments. Standards, scoring transformations, and major practical drawbacks are reviewed.

  • Subjective testing: Subjective testing asks groups of observers to express opinions about image quality and is supported by international standards.
  • Standards: ITU-R BT.500-11 specifies viewing conditions, experiment instructions, test materials, and presentation of subjective results for television pictures.
  • Standards: ITU-T P.910 addresses digital video quality assessment below 1.5 Mbits/sec, while ITU-R BT.814-1 specifies display brightness and contrast.
  • Subjective methods: Single-stimulus testing displays each test image briefly before observers rate it on categories such as excellent, good, fair, poor, or bad.
  • Subjective methods: Double-stimulus testing displays test and reference images together before observers rate the test image using the same abstract scale.
  • Scoring: Raw ratings are unreliable because observers may use different quality scales across scenes and distortion types.
  • Scoring: DMOS uses the difference between reference and test-image raw quality scores, while Z-scores normalize each observer’s mean and variance.
  • Limitations: Subjective methods are accurate and reliable but time consuming, expensive, unsuitable for real-time applications, and dependent on observer and display conditions.

3 Objective image quality assessment

Objective IQA develops mathematical models to predict image quality automatically for monitoring, benchmarking, and optimization. Methods are classified by reference availability and by whether they are general-purpose or application-specific.

  • Purpose and applications: Objective IQA methods aim to predict image quality accurately and automatically, ideally matching average human quality judgments.
  • Purpose and applications: Objective metrics can monitor quality-control systems and automatically adjust image acquisition systems to obtain better image data.
  • Purpose and applications: Objective metrics benchmark image-processing algorithms and help select the algorithm producing higher-quality images.
  • Purpose and applications: In visual communication networks, metrics can optimize encoder pre-filtering and bit assignment alongside decoder post-filtering and reconstruction.
  • Classification: Reference availability defines FR-IQA, RR-IQA, and NR-IQA categories, ranging from a fully available reference to partial or absent reference information.
  • Classification: Methods are also general-purpose, covering varied distortions, or application-specific, targeting particular distortion types such as compression.
  • Classification: NR-IQA is convenient when references are unavailable but is more difficult than RR-IQA and FR-IQA.

3.2.1. Methods based on the models of the image source

Image-quality methods include statistical models of natural-image sources and distortion-oriented metrics, with emphasis on FR-IQA measures such as MSE and SSIM. MSE is computationally simple but poorly aligned with human perception, while SSIM compares luminance, contrast, and structure.

  • Methods based on the models of the image source: Source-model methods capture low-level natural-image statistics and can summarize image information at a low data rate.
  • Methods based on capturing image distortions: Distortion-specific methods are useful when the distortion is known but cannot capture distortions for which they were not designed.
  • Mean squared error (MSE): MSE measures the power of distortion between reference and test images and is simple, inexpensive, physically interpretable, and useful in optimization.
  • Mean squared error (MSE): MSE can poorly predict perceived quality because it ignores important physiological and psychophysical characteristics of the human visual system.
  • Structural similarity index (SSIM): SSIM models perceived degradation as structural-information change through luminance, contrast, and structure comparisons, combined into an overall similarity measure.
  • Structural similarity index (SSIM): SSIM is applied locally because image statistics and distortions vary spatially, producing a quality map with more localized information.

3.3.2.1. Parameter specification in the SSIM algorithm

MS-SSIM extends SSIM across scales to address scale selection under differing viewing conditions. VIF models natural images, distortions, and the human visual system in the wavelet domain to quantify conveyed information.

  • Multi-scale structural similarity index (MS-SSIM): MS-SSIM was motivated by SSIM’s dependence on an appropriate scale, which varies with viewing distance and display resolution.
  • Multi-scale structural similarity index (MS-SSIM): MS-SSIM combines luminance comparison at the coarsest scale with contrast and structure comparisons across multiple scales.
  • Parameter specification in the MS-SSIM algorithm: The MS-SSIM exponents were estimated using synthesized noisy images spanning 5 scales and 12 distortion levels, with judgments from 8 subjects.
  • Parameter specification in the MS-SSIM algorithm: The reported MS-SSIM parameters are α_5 = β_5 = γ_5 = 0.1333.
  • Visual information fidelity (VIF): VIF models natural images in the wavelet domain using Gaussian scale mixtures and comprises source, distortion, and HVS models.
  • Visual information fidelity (VIF): VIF represents distortion as signal attenuation plus additive white Gaussian noise and models the HVS as a noisy channel limiting transmitted information.

3.3.4.5. Calculating test image’s information

VIF estimates how much information from the reference image can be recovered in the test image, using mutual information across independent subbands. Its score summarizes overall or spatially localized quality, with larger values indicating better perceptual quality.

  • VIF information calculation: VIF estimates test-image information relative to the reference by assuming independent subbands and using a ratio of mutual-information terms.The measure can be computed over entire subbands or localized coefficient regions.
  • VIF interpretation: VIF ranges from 0 to 1 for practical distortions, where 0 indicates complete loss of reference information and values near 1 indicate higher perceptual quality.A linear contrast enhancement without added distortion can produce VIF greater than 1.
  • VIF interpretation: VIF greater than 1 indicates that the test image has superior visual quality to the reference after linear contrast enhancement.
  • Parameter estimation: The algorithm estimates model parameters from reference and test coefficients, while σ_n is selected for best overall quality-prediction accuracy.Cu is estimated from reference wavelet coefficients; g_i and σ_v,i are obtained by regression.
  • MAD detection-based strategy: MAD’s detection-based strategy transforms images into perceived luminance, computes an error image, and measures distortion visibility using local contrast.The strategy determines visible-distortion locations before calculating perceived detection distortion.
  • MAD detection-based strategy: The detection distortion score is nonnegative: zero means no visible distortions, while increasing values correspond to increasing perceived distortion and decreasing visual quality.

3.3.5.2. Appearance-based strategy

MAD’s appearance-based strategy models quality judgments for clearly distorted images by comparing multiscale local statistics of reference and test images. It combines this distortion with the detection-based score into an overall measure.

  • Appearance-based strategy: The appearance-based strategy targets low-quality images by modeling how observers look past distortions toward image subject matter.
  • Appearance-based strategy: Reference and test images are decomposed with a 2-D log-Gabor filter bank before local statistical differences are computed.The implementation uses five scales and four orientations, producing 20 subbands per image.
  • Appearance-based strategy: For each 16×16 block, the method compares subband standard deviation, skewness, and kurtosis across scales and orientations.Scale weights account for the HVS preference for coarser over finer scales.
  • Appearance-based strategy: The appearance distortion score ranges from 0 to infinity; zero indicates no perceived distortion, while larger values indicate lower visual quality.
  • Overall MAD measure: MAD combines detection-based and appearance-based distortions using a weighted geometric mean, with the weighting constant reflecting their relative importance.The weighting constant is selected based on detection distortion for good overall performance.
  • Implementation: MAD parameters are tuned for specified display and pixel conditions, with β1 and β2 chosen to optimize performance on the A57 dataset.

   L PC G S S S         x x x (45)

FSIM combines phase-congruency and gradient-magnitude similarities, weighting local similarity by perceptual significance indicated by phase congruency.

  • FSIM similarity measure: FSIM computes local similarity from phase congruency and gradient magnitude, then combines the two components with relative-importance weights.The cited implementation sets α = β = 1.
  • FSIM similarity measure: The final FSIM index aggregates the locally weighted similarity over the whole image spatial domain.
  • FSIM implementation: FSIM parameters are fixed after maximizing Spearman’s rank order correlation coefficient on a subset of TID2008.The subset contains 8 reference images and 544 corresponding test images.
  • FSIM implementation: The implementation uses four scales, four orientations, specified filter constants, and the Scharr operator because it achieved the highest SRCC among the tested gradient operators.

4 Quality assessment of color images

Color IQA metrics are needed because color information affects human quality judgments. FSIMC extends FSIM by separating luminance from chrominance and comparing chromatic components alongside luminance features.

  • Motivation: Color information affects observers’ judgments, motivating objective metrics that assess test color images against reference images.
  • FSIMC construction: FSIMC transforms RGB images into YIQ space so luminance and chrominance components can be treated separately.Y denotes luminance, while I and Q denote chrominance.
  • FSIMC construction: Chromatic similarity is computed separately for the I and Q components and combined by multiplication.The stabilizing constants T6 and T7 are set equal in the cited implementation.
  • FSIMC construction: FSIMC computes phase congruency and gradient magnitude from luminance, using the same calculation process as FSIM.
  • Implementation: FSIMC uses the FSIM parameter settings, with T6 = T7 = 200 and λ = 0.03.

5 Quality assessment of high dynamic range (HDR) images

HDR and tone-mapped image assessment requires objective methods that handle differing dynamic ranges because conventional FR-IQA methods assume similar reference and test ranges. The survey describes DRIM and TMQI as specialized approaches for these settings.

  • Motivation: Tone mapping compresses HDR dynamic range, causing information loss and motivating quality assessment of resulting LDR images.The survey notes that this degradation may not be visible to human observers.
  • TMQI: TMQI combines multi-scale structural fidelity with statistical naturalness in two stages for evaluating tone-mapped images against reference HDR images.Its structural fidelity comparison omits direct luminance comparison because tone-mapping changes local luminance and contrast.
  • DRIM: DRIM evaluates images with arbitrary dynamic ranges by producing distortion maps for visible-feature loss, invisible-feature amplification, and contrast-polarity reversal.Its inputs are luminance maps of reference and test images.
  • DRIM: DRIM models three structural changes: visible contrast becoming invisible, invisible contrast becoming visible, and visible contrast reversing polarity.These changes are associated with tone-mapping detail compression, inverse-tone-mapping contouring, and strong distortions, respectively.
  • DRIM: DRIM computes distortion probabilities across multiple scales and orientations, filters probability maps with corresponding cortex filters, and combines subband detections under an independence assumption.The algorithm visualizes distortion types using an in-context distortion map and selects the highest-probability type at each pixel.

6 Subjective datasets and performance measures in image quality assessment

The survey reviews six widely used subjective IQA datasets and six performance measures used to compare objective metrics with human quality judgments. It also describes score mapping needed to account for nonlinear subjective ratings.

  • Datasets: Six reviewed datasets are Cornell-A57, IVC, TID2008, LIVE, Toyama-MICT, and CSIQ.They differ in reference-image counts, distortion types, image totals, and subjective scoring procedures.
  • Datasets: TID2008 contains 1700 test images from 25 references, 17 distortion types, four distortion levels, and ratings from 654 observers across three countries.Its experiments vary lighting, screen size, monitor type, and color gamma.
  • Performance measures: A nonlinear mapping transforms each objective score before correlation with subjective scores to compensate for nonlinearity introduced during subjective experiments.The mapping parameters are estimated by minimizing squared differences between subjective and mapped scores.
  • Performance measures: The reviewed performance measures include PLCC, SRCC, KRCC, RMSE, MAE, and OR, covering prediction accuracy, monotonicity, and consistency.A good metric has higher PLCC, KRCC, and SRCC and lower RMSE, MAE, and OR.
  • Performance measures: SRCC and KRCC measure prediction monotonicity, whereas PLCC, RMSE, and MAE measure prediction accuracy and OR measures prediction consistency.SRCC is independent of monotonic nonlinear mappings between objective and subjective scores.

7 Evaluation results

The survey evaluates selected FR-IQA algorithms on four subjective quality datasets and separately measures their computation times because the methods target different image categories. DRIM is excluded because it does not produce a single image-wide score and lacks publicly available source code.

  • Evaluation setup: Eight FR-IQA algorithms are evaluated using original MATLAB implementations supplied by their authors.The evaluated methods include PSNR, SSIM, MS-SSIM, VIF, MAD, FSIM, FSIMC, and TMQI.
  • Evaluation setup: DRIM is excluded from performance evaluation because it produces a distortion map rather than a single whole-image quality score.It is also excluded from computation-time evaluation because publicly accessible source code was unavailable to the survey authors.
  • Results: Table 1 reports results on four subjective quality datasets, while Table 2 reports average SRCC, KRCC, PLCC, RMSE, and MAE over three datasets.Averages are computed both directly and with dataset-dependent weights.
  • Results: TMQI is evaluated only with SRCC and KRCC because its subjective experiment ranks images from best to worst rather than assigning scores within a fixed quality range.PLCC, RMSE, and MAE are therefore not calculated for that evaluation.
  • Computation time: Computation time is measured separately for the different image categories using images of size 512 × 512.The survey reports computation-time evaluation in Table 3.

8 Quality assessment of 3-D images

The survey reviews 3-D IQA descriptors, methods, standards, and datasets while emphasizing that 2-D IQA classifications do not transfer directly to perceived 3-D images. Existing 2-D methods perform well only for symmetric stereo images, and standardized descriptor quantification remains unresolved.

  • 3-D IQA methods: 2-D objective IQA methods evaluate 3-D images well only when the stereo images are symmetric, with approximately equal PSNR values for both eyes.This finding directly addresses the applicability of 2-D methods to 3-D images.
  • Open issues: There are no commonly accepted methods for quantifying the reviewed 3-D quality descriptors, although standards have recently been introduced.ITU-R, VQEG, and IEEE efforts address subjective assessment, ground truth, objective evaluation, viewing environments, and human factors.
  • 3-D IQA methods: 2-D FR-IQA, RR-IQA, and NR-IQA classifications do not apply straightforwardly because the perceived Cyclopean image cannot be directly accessed from left and right views.Only the two eye views are available, not the single mental image formed by binocular perception.
  • 3-D IQA methods: 3-D IQA methods are classified by their inputs as color-only methods or methods using both color and disparity information.The survey reviews examples of full-, reduced-, and no-reference methods in these categories.
  • Datasets: Reviewed datasets include LIVE 3-D IQA, IVC 3-D images, and a 3-D HDR tone-mapped dataset.The 3-D HDR dataset contains 9 reference images tone-mapped using 8 operators, for 81 total images.

9 Conclusion

The paper surveys IQA methods for conventional, HDR, color, and 3-D images, with particular emphasis on FR-IQA. It reviews nine FR-IQA measures, evaluates their prediction performance and computation time, and summarizes open issues in 3-D IQA.

  • IQA supports image-based applications by quantifying image quality across processes including compression, transmission, display, and acquisition.
  • The survey covers subjective and objective IQA, three objective-assessment categories, and quality assessment for 3-D, color, and HDR images.
  • The paper thoroughly describes nine FR-IQA methods and evaluates their prediction performance and computation time.
  • Only a small number of IQA methods are discussed in detail, selected mainly for citation, reported performance, and source-code availability.
  • 3-D IQA remains shaped by dependencies among display, content, and viewer factors, as well as individual user constraints and preferences.
Loading 1406.7799v1…