Source-linked AI summary
Massive Online Crowdsourced Study of Subjective and Objective Picture Quality
Deepti Ghadiyaram, Alan C. Bovik
TL;DR
Existing IQA databases emphasize controlled synthetic distortions, while mobile photographs contain diverse authentic mixtures that are less well represented. The paper creates a representative database and crowdsourcing study, collecting hundreds of thousands of ratings and testing blind IQA models in this setting.
Problem
Existing image-quality databases inadequately represent the diverse authentic distortion mixtures found in mobile-device photographs.
Method
The authors build the LIVE In the Wild database and conduct a large online subjective study using a crowdsourcing framework.
Results
350,000 opinion scores on 1,162 naturally distorted images were collected from over 8,100 subjects, while state-of-the-art NR IQA algorithms performed poorly on the database.
Takeaways & Limitations
The database and ratings provide a challenging real-world resource for developing and evaluating blind image-quality prediction models.
Takeaways & Limitations
The crowdsourced study lacks detailed control and data about workers’ display devices and viewing conditions, so additional reliability checks could benefit future studies.
Abstract
from arXiv · showhide
Most publicly available image quality databases have been created under highly controlled conditions by introducing graded simulated distortions onto high-quality photographs. However, images captured using typical real-world mobile camera devices are usually afflicted by complex mixtures of multiple distortions, which are not necessarily well-modeled by the synthetic distortions found in existing databases. The originators of existing legacy databases usually conducted human psychometric studies to obtain statistically meaningful sets of human opinion scores on images in a stringently controlled visual environment, resulting in small data collections relative to other kinds of image analysis databases. Towards overcoming these limitations, we designed and created a new database that we call the LIVE In the Wild Image Quality Challenge Database, which contains widely diverse authentic image distortions on a large number of images captured using a representative variety of modern mobile devices. We also designed and implemented a new online crowdsourcing system, which we have used to conduct a very large-scale, multi-month image quality assessment subjective study. Our database consists of over 350000 opinion scores on 1162 images evaluated by over 7000 unique human observers. Despite the lack of control over the experimental environments of the numerous study participants, we demonstrate excellent internal consistency of the subjective dataset. We also evaluate several top-performing blind Image Quality Assessment algorithms on it and present insights on how mixtures of distortions challenge both end users as well as automatic perceptual quality prediction models.
I. INTRODUCTION
The paper introduces a real-world image-quality resource and crowdsourced study to address the mismatch between synthetic legacy databases and authentic mobile-camera distortions. It evaluates human ratings and blind quality models on this challenging setting.
- Motivation: Automatic quality prediction can help identify low-quality images and support quality-aware capture, transmission, and viewing processes.The paper links objective quality prediction to culling or correcting poor images and improving quality of experience.
- Motivation: Mobile-camera photographs often contain complex mixtures of authentic distortions that synthetic, singly distorted databases omit.This mismatch limits how representative legacy training and evaluation data are for real-world blind IQA.
- Contributions: The authors conducted a very large, multi-month online subjective study using a purpose-built crowdsourcing system.The study was designed to collect rich human opinion data and validate the resulting ratings.
- Contributions: The LIVE In the Wild database contains 1162 authentically distorted images captured by diverse mobile devices without artificially adding distortions.The database was designed to include broad image content, capture conditions, quality types, mixtures, and severities.
- Results: 350,000 opinion scores from over 8,100 subjects were collected, and the crowdsourcing ratings correlated 0.9851 with an independent study’s MOS values.The reported correlation supports the reliability of the online system’s human judgments.
- Results: State-of-the-art no-reference IQA algorithms perform poorly on the database, which contains many authentically multiply distorted images.The result indicates that models successful on legacy databases face difficulty in this more realistic setting.
II. RELATED WORK
The related-work review contrasts controlled, synthetic IQA benchmarks and laboratory studies with the paper’s more diverse authentic-image and Internet-based approach. It also identifies methodological and data-access limitations in prior online studies.
- Benchmark IQA Databases: LIVE, TID2008, and TID2013 are widely used benchmarks built primarily from synthetically distorted images and controlled subjective studies.LIVE models five distortion types, TID2008 models 17 categories, and TID2013 expands this to 24 simulated distortions.
- Test Methodologies: TID2008’s opinion scores were derived from accumulated pairwise winning points, a procedure the authors consider questionable because ranking methods can affect reliability.The review also notes that its presentation format may not reflect common mobile viewing scenarios and may introduce hysteresis effects.
- Benchmark IQA Databases: Legacy databases do not adequately represent complex authentic distortion mixtures produced by mobile-device capture.This gap motivates collecting real images with natural distortion mixtures for robust IQA development.
- Online Subjective Studies: Internet-based studies offer larger and more diverse subject pools, but they sacrifice control over ambient and display conditions.The paper presents this trade-off as a central motivation and challenge for online subjective assessment.
- Online Subjective Studies: Prior online rating studies generally used small sets of synthetically distorted images, and their subjective data were often unavailable publicly.These constraints limited their relevance to authentic mobile-camera imagery and reproducibility for later research.
- Online Subjective Studies: A prior Mechanical Turk study tested only 116 JPEG-compressed images with 40 subjects, whereas this work used 1162 images and more than 8100 subjects.The authors also sought laboratory-like instructions, training, test phases, and continuous ratings through a custom framework.
III. LIVE IN THE WILD IMAGE QUALITY CHALLENGE DATABASE
The LIVE In the Wild database was created to address the limited realism of legacy image-quality databases by collecting images with diverse authentic distortion mixtures from varied mobile devices.
- Authentically distorted mobile-camera images can combine low-light noise, blur, exposure errors, and compression artifacts.
- Existing image-quality databases lack content diversity and mixtures of bona fide distortions, limiting development of models for real-world image distortions.
- The database contains images afflicted by diverse authentic distortion mixtures on a variety of commercial devices.
- Images were captured using a wide variety of mobile device cameras and cover diverse subjects, scenes, compositions, and visual activities.
A. Distortion Categories
The distortion-category study shows that authentic image distortions often resist a single agreed label. Disagreement increases when images contain multiple interacting distortions and viewer sensitivities differ.
- The study asked subjects to select the single category representing the most dominant distortion in each image.
- 100% consensus images had a dominant distortion on which most subjects fully agreed.
- 50% consensus images received approximately equal votes for two different distortion classes.
- No-consensus images had nearly a third of subjects assigning a category different from the two other dominant labels.
- Viewer sensitivities, display device, viewing distance, and image content contribute to disagreement about underlying distortion.
- Real-world images form a multidimensional continuum of interacting distortion perturbations rather than discrete, well-defined categories.
IV. CROWDSOURCED FRAMEWORK FOR GATHERING
The study used crowdsourcing to gather subjective image-quality ratings at scale while acknowledging limited control over participants’ viewing environments and task participation.
- Crowdsourcing platforms enable large numbers of opinions from diverse, geographically distributed participants.
- Crowdsourcing limits control over room illumination and display devices, so the study collected information about these factors in a compulsory survey.
- Online tests must be partitioned into smaller tasks because workers may avoid large, time-consuming assignments.
- Researchers cannot ensure that every participating worker views and rates every image in the dataset.
- Despite crowdsourcing limitations, the study gathered a large number of highly reliable opinion scores across the dataset.
- The study focused on subjective quality scores rather than image aesthetics and identified future value in collecting content and aesthetics information.
B. Instructions, Training, and Testing
The crowdsourced test combined participant screening, standardized instructions, training, randomized rating tasks, and survey collection to obtain quality judgments while addressing worker reliability.
- B. Instructions, Training, and Testing: AMT workers first read task instructions, accepted the HIT, completed the task page, and submitted results.
- B. Instructions, Training, and Testing: Image-quality scoring is more subtle and subjective than object labeling, requiring attention to workers’ limited experience with image-quality concepts.
- B. Instructions, Training, and Testing: Workers were restricted to unique participation and could not proceed after accepting the task previously.
- B. Instructions, Training, and Testing: More than 350,000 ratings were collected after allowing only workers with confidence values greater than 0.75 to participate.
- B. Instructions, Training, and Testing: Each HIT used seven training images followed by 43 randomized test images, for 50 rated images total, with slider scores converted to integers from 1 to 100.
- B. Instructions, Training, and Testing: The study used a single-stimulus continuous procedure in which participants dragged a slider to report image-quality judgments.
- C. Subject Reliability and Rejection Strategies: Crowdsourcing required addressing noisy ratings and the reliability of Mechanical Turk workers.
- 1) Intrinsic metric:: Workers with AMT confidence values greater than 75% were allowed to select the task, and each worker could select it no more than once.
2) Repeated images:
The study used repeated and gold-standard images within a Mechanical Turk HIT to screen unreliable workers and validate collected ratings against established scores.
- Worker screening: Workers whose repeated-image ratings differed beyond a threshold on at least 3 of 5 images were rejected.The rejection threshold was rounded to 20.
- Repeated-image procedure: Workers entered a training phase followed by a randomized test phase containing repeated and gold-standard images.The HIT used 43 test images, including five repeated-image checks and five gold-standard images.
- Gold-standard validation: Gold-standard ratings agreed closely with laboratory MOS values, with a mean absolute difference of 4.65 that was not statistically significant.The gold-standard images came from the LIVE Multiply Distorted Image Quality Database.
- Validation outcome: The crowdsourcing framework was judged effective for gathering high-quality opinion data despite uncontrolled online conditions.The authors describe the obtained subject data as reliable and high quality.
D. Subject-Consistency Analysis
Subjective ratings showed strong consistency across workers and images, while uncontrolled viewing conditions remained a source of possible variability.
- Subject consistency: 0.9896 average Spearman correlation across random rating splits indicates high consistency of image MOS estimates.The correlation was averaged over 25 random splits of each image’s ratings.
- Intra-subject consistency: 0.8721 median SROCC between individual scores and gold-standard MOS values indicates substantial intra-subject reliability.The median was computed over all participating subjects.
- Dataset scale: More than 350,000 ratings from over 8,100 unique subjects formed the database after unreliable participants were rejected.Each image was rated by an average of 175 unique subjects.
- Score distribution: MOS values ranged from 3.42 to 92.43, with an average per-image subjective-score standard deviation of 19.2721.MOS was computed by averaging individual scores from multiple workers.
- Study conditions: Uncontrolled conditions varied in demographics, displays, lighting, viewing distance, and concentration, potentially affecting quality scores.The study did not test participants for vision problems, although corrective-lens compliance was checked.
- Analysis overview: The figures summarize participant demographics, preferred capture devices, MOS distributions, and factors analyzed for their influence on perceived quality.The factor analysis controlled other variables while examining each factor independently.
A. Gender
Under controlled demographic and device conditions, gender and age produced similar ratings on the selected test images, but image content may interact with these factors.
- Gender: Male and female workers rated five randomly chosen images similarly when age, display type, and viewing distance were held constant.The comparison used subjects aged 20–30 using desktops from 15–30 inches away.
- Interpretation: Gender and age did not significantly affect ratings on the randomly chosen images examined.The authors caution that the limited image sample does not exclude content-specific effects.
- Age: Subjects in different age categories also rated the selected images similarly under fixed gender, display, and viewing-distance conditions.This analysis used laptop users aged across the reported categories and the same five images.
- Scope: A systematic study of image content, gender, and age could clarify their interplay in perceptual quality.The paper identifies this interaction as a topic for future study.
D. Display Device
The analysis examined display devices and other viewing factors independently; device type showed little effect in the selected comparisons, while ratings still varied with uncontrolled conditions and image quality.
- Display device: Display device type appeared to have little effect on subjective ratings for the five analyzed images.The comparison focused on 20–30-year-old subjects viewing from 15–30 inches away.
- Display-device caveat: The authors do not claim that display devices generally leave perceived image quality unaffected.They suggest that finer details such as resolution and display technology may matter.
- Distortion sensitivity: Subjects with different annoyance levels toward poor Internet picture quality were almost equally sensitive to distortions.Ratings were grouped by three survey responses about whether poor image quality bothers them.
- Rating convergence: MOS values flattened as more subjects rated an image, with greater consistency for very high and very low MOS images than intermediate-quality images.Viewing conditions and image presentation order contributed to score variability.
- Overall assessment: The study’s diverse, uncontrolled participation supported analysis of viewing factors and produced reliable large-scale data when compared with laboratory gold standards.The authors base this assessment on factor analyses and correlations against controlled-condition MOS values.
F. Limitations of the current study
The study recognizes that crowdsourcing can produce valuable, generalized subjective databases but introduces complexities that may affect result validity. Reliability assessment therefore depends on deeper analysis and could be strengthened through richer participant and viewing-condition data.
- F. Limitations of the current study: Crowdsourcing offers valuable, generalized subjective databases but involves complexities and pitfalls that may affect the veracity of results.The paper flags crowdsourcing as promising while directing readers to analyses of its potential concerns.
- F. Limitations of the current study: AMT confidence values alone are not necessarily reliable indicators for any specific task.The authors base confidence in their results on deep analysis rather than simply screening participants by aggregate AMT confidence.
- F. Limitations of the current study: Future studies could improve reliability assessment by collecting more detailed information about workers’ display devices and viewing conditions.The paper also mentions visual tests and reports of time spent viewing and rating images as potentially useful measures.
- F. Limitations of the current study: The evaluation uses randomly divided, content-separated 80% training and 20% testing sets for learning and validation.This split defines the reported algorithm-comparison protocol rather than a limitation of crowdsourced subjective judgments.
A. New Blind Image Quality Assessment Model
The paper introduces FRIQUEE, a blind IQA model designed for authentic mixtures of picture distortions. It combines 564 statistical features with deep belief network and SVM learning, and performs significantly better than current leading methods on the LIVE In the Wild database.
- A. New Blind Image Quality Assessment Model: FRIQUEE is a blind IQA model designed to address mixtures of authentic picture distortions in the LIVE In the Wild database.Its name expands to Feature maps based Referenceless Image QUality Evaluation Engine.
- A. New Blind Image Quality Assessment Model: 564 statistical features feed a model combining a deep belief net and an SVM for perceptual quality prediction.The deep belief net uses four hidden layers formed by stacking restricted Boltzmann machines.
- A. New Blind Image Quality Assessment Model: FRIQUEE’s learned deep representations generalize across different distortion types, mixtures, and severities.The model builds more complex representations from the statistical FRIQUEE features.
- A. New Blind Image Quality Assessment Model: Significantly better performance on unseen test data distinguishes FRIQUEE from current top-performing state-of-the-art methods on the LIVE In the Wild database.The comparison uses content-separated training and testing data; the reported night-image exclusion leaves 1013 images in one experiment.
- A. New Blind Image Quality Assessment Model: Legacy LIVE contains singly distorted images, whereas the challenge database contains unknown mixtures that substantially challenge top-performing algorithms.The paper compares median correlations on both databases to emphasize this difference.
D. With and Without Night-time Images
Including night-time images tests blind IQA models under severe low-light distortions absent from legacy databases. FRIQUEE remains comparatively strong, supporting training across complex distortion mixtures and varied lighting conditions.
- D. With and Without Night-time Images: 149 of 1,162 images were captured at night and contain severe low-light distortions.The paper uses “low-light images” and “night images” interchangeably.
- D. With and Without Night-time Images: Legacy benchmark databases lack images captured under such low illumination conditions.The paper notes that other models’ NSS features were trained on natural images under normal lighting conditions.
- D. With and Without Night-time Images: FRIQUEE still performed well relative to other state-of-the-art models when night-time images were included in training and testing.The experiments included the night-time pictures in the data pool and trained FRIQUEE alongside other blind IQA models.
- D. With and Without Night-time Images: The results support training generalizable blind IQA models on mixtures of complex distortions under different lighting conditions.The paper frames this as an implication of the day-and-night evaluation.
- D. With and Without Night-time Images: The study reports more than 350,000 subjective judgments and identifies future large-scale crowdsourced video-quality studies as a continuation of this effort.The conclusion connects realistic distortion diversity in images with the need for representative modern video-quality databases.