Source-linked AI summary

KonIQ-10k: An ecologically valid database for deep learning of blind image quality assessment

Vlad Hosu, Hanhe Lin, Tamas Sziranyi, Dietmar Saupe

arXiv:1910.06180v2cs.CVcs.MM

TL;DR

Deep-learning IQA is constrained by small existing datasets and the resources required for authentic image collection and reliable annotation. The paper introduces a scalable, ecologically valid dataset and an end-to-end BIQA model, with KonCept512 generalizing strongly to LIVE-itW. The authors identify training-database size as the main challenge for further improvement.

  • Problem

    Deep-learning IQA methods are limited by small existing datasets, while extensive authentic datasets require substantial resources for content generation and annotation.

  • Method

    The paper creates KonIQ-10k through a systematic, scalable sampling and crowdsourcing approach, then proposes the end-to-end KonCept512 BIQA model.

  • Results

    KonIQ-10k contains 10,073 diverse images with reliable subjective ratings, and KonCept512 generalizes well from it to LIVE-itW.

  • Takeaways & Limitations

    Ecologically valid, reliably annotated, and larger IQA databases support more general BIQA models and remain central to further performance improvement.

  • Takeaways & Limitations

    Further improvement remains constrained by training-database size despite careful image selection and reliable annotation.

Abstract

from arXiv · show

Deep learning methods for image quality assessment (IQA) are limited due to the small size of existing datasets. Extensive datasets require substantial resources both for generating publishable content and annotating it accurately. We present a systematic and scalable approach to creating KonIQ-10k, the largest IQA dataset to date, consisting of 10,073 quality scored images. It is the first in-the-wild database aiming for ecological validity, concerning the authenticity of distortions, the diversity of content, and quality-related indicators. Through the use of crowdsourcing, we obtained 1.2 million reliable quality ratings from 1,459 crowd workers, paving the way for more general IQA models. We propose a novel, deep learning model (KonCept512), to show an excellent generalization beyond the test set (0.921 SROCC), to the current state-of-the-art database LIVE-in-the-Wild (0.825 SROCC). The model derives its core performance from the InceptionResNet architecture, being trained at a higher resolution than previous models (512x384). Correlation analysis shows that KonCept512 performs similar to having 9 subjective scores for each test image.

1 INTRODUCTION

Existing IQA databases are too small, artificial, and costly to annotate for broadly applicable deep-learning BIQA. The paper addresses these limits with a scalable, ecologically valid dataset and an end-to-end model evaluated across databases.

  • IQA supports applications from image compression to machine vision, but subjective assessment is expensive and time-consuming.
  • Existing databases often limit content diversity by degrading a small set of pristine images with restricted distortion combinations.
  • Deep-learning BIQA methods were trained and evaluated on small, artificially distorted databases, motivating larger authentic datasets.
  • KonIQ-10k contains 10,073 images sampled from 10 million YFCC100M entries to scale database creation while supporting content and distortion diversity.
  • 1,459 crowd workers provided 120 reliable quality ratings per image through crowdsourcing.
  • KonCept512 is an end-to-end deep-learning BIQA model evaluated through cross-database testing on another IQA database.

2 RELATED WORK

Prior IQA databases and BIQA methods trade off scale, authenticity, diversity, subjective ratings, and end-to-end learnability. The reviewed work motivates KonIQ-10k as a response to these limitations while situating it among conventional and deep-learning approaches.

  • Earlier benchmark databases such as LIVE, TID2008, and CSIQ were widely used but were too small to substantially support deep-learning IQA.
  • Authentic-distortion databases improved realism, but LIVE in the Wild remained relatively limited in database size and content diversity.
  • The Waterloo Exploration database was large but used artificial distortions and lacked subjective ratings needed for training new subjective-quality models.
  • Crowdsourcing enabled larger image and video databases, and prior studies found reliable results under specific experimental setups.
  • Conventional BIQA methods engineer feature extractors and use regression models such as SVR to predict MOS.
  • Deep-learning BIQA methods instead learn representations from raw images, using feature extraction, sampled patches, transfer learning, or end-to-end architectures.

3 DATABASE CREATION

KonIQ-10k was created by filtering a massive public multimedia collection through scalable tag-based and indicator-based sampling to preserve authentic distortions and diverse content.

  • Motivation: The database targets better deep-learning BIQA by addressing existing datasets’ small scale, limited pristine content, and artificial degradations.The authors emphasize that IQA is content-dependent and requires authentic image degradations for accurate quality prediction.
  • Source selection: Approximately 9,974,030 YFCC100m image records were randomly selected before two-stage filtering produced the final database.The first stage retained images with suitable licenses, machine tags, and resolutions between 960 × 540 and 6000 × 6000.
  • Initial tag-based content sampling: A tag-based heuristic selected one million images from 4,807,816 candidates while targeting quota-based coverage of less-popular tags.Exact quota fulfillment was difficult because images carried 9.2 tags on average.
  • Filtering and diversity sampling: The second stage standardized images to 1024 × 768, used face- and saliency-aware cropping, and sampled 13,000 images with uniform distributions across eight indicators.Duplicates were subsequently removed using category and indicator information.
  • Availability: The database and code were made available at database.mmsp-kn.de.

3.3 Selective image cropping

The creation pipeline combines content-preserving cropping with multi-indicator sampling to standardize images while broadening representation of quality-related attributes.

  • Standardization: Images were cropped to 1024 × 768 because inconsistent resizing changes aspect ratio and can affect perceived quality and collection authenticity.The resolution was chosen because 95% of crowd workers’ devices had at least this resolution.
  • Selective image cropping: The selective crop keeps faces and salient areas inside the frame, avoiding unintended degradations caused by naive central cropping.Its importance map combines multi-pose Viola-Jones face detection with saliency information and sharpness-related relevance.
  • Selective image cropping: The crop location maximizes mean importance using a 1024 × 768 kernel, with frequency-domain convolution reducing average processing time below one second per image.The kernel is 1 except for a 10-pixel border on all sides, where its value is −1.
  • Diversity sampling: Sampling uses brightness, colorfulness, RMS contrast, sharpness, noise, and NR IQA-related indicators to broaden the range of represented distortions.Four retained measures were selected for correlation with human perception: brightness, colorfulness, RMS contrast, and sharpness.
  • Diversity sampling: Extreme indicator values were trimmed using z-scores before sampling, while VGG-16 FC7 features were quantized into 200 content clusters.The sampling procedure jointly optimizes quantized indicator histograms with Mixed Integer Linear Programming.
  • Diversity sampling: The procedure used N = 200 bins for seven scalar indicators and generated 13,000 images to allow duplicate removal and other post-filtering.

3.5 Removal of duplicates and inappropriate content

After diversity sampling, KonIQ-10k removes near-duplicates and unsuitable images, leaving 10,073 images for the database.

  • Duplicate removal: Uniform indicator binning can select identical or near-duplicate images, including slightly different views of the same scene.
  • Duplicate removal: Near-duplicates were identified through pairwise Euclidean distances in an eight-dimensional indicator-plus-content space after scaling indicator values to [0, 1].Content distance was set to 0 within a cluster and 1 across clusters.
  • Inappropriate content removal: Manual filtering removed text screenshots, text scans, heavily under-exposed images, and inappropriate mature-content images.After filtering, 10,073 images remained in KonIQ-10k.

3.6 Diversity analysis

KonIQ-10k exhibits broader quality-indicator and content distributions than the comparison databases LIVE-itW and TID2013.

  • Quality indicators: KonIQ-10k is more diverse than LIVE-itW and TID2013 in brightness, colorfulness, contrast, and sharpness distributions.
  • Content diversity: t-SNE embeddings of VGG-16 features show that LIVE-itW covers only a small region of KonIQ-10k, while TID2013 derives from just 25 reference images.
  • MOS distribution: The comparison also presents MOS distributions across TID2013, LIVE-itW, and KonIQ-10k after aligning their scales.

4 SUBJECTIVE IMAGE QUALITY ASSESSMENT

The study uses filtered crowdsourcing and expert-based checks to obtain quality scores for 10,073 images. Reliability analyses show substantial agreement with experts and strong consistency across crowd groups, while disagreements partly reflect differing interpretations of blur.

  • Crowdsourcing experiment: Workers were screened with quizzes, hidden test questions, outlier removal, and line-clicker detection before their ratings contributed to the final MOS.The quiz and hidden-question stages required accuracy above 70%.
  • Crowdsourcing experiment: 1,459 of 2,302 crowd workers passed the filtering steps, producing more than 1.2 million trusted judgments for 10,073 images.Each image received at least 120 scores.
  • Reliability of the crowd: The crowd MOS error relative to 11 experts converged to an RMSE lower bound of 11.35, while bootstrapped expert MOS had a standard deviation of 6.63.The comparison used bootstrapped crowd subsamples of up to 120 ratings and expert groups of size 11.
  • Reliability of the crowd: Crowd and expert ratings sometimes diverged because workers treated strong blur as degradation, whereas professional photographers interpreted shallow depth of field as an artistic effect.The authors attribute this disagreement at least partly to differing domain knowledge.
  • Reliability of the crowd: Mean inter-group agreement increased with group size, reaching 0.973 SROCC when random halves of contributors were compared.For groups of 700 observers, the average number of ratings per image was 57.68 and agreement was 0.973±0.001 SROCC.

5 FINDING A BETTER END-TO-END DEEP BIQA

The section develops an end-to-end BIQA model by comparing input representations, CNN backbones, prediction targets, and loss functions. It supports both direct MOS prediction and rating-distribution prediction.

  • Design factors: Existing BIQA methods vary in input presentation, base architecture, loss function, and prediction aggregation.The section studies these design factors to improve deep BIQA performance.
  • Architecture: The proposed system uses a CNN body, global average pooling, and fully connected layers to predict either MOS or five-bin rating distributions.The output layer has one unit for MOS prediction or five units for rating distributions.
  • Training strategy: The model uses a pretrained CNN body with a fine-tuned end-to-end prediction head.The fully connected layers use ReLU activations and dropout before the final MOS or distribution output.
  • Loss functions: MAE and MSE are used for MOS prediction, while cross-entropy, Huber, and Earth Mover’s Distance losses train rating-distribution predictions.The loss-function choices correspond to the two prediction targets.

6 EXPERIMENTAL RESULTS

The experiments evaluate the proposed deep BIQA model on KonIQ-10k and LIVE-itW, covering both in-dataset testing and cross-database generalization.

  • Evaluation datasets: The model is evaluated on KonIQ-10k and the LIVE in the Wild database.These databases support evaluation on the proposed dataset and cross-database testing.

6.1 Setup

The experimental setup partitions KonIQ-10k into training, validation, and test sets, then evaluates BIQA models using rank and linear correlation metrics. Training compares MOS and rating-distribution prediction.

  • Data splits: KonIQ-10k is divided into 7,058 training, 1,000 validation, and 2,015 test images.The validation set selects the best generalizing model, while the test set reports final performance.
  • Metrics: BIQA performance is evaluated with SROCC and PLCC.The metrics compare predictions with subjective quality scores.
  • Optimization: Training uses Adam with staged learning rates and selects models using validation-set PLCC.The schedule begins at α = 10^-4, then reduces the rate to 5 × 10^-5 and 10^-5.
  • Prediction targets: The system tests two training targets: direct MOS values and distributions of ratings.This comparison is represented in the end-to-end architecture experiments.
  • Implementation: Batch sizes are chosen separately for each CNN according to available GPU memory.Different model depths and parameter counts prevent using one common batch size.

6.2 Performance evaluation and discussion

KonCept512 combines InceptionResNetV2, MSE loss, and 512 × 384 inputs, achieving strong KonIQ-10k performance and broader generalization than conventional BIQA methods. Larger training sets and human-rating comparisons further contextualize its performance.

  • 6.2.1 Best model selection: KonCept512: 512 × 384 inputs outperform 224 × 224 inputs and generally outperform 1024 × 768 inputs in the resolution study.The authors suggest that down-sampling to 224 × 224 loses quality-related information, while larger inputs may interact poorly with architectures and batch sizes.
  • 6.2.1 Best model selection: KonCept512: InceptionResNetV2 achieves the best performance among the tested CNN backbones.Deeper architectures perform better, while MSE is best on the KonIQ-10k test set with only marginal improvement over MAE and Huber.
  • 6.2.1 Best model selection: KonCept512: 0.836 SROCC is achieved by distributional Huber loss when cross-tested on LIVE-itW.Models trained on KonIQ-10k show about a 0.1 SROCC reduction on LIVE-itW despite larger differences on KonIQ-10k.
  • 6.2.1 Best model selection: KonCept512: KonCept512 uses InceptionResNetV2, MSE loss, and 512 × 384 images, with LIVE-itW images resized to that resolution for cross-testing.This configuration is named the best-performing model.
  • 6.2.2 Comparison with state-of-the-art BIQA methods: Deep-learning BIQA methods lose only around 0.1 SROCC when cross-tested, whereas conventional methods show substantial performance drops.KonCept512 also improves SROCC by around 0.02 over local patch-based DeepBIQ on both datasets.
  • 6.2.3 Effects of the training set size: 0.965 ± 0.025 SROCC and 0.895 ± 0.021 SROCC are extrapolated for 100,000 virtual training images on KonIQ-10k and LIVE-itW, respectively.The estimates use 95% observational confidence bounds from bootstrapping.
  • 6.2.4 On the prediction power of IQA methods: KonCept512 matches the average SROCC of approximately nine subjective ratings per test image.The comparison uses an SROCC of 0.918 between predictions and MOS from a random half of participants.

7 CONCLUSION

KonIQ-10k combines ecological validity, diversity, and reliable subjective scoring with a deep model designed for strong IQA generalization. The authors identify training-database size as the main remaining challenge for improving deep BIQA.

  • Generalization: KonIQ-10k generalizes well to LIVE-itW, which the authors attribute to representative training data, reliable crowdsourcing, and architecture improvements.The stated explanation combines dataset diversity, reliable experiments for both databases, and the technical design of KonCept512.
  • Dataset: KonIQ-10k contains diverse, representative images spanning categories, quality indicators, and technical parameters.The 10,073 images come from 1,265 camera models across about 100 manufacturers.
  • Model: KonCept512 uses InceptionResNetV2 and improves IQA through higher-resolution training, model selection by correlation, and a multi-resolution fully connected head.It trains at 512 × 384 rather than the typically used 224 × 224 resolution and uses a GAP layer.
  • Outlook: The authors identify training-database size, alongside careful image selection and reliable annotation, as the main challenge for further deep BIQA improvement.They predict that similarly built datasets of about 100,000 images will close the gap with aggregated opinions from large observer groups in the wild.

8 APPENDIX

The appendix describes an efficient heuristic for tag-based image sampling and provides biographical information about the paper’s authors.

  • Sampling: The sampling heuristic uses tag frequencies and a quota to select images efficiently from a large source set.Tags occurring fewer than the quota are fully included before remaining tags are handled approximately.
  • Author biographies: Vlad Hosu’s research includes visual quality assessment, image enhancement, crowd-sourcing strategies, and human visual perception.
  • Author biographies: Hanhe Lin’s research focuses on machine learning, deep learning, visual quality assessment, and crowd-sourcing.
  • Author biographies: Tamas Sziranyi is a professor and researcher whose work includes image processing and pattern recognition.The appendix also notes his editorial, professional, and award histories.
  • Author biographies: Dietmar Saupe is a computer science professor leading research in multimedia signal processing and related areas.His listed interests span compression, computer vision, medical image processing, and sports informatics.
Loading 1910.06180v2…