Source-linked AI summary

Large-Scale Study of Perceptual Video Quality

Zeina Sinno, Alan C. Bovik

arXiv:1803.01761v2eess.IV

TL;DR

Existing VQA databases inadequately represent diverse real-world videos and authentic mixed distortions, limiting evidence for no-reference quality prediction. The paper constructs LIVE-VQC and collects large-scale crowdsourced quality judgments, then evaluates blind predictors on it. The database supports this evaluation, while results show demographic effects in opinion scores and substantial room for better prediction.

  • Problem

    Existing VQA databases provide limited content, capture diversity, resolutions, users, and distortion types compared with real-world videos containing authentic mixed distortions.

  • Method

    The paper combines a crowdsourced VQA framework with LIVE-VQC, a database of authentic videos collected from diverse contributors and devices.

  • Results

    V-BLIINDS achieved the best performance across the three reported performance metrics, while the results still indicate room for improved NR VQA models.

  • Takeaways & Limitations

    LIVE-VQC provides a large, diverse resource for evaluating NR VQA models on complex authentic video distortions.

Abstract

from arXiv · show

The great variations of videographic skills, camera designs, compression and processing protocols, and displays lead to an enormous variety of video impairments. Current no-reference (NR) video quality models are unable to handle this diversity of distortions. This is true in part because available video quality assessment databases contain very limited content, fixed resolutions, were captured using a small number of camera devices by a few videographers and have been subjected to a modest number of distortions. As such, these databases fail to adequately represent real world videos, which contain very different kinds of content obtained under highly diverse imaging conditions and are subject to authentic, often commingled distortions that are impossible to simulate. As a result, NR video quality predictors tested on real-world video data often perform poorly. Towards advancing NR video quality prediction, we constructed a large-scale video quality assessment database containing 585 videos of unique content, captured by a large number of users, with wide ranges of levels of complex, authentic distortions. We collected a large number of subjective video quality scores via crowdsourcing. A total of 4776 unique participants took part in the study, yielding more than 205000 opinion scores, resulting in an average of 240 recorded human opinions per video. We demonstrate the value of the new resource, which we call the LIVE Video Quality Challenge Database (LIVE-VQC), by conducting a comparison of leading NR video quality predictors on it. This study is the largest video quality assessment study ever conducted along several key dimensions: number of unique contents, capture devices, distortion types and combinations of distortions, study participants, and recorded subjective scores. The database is available for download on this link: http://live.ece.utexas.edu/research/LIVEVQC/index.html .

I. INTRODUCTION

Video quality assessment models must reflect diverse real-world content, capture conditions, and authentic mixed distortions, but existing databases and crowdsourced studies provide limited coverage or control. The paper addresses these gaps with a robust crowdsourcing framework and the LIVE-VQC database.

  • Motivation: Existing VQA databases contain limited content, few capture users and devices, fixed conditions, and mostly synthetic distortions, restricting generalization to diverse videos.Real-world videos often contain authentic, commingled distortions that are difficult to reproduce synthetically.
  • Motivation: Crowdsourced video studies must address participant reliability, display variation, playback performance, bandwidth, and fatigue.Video stalls can heavily affect experienced quality, while long sessions can reduce focus and performance.
  • Contributions: The paper introduces a robust framework for conducting crowdsourced video quality studies and collecting quality scores through Amazon Mechanical Turk.The framework was designed to account for technical and participant-related difficulties in uncontrolled viewing environments.
  • Contributions: LIVE-VQC contains 585 unique-content videos captured by 80 users and 101 devices, with authentic real-world distortions and associated mean opinion scores.The database is made available to the research community at no charge.
  • Study scope: The study collects more than 205000 opinion scores on 585 diverse videos and evaluates prominent blind VQA algorithms on the resulting database.The paper presents the database, crowdsourcing methods, post-processing, and predictor evaluation in subsequent sections.

II. LIVE VIDEO QUALITY CHALLENGE (LIVE-VQC) DATABASE

LIVE-VQC was built from videos contributed by diverse everyday users under largely unconstrained capture conditions. The resulting collection spans varied content, devices, environments, and authentic distortions.

  • Purpose: The database targets mobile- and digital-camera videos captured by casual users, providing authentically captured and distorted content with psychometric quality scores.The stated objective is to support research on this comparatively less-studied class of video.
  • Content Collection: 80 largely naïve contributors supplied videos captured with mobile cameras across diverse ages, genders, social backgrounds, cultures, and geographic locations.Contributors ranged from 11 to 65 years old and came from many countries and populated continents.
  • Content Collection: Contributors uploaded videos as captured, without processing-app alterations, and received no instructions about content or capture style beyond normal use.Only videos lasting at least 10 seconds were accepted.
  • Content Collection: The content includes sports, concerts, nature, human activities, indoor and outdoor scenes, varied lighting, and diverse camera and in-frame motion.These conditions contribute to complex, space-variant distortions that are difficult to categorize precisely.
  • Capture Devices: 101 devices, including 43 models released from 2009 to 2017, were used to capture the videos, with smartphones forming the majority.Most videos were captured using devices released after 2015.
  • Capture Devices: Approximately 74% of viewed videos were captured using Apple and Samsung devices.The distribution is grouped by device brand in Figure 2.

C. Video Orientations and Resolutions

The study accommodates heterogeneous display capabilities and video orientations by selectively downscaling high-resolution material and assigning videos according to detected display resolution. Most database videos use three predominant resolutions.

  • Video Orientations and Resolutions: 23.2% of videos were captured in portrait mode and 76.2% in landscape mode, with some high-resolution portrait videos downscaled for display compatibility.Downscaling used bicubic interpolation for specified portrait resolutions.
  • Video Orientations and Resolutions: Only 10–20% of global web viewers were estimated to have displays at least 1920×1080, motivating selective downsampling of high-resolution landscape videos to 1280×720.110 videos remained at 1920×1080 while other 1920×1080-and-higher videos were downsampled.
  • Video Orientations and Resolutions: The resolutions 1920×1080, 1280×720, and 404×720 together account for 93.2% of the database.All other resolutions combined account for 6.8%.
  • Viewing Assignment: Workers with displays at least 1920×1080 viewed half their videos at that resolution, while other workers viewed randomly selected videos below 1920×1080.The assignment strategy used detected display resolution to distribute viewing tasks.
  • Viewing Assignment: The study used a single-stimulus presentation because real users typically view one video at a time and each content had one authentic distorted version.Crowdsourcing conditions prevented applying all ITU standard recommendations.

A. Participation Requirements

Participation required reliable, unique workers using non-mobile displays with at least 1280×720 resolution and a supported browser. These constraints addressed rating independence, playback compatibility, and display consistency.

  • Eligible workers needed an AMT reliability score above 90%, prior nonparticipation, a non-mobile device, at least 1280×720 resolution, and a supported browser.
  • The unique-worker condition avoided judgment biases from rating videos more than once.
  • Mobile devices were excluded because browsers lacked video preloading and native players could introduce uncontrolled scaling artifacts.
  • Browser compatibility and preloading were verified to support smooth playback during sessions.
  • Sessions were limited to 30 minutes, with progress checkpoints to filter cases involving very slow loading.

7) Hardware constraint:

Each session combined training with a controlled set of test videos, including database samples, repeated controls, golden videos, and videos rated by every worker.

  • Viewed Content: Each subject viewed 50 videos: 7 training videos and 43 videos during the rating process.
  • Viewed Content: Training videos spanned the database’s quality range and mixed resolutions to prepare subjects for later viewing.
  • Viewed Content: The 43 test videos included 4 golden videos, 31 randomly selected database videos, 4 repeated videos, and 4 videos shown to all workers.
  • Viewed Content: Golden videos used prior human ratings from a tightly controlled study as a control for validating subjects’ ratings.
  • Viewed Content: Test videos were placed in re-randomized order for each subject.

Step 1: Overview:

The study introduced the task with instructions and examples, then required full video playback before quality ratings through a controlled interface.

  • Overview: Workers previewed task requirements and were instructed to compare each video with an ideal video of the same content.
  • Overview: Examples demonstrated underexposure, stalls, shakes, blur, and poor color representation while instructing workers to rate all distortions.
  • Training: The training phase consisted of 7 videos and used the same rating interface shown in Fig. 6.
  • Training: Video controls were hidden, videos played in their entirety, and loading progress was shown before playback.
  • Training: After playback, workers rated quality on a continuous bar guided by five labels from Bad to Excellent.
  • Training: Training playback also measured play duration to assess workers’ ability to play videos under possible stall conditions.

Step 5: Testing:

Testing followed the training procedure for 43 videos, included motivational progress feedback, and collected participant demographics and display information afterward.

  • Testing: The testing phase displayed, controlled, and rated videos like training, but required 43 videos instead of 7.
  • Testing: After 10 testing videos, workers received motivational feedback when progress was sufficient, while slower progress triggered a different message.
  • Exit Survey: The exit survey collected workers’ gender, age, country, corrective-lens use, comments, and automatically recorded display resolution.
  • Participants: The experiment included 4776 AMT subjects from 56 countries, with 91% located in the United States and India.
  • Participants: Participants were 46.4% male and 53.6% female.

2) Viewing Conditions:

The crowdsourced study accommodated diverse participant viewing conditions and used screening, consistency checks, and worker feedback to support reliable ratings.

  • Viewing conditions: Participants viewed videos under diverse locations, visual correction, distances, browsers, displays, resolutions, and lighting conditions.Most had normal or corrected-to-normal vision; 2.5% with abnormal, uncorrected vision were excluded.
  • Study participation: The study averaged 16.5 minutes per participant and paid one US dollar per completed HIT.The researchers sought high-quality workers partly by maintaining a good AMT reputation.
  • Quality control: The framework addressed participant reliability, hardware variation, and bandwidth-related stalls through training and subject rejection procedures.These issues were more challenging than in the preceding crowdsourced image-quality study.
  • Consistency control: Four repeated videos measured intra-subject consistency, while stalls complicated comparisons between original and time-displaced repeats.Ratings were generally similar when repeated videos played without stalls.
  • Participant feedback: Exit-survey feedback was generally positive, and 32% of the 4776 completing workers submitted additional comments.Among commenters, 31% described the test positively, while 55% reported no additional comments.

A. Video Stalls

The study reduced but did not eliminate playback stalls through training and exclusion procedures; after cleansing, the retained ratings covered the quality spectrum, and golden-video agreement validated the protocol.

  • Stall prevalence: 92% of videos had no stalls or stalls shorter than 1 sec, while 77% played without any stalls.Stalls were substantially mitigated but remained possible because participant-device computation was stochastic.
  • Subject rejection: 11.5% of subjects were excluded because at least 75% of their viewed videos stalled, while 2% circumvented the experiment and 2.5% omitted corrective lenses.Only 23 subjects, or 0.5%, were later identified as outliers under the standard rejection procedure.
  • Cleansed ratings: About 205 stall-free opinion scores remained per video after outlier rejection, and the MOS distribution substantially spanned the quality spectrum.Videos were denser in the MOS range 60-80.
  • Protocol validation: The golden-video comparison produced a mean SROCC of 0.99 between worker MOS and LIVE VQA ground-truth MOS.The authors use this excellent agreement to validate the crowdsourced experimental protocol.

2) Overall inter-subject consistency:

The authors assessed consistency by split-sample MOS comparisons and examined stalls through DMOS, finding stable MOS estimates beyond roughly 200 ratings and lower scores for stalled playback.

  • Overall inter-subject consistency: The average SROCC between MOS values from two equal, disjoint score halves was 0.984 across 100 repetitions.This experiment measured overall subject consistency across all videos.
  • Impact of experimental parameters: Increasing the sample size beyond 200 ratings did not improve or otherwise affect the MOS error-bar and standard-deviation figures.Similar behavior was observed across the rest of the videos.
  • Stall comparison: DMOS was computed as MOS without stalls minus MOS with stalls and plotted against video index.The comparison separated playback-related quality changes from the video’s non-stalled ratings.
  • Stall impact: Stalls lowered MOS for more than 95% of videos.Rare small increases were treated as noise because their causes were difficult to establish.
  • Scope boundary: The study collected per-video total stall durations but not the number or locations of stalls, leaving stalled-rating analysis for future work.The authors nevertheless regarded the stalled data as valuable for future stall-aware quality models.

E. Worker Parameters

Display resolution, device, viewing distance, and gender showed high agreement or small MOS differences, whereas age affected MOS distributions, with younger participants assigning lower scores.

  • High vs low resolution pools: The high- and low-resolution participant pools achieved SROCC 0.97 on 475 common videos, with a mean MOS difference close to 1.The high-resolution pool represented 31.15% of participants and rated all 585 videos.
  • Participants’ Resolution: The two dominant display resolutions, 1366x768 and 1920×1080, produced SROCC 0.95 with a mean MOS difference close to zero.This supported the conclusion that participant display resolution did not significantly affect ratings.
  • Display device: Laptop and computer-monitor users produced SROCC 0.97 between display-device groups, with a mean MOS difference close to 1.Display device did not noticeably impact MOS.
  • Viewing distance: Viewing-distance categories yielded SROCC values from 0.91 to 0.97, while average MOS differences remained below 1.The categories were small, medium, and large viewing distances.
  • Other demographic information: Gender groups had SROCC 0.97 and an average MOS difference of about 2, with female participants tending to give slightly lower scores.The authors reported no noticeable difference between the MOS distributions.
  • Other demographic information: Age affected MOS distributions: younger participants gave lower scores than older participants, and larger age gaps corresponded to lower SROCC.Participants younger than 20 tended to assign lower quality scores than participants older than 40.

V. PERFORMANCE OF VIDEO QUALITY PREDICTORS

The study evaluates leading blind VQA models on the LIVE-VQC database using scatter plots and three performance metrics. V-BLIINDS performs best across PLCC, RMSE, and SROCC, but substantial room for improvement remains.

  • Database and evaluation: NIQE and VIIDEO provide training-free quality scores, whereas V-BLIINDS and BRISQUE use SVR mappings learned from feature spaces to ground-truth MOS.The compared models differ in whether they require training before producing quality predictions.
  • Prediction behavior: VIIDEO predictions correlate poorly with ground-truth MOS, while NIQE, BRISQUE, and V-BLIINDS follow more regular trends against MOS.Figure 13 presents scatter plots for the four models; -NIQE is used because NIQE increases with distortion.
  • Database and evaluation: The evaluation uses PLCC, RMSE, and SROCC to quantify agreement between predicted quality values and MOS distributions.PLCC is computed after applying a prescribed nonlinear mapping to predicted quality values.
  • Prediction behavior: V-BLIINDS supplies the best performance across all three metrics, although existing models still leave ample room for improvement on authentic real-world distortions.Results exclude three videos for VIIDEO and are reported on a 553-video subset for some analyses; NIQE and BRISQUE also run on the full database.
Loading 1803.01761v2…