Source-linked AI summary

VideoSet: A Large-Scale Compressed Video Quality Dataset Based on JND Measurement

Haiqiang Wang, Ioannis Katsavounidis, Jiantong Zhou, Jeonghoon Park, Shawmin Lei, Xin Zhou, Man-On Pun, Xin Jin, Ronggang Wang, Xu Wang, Yun Zhang, Jiwu Huang, Sam Kwong, C. -C. Jay Kuo

arXiv:1701.01500v2cs.MM

TL;DR

Existing video coding quality models do not capture nonlinear human perception, motivating a larger statistical JND dataset. The paper constructs VideoSet and characterizes its measurements, showing how the dataset supports perceptual quality analysis and data-driven coding research.

  • Problem

    Conventional video coding quality models use continuous rate-distortion functions that do not account for nonlinear human perception, while prior JND datasets were small.

  • Method

    The paper builds VideoSet from 220 five-second sequences at four resolutions, encodes 880 clips with H.264 across QP = 1, …, 51, and measures three JND points with 30+ subjects.

  • Results

    The collected JND data reveal that the first JND point has significantly greater variability than the second and third, which are judged more confidently.

  • Takeaways & Limitations

    VideoSet provides a large-scale basis for perceptual video coding research and points toward data-driven perceptual coding.

  • Takeaways & Limitations

    Standard subject-screening procedures do not apply properly to the collected JND data because they were designed for conventional quality scores rather than JND QP values.

Abstract

from arXiv · show

A new methodology to measure coded image/video quality using the just-noticeable-difference (JND) idea was proposed. Several small JND-based image/video quality datasets were released by the Media Communications Lab at the University of Southern California. In this work, we present an effort to build a large-scale JND-based coded video quality dataset. The dataset consists of 220 5-second sequences in four resolutions (i.e., $1920 \times 1080$, $1280 \times 720$, $960 \times 540$ and $640 \times 360$). For each of the 880 video clips, we encode it using the H.264 codec with $QP=1, \cdots, 51$ and measure the first three JND points with 30+ subjects. The dataset is called the "VideoSet", which is an acronym for "Video Subject Evaluation Test (SET)". This work describes the subjective test procedure, detection and removal of outlying measured data, and the properties of collected JND data. Finally, the significance and implications of the VideoSet to future video coding research and standardization efforts are pointed out. All source/coded video clips as well as measured JND data included in the VideoSet are available to the public in the IEEE DataPort.

1 Introduction

The paper argues that conventional video coding quality models overlook nonlinear human perception and introduces VideoSet to measure perceptual quality statistically through JND points.

  • Video traffic is projected to grow from about 70% of Internet traffic to 80–90%, motivating advances in video coding technology.
  • Traditional R-D modeling assumes a continuous, convex relationship and does not account for nonlinear human perception.
  • Human observers distinguish only discrete distortion levels, producing a stair-shaped perceived R-D curve with JND points.JND is defined as the maximum difference unnoticeable to a human being.
  • Group-based subjective testing yields a statistically aggregated QoE function that is more meaningful than judgments from a few expert viewers.
  • VideoSet contains 220 five-second sequences at four resolutions, with 880 clips encoded across QP = 1, …, 51 and the first three JND points measured using 30+ subjects.The source and coded clips and measured JND data are publicly available through IEEE DataPort.

2 Source and Compressed Video Content

The VideoSet source material was selected and normalized to provide diverse, compatible content across four resolutions, then encoded with controlled H.264 quantization for subjective testing.

  • Source Video: The dataset starts with 220 diverse five-second source clips selected from publicly available datasets.The source material spans multiple original resolutions, frame rates, and color formats, with selection aimed at avoiding redundancy.
  • Source Video: 60 fps sources are converted to 30 fps, while sources at 30 fps or below retain their frame rate for smooth playback.
  • Source Video: Each clip is down-sampled to 1920×1080, 1280×720, 960×540, and 640×360 using Lanczos interpolation and 4:2:0 chroma sampling.
  • Source Video: The processing produces 880 uncompressed sequences, including dominant web formats and lower resolutions representing tablet or mobile viewing.
  • Compressed Video: The 880 sequences are encoded with H.264/AVC high profile using constant QP to isolate the relationship between quantization and perceptual quality.
  • Compressed Video: The practical subjective-test range is QP [8, 47], while QP values below 8 and above 47 are substituted by QP = 0 and QP = 47, respectively.The paper states that this modification does not influence subjective test results.

3 Subjective Test Environment and Procedure

The study used controlled subjective comparisons across 58 stations to locate three JND points per coded video clip through a robust, anchor-based QP search. The procedure improved resistance to uncertain responses by retaining a buffer in the search interval, while using quartile-based anchors for successive JND searches.

  • Test environment: The subjective test used 58 stations in six Shenzhen universities, with controlled laboratory conditions and viewing distance following ITU-R BT.2022.Monitors were not cross-calibrated but were adjusted to comfortable settings; monitor profiling recorded chromaticity, color difference, peak luminance, and luminance ratio.
  • Test procedure: The test partitioned 880 clips into 58 packages, each containing source clips and all coded versions for assigned content-resolution pairs.Sequences were displayed at native resolution without scaling, with inactive screens set to light gray.
  • Test procedure: Around 800 students received training before comparing source and coded clips in sequential YES/NO noticeable-difference judgments.Each subject could replay the pair once, and sessions lasted about 35 minutes with a five-minute break.
  • JND search procedure: The revised search requires eight comparisons instead of six, trading slightly greater cost for increased robustness.The procedure was designed to avoid fixing the JND at an incorrect boundary after an unconfident initial decision.
  • JND search procedure: The robust binary search updates a fixed anchor and comparison QP within a search interval, using quartile updates rather than discarding an entire half after an uncertain response.The new procedure removes only the farthest quarter, preserving a buffer for possible decision errors.
  • JND search procedure: Each clip receives three JND points, whose locations vary with video content, subject discrimination, and viewing environment.For successive JND searches, the current-point histogram’s first quartile becomes the next anchor, representing a threshold that 75% of subjects cannot notice.

4 JND Data Post-Processing via Outlier Removal

VideoSet post-processing removes unreliable subjects and content-specific outlying JND samples before assessing whether the remaining JND measurements are approximately Gaussian. Subject screening uses z-score consistency and dispersion, while sample screening iteratively applies Grubbs’ test.

  • Outlier removal targets both unreliable subjects and collected JND samples to support more reliable conclusions.
  • 4.1 Unreliable Subjects: Samples falling in the lossless QP = [1, 7] interval identify the subject as an outlier, causing all samples from that subject to be removed.
  • 4.1 Unreliable Subjects: The standard ITU-R BT 1788 screening procedure is unsuitable for these JND data because it assumes score types and comparisons that do not apply to QP-based JND measurements.
  • 4.1 Unreliable Subjects: Subject z-scores measure each raw JND sample’s distance from the population mean in standard-deviation units, with their dispersion indicating individual consistency.
  • 4.1 Unreliable Subjects: A subject is flagged when both the range and standard deviation of the z-score vector are large, indicating inconsistent evaluations.
  • 4.2 Outlying Samples: For content-specific outliers, Grubbs’ test removes one sample at a time and repeats until no outliers remain.
  • 4.2 Outlying Samples: With about N = 30 samples and α = 0.05, a sample is identified as an outlier when its distance from the mean exceeds 2.9085 standard-deviation units.
  • 4.3 Normality of Post-processed JND Samples: After post-processing, a great majority of JND points pass the normality test and follow a Gaussian distribution.

5 Discussion

The VideoSet reveals substantial sequence-dependent variation in JND measurements, with content characteristics influencing both dispersion and detectability. Across resolutions, later JND points are more consistent than the first, and the three JND distributions center near QP 27, 31, and 34.

  • JND samples vary substantially across different video sequences.
  • Fast motion and changing backgrounds produce strong masking, yielding large JND variation for sequence #15, whereas salient facial content produces a compact distribution for sequence #37.Sequence #15 has the largest deviation, while sequence #37 has the smallest SD among the 50 sequences.
  • The first, second, and third JND histograms across 220 sequences center around QP 27, 31, and 34, respectively.The authors argue that measuring three JND points is sufficient because quality beyond the third point is too poor for practical streaming.
  • Across all four resolutions, the second and third JND points have significantly smaller SDs than the first JND point.The first JND is harder to determine because slight blurriness is subtle, whereas later judgments involve noticeable blockiness.
  • Masking explains why viewers disagree more on sequences with larger SD values: artifacts are visible to some viewers but hidden from others.

6 Significance and Implications of VideoSet

The VideoSet supports SUR as a perceptual quality measure that can compare coded videos across content and select QP values for a target fraction of viewers. Its larger-scale JND data also motivate data-driven perceptual coding and artifact-masking strategies beyond PSNR.

  • VideoSet construction enables machine-learning prediction of JND values over a short interval, supporting data-driven perceptual coding.
  • Measured JND samples can be converted into an SUR curve, whose target percentile identifies a QP satisfying a chosen percentage of viewers.
  • Under a normal JND model, the SUR curve becomes a Q-function; its top quartile gives a QP whose coded quality is perceptually lossless for 75% of viewers.
  • SUR provides a universal quality metric for comparing coded videos with different content, unlike PSNR values.
  • SUR can reveal artifact causes and guide masking methods intended to shift the first JND point toward a larger QP value.

7 Conclusion and Future Work

The paper constructs the large-scale VideoSet and presents its subjective testing, outlier-processing, and collected JND data. It identifies predicting JND points from video content and understanding coding artifacts as key future steps toward practical data-driven perceptual coding.

  • The paper details construction of the large-scale JND-based compressed video quality dataset called VideoSet.
  • It reports the subjective test procedure, detection and removal of outlying measurements, and properties of the collected JND data.
  • The paper presents VideoSet as a clear path toward data-driven perceptual coding.
  • Predicting the mean and variance of the first, second, and third JND points from video content remains an essential challenge for practical data-driven perceptual coding.
  • Identifying coding artifacts to which humans are sensitive could enable masking methods that shift the first JND point to a larger QP value.
Loading 1701.01500v2…