Source-linked AI summary

How Accurate are Video Quality Models for Diffusion-Based Video Super-Resolution?

Benjamin Herb, Steve Göring, Alexander Raake, Rakesh Rao Ramachandra Rao

arXiv:2605.25940v1eess.IVcs.CV

TL;DR

The paper asks whether existing video quality models can reliably evaluate diffusion-based VSR, a setting not covered by prior evaluations. It compares model predictions with subjective ratings across six upscaling methods, compressed and uncompressed sources, and UHD-1 playback. CNN-based full-reference models correlate best within sequences, but all tested models remain insufficient to replace subjective testing.

  • Problem

    Prior evaluations omitted diffusion-based VSR and resolutions above 1080p, leaving the accuracy of existing quality models for these outputs uncertain.

  • Method

    The study compares six upscaling methods on uncompressed, AV1-compressed, and DCVC-RT-compressed low-resolution videos using subjective ratings and full- and no-reference quality models.

  • Results

    CNN-based full-reference models, including LPIPS, DISTS, and CVQA-FR, outperform conventional models and no-reference models in within-sequence comparisons; LPIPS reaches an SRCC of 0.88.

  • Takeaways & Limitations

    None of the tested quality models performs well enough to validate VSR methods or replace complementary subjective testing.

  • Takeaways & Limitations

    The findings require further investigation across broader ranges of source videos, encoding quality levels, and VSR methods.

Abstract

from arXiv · show

Recent video super-resolution (VSR) approaches use deep neural networks to enhance low-quality input videos and recover visual detail, with diffusion-based methods in particular showing promising results. In this paper, we investigate whether existing video quality models can be used to assess the performance of these diffusion-based VSR methods, by comparing model predictions with results from a subjective test. The study compares six upscaling methods (Lanczos, Rhea, SCST, DOVE, SeedVR2, Starlight Mini) applied to both compressed (AV1 and DCVC-RT) and uncompressed low-resolution videos considering the play-out on a UHD-1/4K screen. A range of full- and no-reference quality models are used to assess their applicability to this new type of quality degradation, focusing on within-sequence performance. The results highlight that CNN-based full-reference models, such as LPIPS, DISTS, and CVQA-FR show significantly higher correlation coefficients than both conventional full- as well as the tested no-reference models. Most overestimate the overly sharp results of SCST, with VMAF mainly failing due to spatial inconsistencies introduced by Starlight Mini. None of the tested video quality models reach sufficient accuracy so as to replace complementary subjective testing. The reference, degraded and upscaled videos, as well as the user ratings and model scores are made available with the paper at https://github.com/Telecommunication-Telemedia-Assessment/AVT-VQDB-UHD-1-VSR as open data.

I. INTRODUCTION AND RELATED WORK

Existing VSR quality studies have not covered diffusion-based methods or resolutions above 1080p. This paper addresses those gaps with subjective evaluation of recent VSR methods across diverse degradations and UHD-1 videos, then assesses existing quality models.

  • Deep-learning VSR methods increasingly use architectures including 3D CNNs, encoder-decoder structures, recurrent networks, and generative adversarial networks.
  • A prior expert study found that subjective results did not particularly align with model results.
  • Prior VSR evaluations did not include diffusion-based methods or resolutions above 1080p.
  • The paper evaluates recent VSR methods using subjective testing, diverse source degradations, and high-resolution 4K/UHD-1 videos.

II. TEST DESIGN

The test design applies VSR to both uncompressed and compressed source videos, using conventional and neural codecs. Quality models are assessed for both within-sequence validation and overall comparison.

  • The subjective test applies VSR approaches to both uncompressed and compressed source videos.
  • Both conventional and neural video codecs are used to examine potential differences in upscaling performance.
  • Six upscaling methods, including Lanczos as a comparison reference, are evaluated.
  • Quality-model suitability is assessed for within-sequence validation and overall evaluation.

A. Videos

The study uses six 4K/UHD-1, 60 fps source clips and creates high-quality uncompressed baselines alongside AV1 and DCVC-RT compressed versions.

  • Six 8–10-second, 4K/UHD-1, 60 fps source clips are selected from the AVT-VQDB-UHD-1 dataset.
  • Uncompressed baselines are created by directly upscaling 360p and 720p sources to 4K/UHD-1 at 3× and 6×.
  • AV1 provides the conventional codec baseline, while DCVC-RT represents a recent neural video codec.

B. Upscaling Methods

Six methods upscale low-resolution videos to 2160p: Lanczos, two commercial methods, and three literature-based VSR approaches. The methods differ in detail recovery, stability, temporal consistency, and processing speed.

  • Six methods upscale low-resolution videos to 2160p, with Lanczos serving as the conventional reference.
  • The VSR methods comprise SCST, DOVE, SeedVR2, TopazLab Rhea, and TopazLab Starlight Mini.
  • SCST often produces overly sharpened results and noticeable temporal consistency issues, while its processing speed averages 96 s/frame.
  • DOVE is a one-step diffusion model using a text-to-video model as its prior and latent- then pixel-space refinement.
  • SeedVR2 converts a 64-step teacher diffusion model into a one-step generator and processes video at 11 s/frame.
  • Rhea generally produces the most stable results but offers less detail-recovery potential than diffusion-based methods.

C. Quality Models

The study evaluates conventional, improved conventional, and deep learning-based full- and no-reference image and video quality models for video assessment. The models span natural-scene statistics, transformers, CNNs, recurrent networks, ensembles, CLIP embeddings, and an LLM.

  • PSNR, SSIM, and MS-SSIM provide the conventional full-reference baseline for quality assessment.
  • Improved conventional IQA models include PSNR-HVS, SSIMULACRA2, and Butteraugli.
  • Natural-scene-statistic no-reference IQA models include BRISQUE and NIQE.
  • Deep no-reference models cover transformer-based IQA, CNN-recurrent VQA, CNN ensembles, and CNN-latent approaches.The listed models include MUSIQ, FAST-VQA, FasterVQA, MDTVSFA, UVQ, and CVQA-NR.
  • CLIP-IQA+, MaxVQA, and Q-Align use CLIP embeddings or an LLM for quality prediction.

D. Experimental Procedure

The subjective study used 32 participants who rated videos on a UHD monitor using a five-point ACR procedure in a controlled environment. Testing lasted 45–60 minutes per participant and included vision screening.

  • 32 participants completed the subjective quality study in a controlled environment.
  • Participants used the five-point absolute category rating method, with testing lasting 45–60 minutes and including a short break.
  • Videos were displayed on a 43-inch Asus XG43UQ UHD monitor at a fixed viewing distance of 1.5H.
  • Participants completed a FrACT10 vision test before rating the videos and were compensated for their participation.

III. SUBJECTIVE QUALITY ASSESSMENT

The subjective assessment examined rating consistency and the perceived quality of six upscaling methods across compressed and uncompressed videos. SeedVR2, DOVE, and Starlight Mini performed best overall, while SCST performed worst and improvements varied by codec, resolution, and sequence complexity.

  • The rating distribution was approximately normal with a tendency toward lower ratings, and SOS analysis estimated a value of 0.254.
  • SeedVR2, DOVE, and Starlight Mini achieved the best overall upscaling performance without significant differences across all tested settings.
  • SCST performed worst overall, with better results on lower-quality source videos than on higher-quality ones.The paper suggests that greater noise at lower quality levels may mask artifacts.
  • All methods performed significantly better on uncompressed low-resolution videos, with SeedVR2 reaching source-comparable quality when upscaling from 360p.
  • Rating gains from Lanczos were higher for AV1 than DCVC-RT, and sequence complexity strongly affected improvements.At 360p, higher-temporal-complexity sequences improved by less than 0.5, whereas some less-complex AV1 sequences improved by more than 1.0.
  • For uncompressed videos, multiple methods matched or surpassed the perceived quality of original UHD-1 sequences from 720p.Results from 360p sources were only slightly lower.

IV. OBJECTIVE QUALITY ASSESSMENT

Within-sequence comparisons favor CNN-based full-reference models, but both full- and no-reference models retain method-specific biases that limit reliable VSR validation.

  • LPIPS achieves the highest within-sequence SRCC of 0.88, with CVQA-FR and DISTS showing comparable results.
  • CNN-based full-reference models significantly outperform conventional models, likely because they are more invariant to slight texture changes introduced by upscaling.
  • Full-resolution pixel-space models such as PSNR, SSIM, Butteraugli, and VMAF degrade when Starlight Mini introduces minor spatial inconsistencies.
  • VMAF overpredicts SCST's oversharpened results, while its NEG variant reduces this effect.
  • Despite relatively high CNN-based full-reference correlations, method-dependent biases make these models unreliable for validation, and no-reference models are not accurate enough either.
  • FasterVQA has the highest no-reference mean SRCC of 0.68, while most no-reference models struggle with SCST and scene-complexity differences.

V. CONCLUSION

The study finds substantial gaps in current quality models for diffusion-based VSR evaluation. CNN-based full-reference models perform best within sequences, but subjective testing remains necessary and broader generalization is unresolved.

  • All tested models show relatively weak overall correlation, while CNN-based full-reference models outperform other architectures for within-sequence comparisons.
  • Current models are insufficient for validating new VSR methods without additional subjective testing.
  • Future work should test whether these findings generalize across broader source-video sets, encoding-quality levels, and VSR methods.
Loading 2605.25940v1…