Source-linked AI summary
MDFI: A Multi-Domain Features Integration for Compressed Video Quality Enhancement
Sang NguyenQuang, Hieu Bui Minh, Dang BuiDinh, Xiem HoangVan
TL;DR
H.266/VVC improves compression efficiency but still produces blocking and blurring artifacts at low bitrates. MDFI addresses this with prediction-guided multi-domain feature integration and FPFT, and experiments report superior enhancement performance over state-of-the-art methods.
Problem
H.266/VVC still produces noticeable blocking and blurring artifacts under aggressive compression, limiting perceptual quality and downstream processing reliability.
Method
MDFI jointly leverages compressed-domain prediction signals, spatiotemporal characteristics, and frequency information through FPFT and a multi-recursive propagation architecture.
Results
0.85 dB in ∆PSNR and 1.68 in ∆SSIM are achieved by MDFI at QP = 37, surpassing OVQE and Wang et al. by 15% and 13%.
Takeaways & Limitations
The MDFI framework and its dedicated dataset provide a foundation for compressed video quality enhancement using prediction information from H.266/VVC bitstreams.
Abstract
from arXiv · showhide
The latest video coding standard, H.266/VVC, has demonstrated significant improvements in compression efficiency compared to H.265/HEVC. Despite its advanced coding techniques, H.266/VVC still faces challenges in meeting the increasing demand for higher perceptual quality and enhanced compression performance. To address these limitations, we propose MDFI (Multi-Domain Features Integration), a compressed video quality enhancement approach that features a novel Frame-Prediction Feature Transform (FPFT) module to process prediction information. Moreover, MDFI integrates a multi-domain feature fusion strategy that effectively combines spatiotemporal characteristics, cross-frequency representations, and compressed-domain prediction information to enhance decoded video quality. Additionally, we introduce a comprehensive dataset that encompasses uncompressed video sequences, corresponding reconstructed versions at multiple QP levels, and predicted frames generated from H.266/VVC compressed bitstreams, providing essential resources for developing and benchmarking video enhancement approaches. Extensive experiments demonstrate that our MDFI approach achieves superior performance to state-of-the-art methods in both objective metrics and visual quality, effectively mitigating video compression artifacts. The code is available at: https://github.com/dangdinh17/MDFI.git.
I. Introduction T
H.266/VVC improves compression efficiency but still produces artifacts, while existing enhancement methods underuse prediction information and lack specialized datasets. MDFI addresses these gaps by integrating prediction, spatiotemporal, and frequency information and introducing a dedicated dataset.
- Motivation: H.266/VVC achieves approximately 50% bitrate reduction versus H.265/HEVC at comparable perceptual quality but still produces blocking and blurring at low bitrates.Advanced partitioning, prediction, transform, and entropy-coding tools support this efficiency improvement.
- Motivation: Effective artifact removal is essential for improving visual quality and downstream processing reliability.
- Research Gap: Existing methods underuse prediction information in compressed bitstreams, lack H.266/VVC specialization, and lack comprehensive datasets exposing internal encoding information.These limitations hinder methods that leverage compressed-domain features.
- Proposed Approach: MDFI jointly leverages prediction, spatiotemporal, and frequency-domain information through a multi-recursive propagation architecture.The framework performs prediction-guided feature learning in the pixel domain without decoder modification and establishes long-range dependencies across frames and domains.
- Proposed Approach: The FPFT module processes prediction signals extracted from the bitstream and adaptively fuses them with current-frame features.
- Dataset: The proposed dataset contains raw sequences, decoded videos at multiple quality levels, and prediction frames extracted from H.266/VVC bitstreams.It provides multi-domain training data with access to internal coding information.
II. Related works
Related work progresses from single-frame spatial enhancement toward multi-domain representations, but spatial methods remain limited because they ignore correlations between adjacent frames. This omission can cause temporal inconsistency, especially under complex motion or rapid scene changes.
- Spatial Methods: Early compressed-video enhancement methods adapt single-image models and process each frame independently using within-frame spatial correlations.Typical tasks include denoising, compression-artifact removal, and deblocking.
- Spatial Methods: Some single-frame models combine pixel-domain and frequency-domain features extracted from DCT coefficients to characterize transform-domain artifacts.
- Limitations: Single-frame spatial methods ignore adjacent-frame spatiotemporal information and therefore cannot exploit inter-frame redundancy or motion continuity.
- Limitations: This limitation is particularly severe with complex motion, rapid scene changes, or long-term temporal dependencies, where temporal flickering and inconsistent visual quality may result.
B. Spatiotemporal Methods
Spatiotemporal methods improve enhancement by extracting correlations across adjacent frames, with deformable sampling and recursive fusion improving motion handling and temporal propagation. However, recursive approaches do not exploit the entire video's spatiotemporal information.
- Spatiotemporal Methods: Spatiotemporal methods use multiple adjacent frames instead of a single frame to extract spatial and temporal correlations.
- Spatiotemporal Methods: MFQE introduced a Multi-Frame CNN, while MFQEv2.0 added a Bidirectional Long Short-Term Memory network and became a benchmark for later work.
- Motion Modeling: Deformable convolution captures complex motion across multiple frames more effectively than standard CNNs, supporting later models such as TGAF and TVQE.
- Recursive Methods: RFDA recursively combines compensated features with current features for performance gains, but does not exploit spatiotemporal information from the entire video.
C. Multi-Domain Methods
Multi-domain methods address the omission of coding information in decoded-frame enhancement by incorporating priors embedded in compressed bitstreams. These priors provide explicit cues about inter-frame dependencies and compression distortions.
- Motivation: Many spatiotemporal models operate only on decoded frames and neglect coding information embedded in compressed bitstreams.This omission can limit further performance gains because bitstreams contain temporal and spatial priors from encoding.
- Coding Priors: Coding priors include motion vectors, prediction modes, residual signals, and reference-frame information.These cues explicitly describe inter-frame dependencies and compression distortions.
- Prior Work: Prior studies showed that codec-related information can enhance reconstruction quality in video super-resolution when properly exploited.
A. System Overview
MDFI enhances decoded H.266/VVC frames by combining prediction information with neighboring decoded frames and frequency-aware refinement. The pipeline progressively fuses these representations and reconstructs the enhanced frame from a predicted residual.
- Enhancement target: H.266/VVC reconstruction can contain blocking artifacts, blurring, and lost fine texture because compression is lossy.MDFI targets quality degradation introduced between the original and reconstructed sequences.
- System inputs: MDFI takes the target decoded frame, seven-frame temporal context, and its corresponding prediction frame as inputs.The neighboring-frame radius is R = 3, giving 2R+1 = 7 decoded frames.
- Feature pipeline: The framework extracts prediction, spatiotemporal, and frequency-aware features through FPFT, STFF, GMFF, and FQE stages.FPFT processes prediction information; subsequent modules aggregate temporal context, refine global features, and estimate the enhancement residual.
- Feature pipeline: GMFF refines spatiotemporal features into globally enriched representations before FQE estimates the residual between decoded and high-quality content.The globally enriched representation is passed to FQE for residual estimation.
- Output reconstruction: The enhanced frame is reconstructed by adding the predicted residual to the decoded frame.This residual reconstruction completes the quality-enhancement process.
B. Prediction Information Exploitation
MDFI exploits prediction information recreated from H.266/VVC bitstreams alongside decoded-frame context. Its FPFT module uses prediction-guided adaptive transformations, while later stages integrate spatiotemporal and frequency features for residual enhancement.
- Prediction Information Exploitation: H.266/VVC prediction information estimates current blocks from reconstructed samples and is recreated at the decoder using bitstream control data.The control data includes coding modes and motion vectors, enabling reconstruction of prediction frames.
- Prediction Information Exploitation: FPFT jointly processes the decoded target frame and corresponding prediction frame to produce prediction features for later enhancement stages.These prediction features provide structural and motion priors for restoring details and aligning temporal features.
- Prediction Information Exploitation: Spatially adaptive affine transformations use prediction information to modulate feature maps through scaling and shifting parameters generated by convolutions.The transformation is expressed as ˆf^j_i = γ ⊙ F^j_i + β, with γ and β representing scaling and shifting.
- Spatio–Temporal–Frequency Exploitation: STFF concatenates refined prediction features with neighboring decoded frames, then uses multi-scale fusion and deformable sampling to capture spatiotemporal dependencies.Its U-Net-based structure aggregates complementary features through SKFF modules and refines them with DCN.
- Spatio–Temporal–Frequency Exploitation: GMFF uses bidirectional grid-based propagation, integrating past, present, and future context across four propagation stages.Each stage applies STFF for alignment and OFAE for omni-frequency enhancement.
- Spatio–Temporal–Frequency Exploitation: FQE estimates a residual from propagated features and reconstructs the enhanced frame by adding it to the decoded frame.OFAE blocks support frequency-component reconstruction within the lightweight quality-enhancement network.
- Dataset construction: The MDFI dataset contains raw and reconstructed sequences plus prediction information from H.266/VVC bitstreams.It includes 126 sequences, split into 108 training and 18 testing sequences, encoded at QPs 22, 27, 32, and 37.
A. Implementation Details
The implementation trains MDFI on cropped multi-frame clips with standard augmentation and Adam optimization. Evaluation focuses on Y-channel quality gains and rate–distortion performance, while the dataset is publicly identified for research use.
- Implementation Details: Training uses seven adjacent input frames and randomly cropped 128 × 128 sub-frame clips from raw, reconstructed, and predicted videos.Each training sample contains 15 consecutive frames with random flipping and rotation augmentation.
- Implementation Details: The models are optimized with Adam using β1 = 0.9, β2 = 0.999, and ϵ = 10^-6.The supplied passage also specifies an initial learning rate of 1 × 10...
- Evaluation Methodologies: Evaluation enhances only the Y-channel in YUV 4:2:0 and reports ΔPSNR, ΔSSIM, and BD-rate.These measures capture quality improvement and rate–distortion performance.
- Dataset availability: The MDFI dataset is made available through Kaggle.The dataset link is identified in the implementation details.
- Dataset comparison: The dataset comparison uses the commonly adopted MFQEv2 subset of 108 training sequences and 18 testing videos.The cited passage contrasts this subset with the original MFQEv2 split.
B. Experimental Results
MDFI consistently improves objective, perceptual, visual, and rate–distortion performance over competing methods, while exposing a quality–complexity trade-off. Its complete configuration performs best, whereas lighter variants reduce computation at modest quality cost.
- Overall Quality Enhancement: At QP = 37, MDFI achieves the best averages of 0.85 dB in ∆PSNR and 1.68 in ∆SSIM, surpassing OVQE and Wang et al. by 15% and 13%.It also improves PSNR by 0.25 dB over TVQE and 0.23 dB over TGAF.
- Subjective Quality Performance: MDFI achieves superior visual quality by preserving fine details, structural boundaries, and texture consistency across representative sequences.In BasketballPass, competing methods over-smooth hand–ball boundaries, whereas MDFI preserves sharper details.
- Rate–Distortion Performance: MDFI achieves a nearly 24.66% BD-rate reduction, approximately 47% higher than TVQE and TGAF.Rate–distortion curves further show higher PSNR at similar bitrate on four test sequences.
- Model Complexity and Ablation Study: Increasing SFTRBs and channels improves ∆PSNR but raises computational cost; MDFI S2 uses 24.4% fewer FLOPs than the baseline, while MDFI S3 halves cost and latency.The full model maximizes enhancement quality but has the highest efficiency ratio and largest parameter and FLOP requirements.
- Ablation Study: The complete MDFI configuration improves ∆PSNR by 0.85 dB and ∆SSIM by 1.68 × 10^-2, outperforming the variant without FPFT by 0.11 dB and 0.19.The FQE-only design still provides a 0.35 dB PSNR gain.
- Perceptual Quality Analysis: MDFI consistently achieves the highest ∆VMAF and lowest ∆LPIPS, with stronger perceptual gains under high compression.The evaluation covers perceptual quality alongside distortion-based metrics.
V. Conclusions
MDFI combines spatiotemporal correlations with compressed-domain prediction information for H.266/VVC quality enhancement. Its FPFT module and dedicated dataset support improved quality and reduced artifacts, with gains over existing methods in objective and subjective evaluations.
- MDFI captures correlations between adjacent reconstructed frames while directly leveraging H.266/VVC compressed-domain prediction information.
- The FPFT module integrates multi-domain features and exploits long-term dependencies across multiple frames and prediction signals.
- The dedicated dataset includes raw–reconstructed sequence pairs and prediction frames extracted from H.266/VVC bitstreams.
- Extensive experiments show that MDFI improves video quality and reduces compression artifacts, outperforming existing state-of-the-art methods in objective and subjective evaluations.