Source-linked AI summary
Video Coding for Machines: A Paradigm of Collaborative Compression and Intelligent Analytics
Ling-Yu Duan, Jiaying Liu, Wenhan Yang, Tiejun Huang, Wen Gao
TL;DR
The paper addresses the gap between pixel-oriented video coding for human vision and compact feature coding for machine analytics. It formulates Video Coding for Machines (VCM), reviews video and feature compression, and proposes collaborative architectures; preliminary experiments report efficiency and performance advantages, including a 5.2 Kbps action-recognition result.
Problem
Separate video and feature coding leaves collaborative compression between human-vision and machine-vision streams insufficiently explored for large-scale intelligent analytics.
Method
The paper defines and formulates VCM, reviews standardized and learning-based compression, and proposes collaborative video-feature architectures using feedback and multiple streams.
Results
Preliminary VCM experiments demonstrate improved coding efficiency and analytics performance, with action recognition achieved at 5.2 Kbps using the proposed method.
Takeaways & Limitations
VCM provides an initial framework for jointly optimizing video and feature streams across human- and machine-oriented tasks.
Abstract
from arXiv · showhide
Video coding, which targets to compress and reconstruct the whole frame, and feature compression, which only preserves and transmits the most critical information, stand at two ends of the scale. That is, one is with compactness and efficiency to serve for machine vision, and the other is with full fidelity, bowing to human perception. The recent endeavors in imminent trends of video compression, e.g. deep learning based coding tools and end-to-end image/video coding, and MPEG-7 compact feature descriptor standards, i.e. Compact Descriptors for Visual Search and Compact Descriptors for Video Analysis, promote the sustainable and fast development in their own directions, respectively. In this paper, thanks to booming AI technology, e.g. prediction and generation models, we carry out exploration in the new area, Video Coding for Machines (VCM), arising from the emerging MPEG standardization efforts1. Towards collaborative compression and intelligent analytics, VCM attempts to bridge the gap between feature coding for machine vision and video coding for human vision. Aligning with the rising Analyze then Compress instance Digital Retina, the definition, formulation, and paradigm of VCM are given first. Meanwhile, we systematically review state-of-the-art techniques in video compression and feature compression from the unique perspective of MPEG standardization, which provides the academic and industrial evidence to realize the collaborative compression of video and feature streams in a broad range of AI applications. Finally, we come up with potential VCM solutions, and the preliminary results have demonstrated the performance and efficiency gains. Further direction is discussed as well.
I. INTRODUCTION
Massive-scale video analytics exposes a tension between compressing pixels for human viewing and transmitting compact features for machine tasks. The paper formulates Video Coding for Machines (VCM) as collaborative optimization of video and feature streams, reviews relevant standards and techniques, and proposes initial architectures.
- Massive video analytics must address low latency and high accuracy while efficiently managing data from smart-city and IoT applications.
- Traditional video coding reconstructs pixels for human vision, whereas feature compression transmits compact information for machine-oriented analysis.
- Full-resolution video storage and decompression can be prohibitive at scale, while lowering compressed-video quality risks degraded analytics performance.
- Separate optimization of video and feature coding is sub-optimal for applications serving both human and machine vision.
- The paper reviews standardized video and feature compression, deep prediction and generation models, and exemplar VCM architectures with preliminary experiments.
- VCM identifies tasks, features, and resources, introducing feedback for collaborative and scalable modes to improve coding efficiency and analytics performance.
B. Deep Learning Based Video Coding
Deep learning advances video coding through data-driven prediction, generation, and end-to-end optimization, but reconstructing complete pictures remains inefficient for large-scale analytics. These capabilities motivate collaborative approaches that transmit task-relevant features at low bitrate.
- Deep learning improves video coding by learning data-driven priors with powerful architectures, large-scale data, and end-to-end optimization.
- Learned coding tools replace fixed partition-dependent processing with broader-context, hierarchical representations and naturally avoid blocking artifacts.
- Neural prediction models support intra-prediction, inter-prediction, deblocking, and fast mode decision, achieving superior rate-distortion performance.
- Reconstructing whole pictures remains unsuitable for real-time analytics on massive datasets because low-value data constitutes much of the volume.
- Feature compression targets low-bitrate recognition, classification, and retrieval when raw visual-feature transmission bandwidth is limited.
- MPEG-7 descriptors evolved from low-level tools toward CDVS and CDVA compact descriptors enabled by advances in computer vision and deep learning.
A. CDVS
MPEG’s compact visual descriptor standards target efficient machine-oriented search and video analysis under stringent bitrate, memory, and computational constraints. CDVS and CDVA establish standardized feature pipelines and descriptors, while their limitations motivate collaborative video-feature compression.
- CDVS origins and standardization: CDVS standardizes compact visual descriptors for mobile visual search and augmented reality, reducing transmitted data and potentially latency.Early research reported at least an order-of-magnitude transmission reduction by extracting descriptors on-device; the standard was published in 2015.
- CDVS performance: CDVS supports six predefined descriptor lengths from 512B to 16KB while satisfying stringent memory and computational complexity requirements.The standard was evaluated on a large-scale image matching and retrieval benchmark.
- CDVS pipeline: CDVS defines a normative feature-extraction pipeline spanning interest-point detection, descriptor construction, aggregation, and compression.The pipeline is designed to guarantee interoperability across implementations.
- CDVS limitation: CDVS’s fixed encoder can underperform task-specific end-to-end features on complex autonomous-driving and video-surveillance analyses.Collaborative video-feature compression is proposed to combine normative feature encoding with normative video decoding.
- CDVA scope and challenge: CDVA standardizes neural-network-based video feature descriptors, initially narrowing its scope to video matching and retrieval.Its primary challenge is the absence of a generic deep model that serves a broad range of analytics tasks.
- CDVA technologies and results: NIP global descriptors combined with CDVS descriptors significantly improve performance at comparable descriptor lengths.CDVA also integrates temporal coding and hand-crafted and deep-learning features; the standard was published in 2019.
- CDVA extensions: CDVA extensions address higher-resolution analysis and raw Bayer-pattern processing, supporting machine-only and hybrid human-machine encoding formats.The stated benefits include reduced network and storage burdens, lower latency, and reduced cost across applications.
IV. VIDEO CODING FOR MACHINES: DEFINITION, FORMULATION AND PARADIGM
VCM formulates collaborative compression across tasks, features, and resources to support human perception and machine intelligence. Its optimization jointly balances task performance against communication, computation, storage, bandwidth, and model-related costs.
- A. Definition and Formulation: VCM addresses the gap between full-resolution video coding for human vision and compact feature coding for machine tasks.Traditional video coding prioritizes visual fidelity, whereas feature coding prioritizes task performance at very low bitrates.
- A. Definition and Formulation: VCM jointly optimizes multiple low-level and high-level vision tasks while minimizing communication and computational resources.The framework is intended to economize a shared bit budget across human and machine-oriented tasks.
- A. Definition and Formulation: VCM includes task objectives spanning low-level signal fidelity and high-level semantics, requiring an efficient coding scheme across tasks.The formulation treats multiple task requirements within a complex optimization objective.
- A. Definition and Formulation: VCM treats bandwidth, computation, storage, and energy as practical resource constraints that traditional video or feature coding alone may not solve well.Resource usage is part of the problem definition rather than an afterthought.
- A. Definition and Formulation: VCM unifies features at different granularities, including pixels for human observers and semantic or syntactic features for machine tasks.The pixel feature is associated with visual signal fidelity, while other features serve task-specific functions.
- A. Definition and Formulation: A smaller feature index i denotes a more abstract feature, and each task’s performance is represented by a quality metric over its reconstructed feature.The reconstructed feature undergoes the relevant encoding and decoding processes.
- A. Definition and Formulation: The formulation models compression, decompression, and feature prediction through parameterized processes C, D, and G.G projects an input feature into the space of a target feature, while the hat notation denotes compression loss after decoding.
- A. Definition and Formulation: The resource constraint includes feature compression and transmission, video coding with prediction, model overhead, and complexity costs such as training time and feedback delay.The total resource budget also accounts for larger training datasets associated with more complex models.
B. A VCM Paradigm: Digital Retina
Digital Retina coordinates video, model, and feature streams to support human and machine vision, while VCM adds joint optimization and feedback for collaborative, scalable compression.
- B. A VCM Paradigm: Digital Retina: Digital Retina establishes video, model, and feature streams for human and machine vision.The video stream supports reconstructed human viewing, the model stream guides updating and prediction, and the feature stream carries task-specific information.
- B. A VCM Paradigm: Digital Retina: The feature stream transmits task-specific semantic or syntactic features from front-end devices to servers.It can reduce bandwidth by sending task-specific features and shift part of feature computation to the front end.
- B. A VCM Paradigm: Digital Retina: Traditional video and feature coding use separate, largely uni-directional optimization, limiting cross-stream utilization and optimization performance.The stated limitations include separate handling of streams and the absence of feedback mechanisms.
- B. A VCM Paradigm: Digital Retina: VCM jointly optimizes feature, video, and model streams through feedback mechanisms in collaborative and scalable modes.Collaborative feedback enables cross-stream optimization, while scalable feedback addresses insufficient bit budgets.
V. NEW TRENDS AND TECHNOLOGIES
Recent VCM-relevant technologies draw on hierarchical deep networks, generative models, and temporal video-analysis architectures to represent and transform visual information.
- 1) Deep Networks: Deep convolutional networks extract discriminative features for image and video tasks including classification, detection, segmentation, and pose estimation.Their hierarchical structures progressively process visual information for semantic prediction.
- 1) Deep Networks: Video-analysis networks model temporal dynamics and joint spatial-temporal correlations for semantic label prediction.Representative architectures extend CNN connectivity, combine appearance and motion, or use CNN-RNN and 3D convolutional designs.
- 1) Generative Adversarial Networks: Generative adversarial networks support supervised and unsupervised image generation, including image-to-image translation and cross-domain mapping.Supervised methods can use GANs as learned loss functions, while unsupervised methods use cycle reconstruction consistency.
- 1) Generative Adversarial Networks: Guided image-generation methods use semantic inputs such as human pose or semantic maps to generate and refine images.Pose-guided approaches progressively transform high-level body-part features to model shapes and appearances.
3) Video Prediction and Generation:
Video prediction and generation provide temporal modeling tools for VCM, while progressive and scalable representations support feature granularity ranging from compact machine information to richer human-oriented reconstruction.
- 3) Video Prediction and Generation: Video prediction generates future frames deterministically from previous frames, commonly using recurrent networks to model temporal dynamics.Other approaches use 3D convolutional networks or estimate local and global transforms.
- 3) Video Prediction and Generation: Video generation produces visually authentic sequences probabilistically using GAN- and VAE-based methods.Later work extends prediction toward long-term generation, including up to 100 future Atari frames.
- 3) Video Prediction and Generation: Prediction and generation models can leverage compact features to propagate video context and improve coding efficiency.This connects temporal modeling with inter prediction in video coding.
- 3) Video Prediction and Generation: VCM requires scalable feature prediction that ranges from extreme compactness to redundant visual data for human and machine tasks.Hierarchical deep processing and progressive representation are identified as relevant to this scalability.
- 3) Video Prediction and Generation: Progressive image and feature methods refine representations from coarse to fine through pyramids, residual networks, or multi-level feature fusion.Related applications include super-resolution, dehazing, inpainting, artifact removal, and deblurring.
- 3) Video Prediction and Generation: Scalable image and video coding forms bitstreams with a base layer and enhancement layers that progressively improve reconstruction quality.VCM extends this scalability to collaborative coding of multiple task-specific streams.
- 3) Video Prediction and Generation: The paper presents exemplar VCM solutions spanning intermediate feature compression, predictive coding with collaborative feedback, and scalable feedback.These solutions use deep learning features and prediction or generation mechanisms to connect machine and human-oriented coding.
A. Deep Intermediate Feature Compression
Deep intermediate feature compression transmits task-general representations and combines sparse-point prediction, generated frames, and residual coding to serve machine analysis and human vision.
- A. Deep Intermediate Feature Compression: Intermediate features are compressed instead of original video or highly task-specific top-layer features to support a wider range of machine vision tasks.Shallow intermediate layers generally retain more informative cues, while deep features are more task-specific.
- A. Deep Intermediate Feature Compression: Feature recomposition seeks to evaluate and organize features across tasks by sharing common representations in a scalable manner.The goal is to address differing feature granularities across tasks.
- A. Deep Intermediate Feature Compression: The joint compression pipeline selects key frames, extracts sparse points describing temporal changes and object motion, and uses prediction and generation to synthesize non-key frames.Traditional video codecs compress the key frames, while learned sparse points guide reconstruction of the remaining frames.
- A. Deep Intermediate Feature Compression: A learnable prediction model produces a compact feature F, which is compressed into a feature bitstream under rate control.The compression model encodes F into BF for transmission and storage.
- A. Deep Intermediate Feature Compression: Residual video is formed from the original and synthesized frames, then compressed alongside sparse-point descriptors into video streams.The total bitrate can be adjusted by controlling the feature and residual-video bitrates.
- A. Deep Intermediate Feature Compression: At the decoder, key frames and sparse-point representations are reconstructed, and the final video combines synthesized frames with decoded residuals.The reconstructed feature stream serves machine analysis, while reconstructed video serves human vision.
C. Enhancing Predictive Coding with Scalable Feedback
The scalable-feedback architecture adds guidance features when decompressed video or feature quality is insufficient, enabling incremental refinement of both streams.
- The optimized VCM architecture extracts key points, removes video redundancy through feature-guided prediction, and compresses the residual video.
- When decompressed video and feature quality fails to meet requirements, scalable feedback introduces additional guidance features.
- The updated feature is formed by generating residual features from base-layer key points and the original video to use additionally allocated bits.
- A refinement generator uses reconstructed key points and video to refine the reconstructed video.
- The incremental residue video is compressed into an additional bitstream, and decoding combines it with existing reconstructed components for higher-quality video.
VII. PRELIMINARY EXPERIMENTAL RESULTS
The experiments evaluate intermediate feature compression and machine-human collaborative compression, finding strong lossy feature-compression potential and distinct packaging strategies for deep features.
- The preliminary experiments cover intermediate deep feature compression and collaborative compression for machine and human vision.
- Compression Results: Lossy compression can reach at least 1000× compression for an image-retrieval ResNet conv2 feature at QP 42, versus 2–5× for lossless methods.
- Compression Results: Feature fidelity generally decreases as QP increases, while QP 22 offers high fidelity with a fair compression ratio.
- Compression Results: Upper-layer features from conv4 to pool5 tend to remain more robust under heavy compression.
- Channel Packaging: Deep features are packaged for video codecs through channel concatenation, descending-difference concatenation, or channel tiling.
- Joint Compression: The collaborative evaluation includes action recognition, human detection, and video reconstruction, with results reported in Tables III and IV.
1) Experimental Settings:
The evaluation uses PKU-MMD clips and HEVC anchors to compare feature-assisted coding across action recognition, human detection, and video reconstruction. The proposed method reports better task or reconstruction outcomes at lower bitrate costs.
- Experimental Settings: The setup uses 3317 training clips and 227 testing clips from PKU-MMD, with 32 frames per clip cropped and resized to 512 × 512.
- Experimental Settings: HEVC in FFmpeg 2.8.15 serves as the comparison anchor by compressing frames in constant-rate-factor mode.
- Experimental Settings: Key-point heatmaps from a U-Net and covariance matrices represent sparse motion patterns for machine vision.
- Action Recognition: The proposed method achieves considerable action-recognition accuracy at only 5.2 Kbps, outperforming HEVC.
- Human Detection: Human detection evaluation compares IoU across 64×64, 128×128, 256×256, and 512×512 HEVC inputs at constant rate factor 51.
- Human Detection: The proposed method achieves better human-detection accuracy with fewer bit costs in quantitative and subjective comparisons.
- Video Reconstruction: For reconstruction, key frames use constant rate factor 32 while HEVC uses 44 to obtain approaching bitrate conditions.
- Video Reconstruction: The proposed method achieves better SSIM reconstruction quality than HEVC with lower bitrate cost and better quality across all compared bitrates.
VIII. DISCUSSION AND FUTURE DIRECTIONS
The discussion identifies unresolved questions about task-feature relationships and domain robustness, while positioning VCM as a collaborative framework supported by preliminary results and future research needs.
- High-level tasks favor discriminative compact features, whereas low-level tasks require abundant pixels for fine modeling.
- There is no theoretical evidence yet to measure the information associated with any given task, motivating extensive investigation of task relationships.
- Bio-inspired spike cameras offer a frame-free way to capture fast-moving scenes while reconstructing full texture, suggesting another route across human-machine vision.
- Domain Shift in Prediction and Generation: Data-driven VCM may suffer over-fitting under domain shift, making domain generalization and online domain adaptation important topics.
- Conclusion: The paper formulates VCM as collaborative optimization of video and feature coding for human and/or machine vision, and presents solutions, preliminary results, reviews, and future directions.