Source-linked AI summary
Wireless Semantic Communications for Video Conferencing
Peiwen Jiang, Chao-Kai Wen, Shi Jin, Geoffrey Ye Li
TL;DR
The paper addresses limited understanding of physical-channel errors in semantic video conferencing. It establishes SVC and extends it with semantic-error-aware IR-HARQ and CSI-based allocation, with simulations reporting flexible bit consumption, fewer transmitted bits, and improved performance under varying channels.
Problem
Physical-channel transmission errors on keypoints in semantic video conferencing have been insufficiently studied, despite the need for acceptable video under varying channels.
Method
The paper builds a keypoint-based SVC network and adds SVC-HARQ with a semantic error detector plus SVC-CSI for channel-aware keypoint allocation.
Results
The proposed system significantly improves transmission efficiency, while SVC-HARQ adapts to different BERs and transmits fewer bits than competing methods.
Takeaways & Limitations
Semantic video conferencing can combine reduced keypoint-based transmission with feedback and CSI mechanisms to balance performance, bit consumption, and robustness.
Abstract
from arXiv · showhide
Video conferencing has become a popular mode of meeting even if it consumes considerable communication resources. Conventional video compression causes resolution reduction under limited bandwidth. Semantic video conferencing maintains high resolution by transmitting some keypoints to represent motions because the background is almost static, and the speakers do not change often. However, the study on the impact of the transmission errors on keypoints is limited. In this paper, we initially establish a basal semantic video conferencing (SVC) network, which dramatically reduces transmission resources while only losing detailed expressions. The transmission errors in SVC only lead to a changed expression, whereas those in the conventional methods destroy pixels directly. However, the conventional error detector, such as the cyclic redundancy check, cannot reflect the degree of expression changes. To overcome this issue, we develop an incremental redundancy hybrid automatic repeat-request (IR-HARQ) framework for the varying channels (SVC-HARQ) incorporating a novel semantic error detector. The SVC-HARQ has flexibility in bit consumption and achieves good performance. In addition, SVC-CSI is designed for channel state information (CSI) feedback to allocate the keypoint transmission and enhance the performance dramatically. Simulation shows that the proposed wireless semantic communication system can significantly improve the transmission efficiency.This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
I. INTRODUCTION
The paper develops semantic video conferencing that transmits facial keypoints to reduce resources, then adds error-aware HARQ and CSI-based allocation for varying wireless channels.
- Background: Semantic communications use deep learning to extract semantic information and encode shared and local knowledge in trained parameters.The approach has been applied to image, video, speech, and text transmission, with overhead reductions reported for specific tasks.
- Motivation: High-resolution video conferencing requires substantial transmission resources, particularly over mobile phones and bandwidth-limited channels.Conventional source-coding approaches do not account for varying wireless-channel effects.
- Motivation: Existing semantic video studies mainly address source coding, leaving the impact of physical-channel transmission errors on keypoints unclear.A practical system must protect key information and maintain acceptable received video under varying channels.
- SVC framework: The basal SVC framework represents facial motion with a small number of keypoints, reducing transmission resources while potentially losing detailed expressions.Unlike conventional methods, channel errors change keypoint locations and expressions rather than directly destroying pixels.
- SVC-HARQ: SVC-HARQ combines ACK feedback, incremental redundancy, and a semantic error detector that evaluates video fluency instead of relying only on CRC.The detector identifies inaccurate keypoints associated with reduced fluency and supports retransmission decisions across varying BERs.
- SVC-CSI: SVC-CSI allocates more important information to subchannels with higher SNRs, while testing-channel mismatch can worsen its performance.The framework is intended to balance performance, bit consumption, and robustness across channel environments.
II. SYSTEM MODEL AND RELATED WORKS
The section reviews semantic communication frameworks and their use of shared knowledge, then identifies limited adaptation to changing wireless channels as a key challenge.
- Semantic communication extracts source meaning before channel encoding, with semantic and channel processing performed separately or jointly.
- A. Semantic Frameworks: Implicit semantic knowledge is embedded in trainable parameters, enabling end-to-end learning of semantic features and physical-channel distortion handling.
- A. Semantic Frameworks: A limitation of implicit knowledge is that trained parameters are difficult to adjust under changing transmit sources or physical environments.
- A. Semantic Frameworks: Explicit semantic knowledge shares task-relevant information, such as important image features or a speaker photo, and can be adjusted when source information changes.
- A. Semantic Frameworks: Existing semantic methods focus on source-coding frameworks and do not adjust settings for different channel environments, limiting adaptation to physical-channel variation.
B. Link settings in the conventional methods
The section introduces conventional link-adaptation techniques, including IR-HARQ and OFDM-based channel adaptation, while highlighting challenges in applying them to semantic transmission.
- IR-HARQ: IR-HARQ splits coded symbols into initially transmitted high-code-rate symbols and incremental symbols used after an error indication.
- Error detection: The conventional CRC detector generates ACK = 1 when no error is found and ACK = 0 otherwise, but it does not represent semantic error severity.
- IR-HARQ: IR-HARQ combines received coded symbols and decodes again, repeating retransmission when errors remain to handle varying channels across time slots.
- OFDM adaptation: OFDM divides bandwidth into L parallel flat-fading subchannels with different SNRs, enabling modulation to adapt to subchannel gains when channel information is available.
- Challenges: Combining semantic networks with conventional link adaptation is challenging because semantic transmission introduces a novel transmission mechanism and must remain reliable across physical environments.
III. TRANSCEIVER DESIGN FOR SEMANTIC VIDEO CONFERENCING
The proposed SVC represents facial motion with transmitted keypoints while sharing a speaker image, then adds semantic error detection and wireless coding components.
- Motivation: The approach is intended to reduce transmission resources by encoding changing facial information rather than relatively stable appearance information.
- Basic SVC framework: SVC uses three levels—effectiveness, semantic, and technical—to deliver speaker motion and expression through semantic keypoint transmission.
- Basic SVC framework: The receiver shares a speaker photo in advance, combines it with received keypoints, and reconstructs each video frame through a generator.
- Basic SVC framework: The SVC architecture contains a keypoint detector and generator at the semantic level, plus an encoder-decoder at the technical level.
- Keypoint processing: The keypoint detector extracts n coordinates from each frame, selecting normalized grid points to represent facial movement.
- Technical-level transmission: The technical encoder processes 2n keypoint coordinates into m symbols and applies two-bit quantization to generate 2m transmitted bits.
- Training and reconstruction: The generator reconstructs frames from the shared image and received keypoints, while perceptual, discriminator, equivariance, and MSE losses train the system.
L (pi, G(p0, k0, KD(pi; WKD); WG)) . (14)
The training procedure first restores keypoints under channel distortion, then fine-tunes the complete SVC; the framework is designed to support semantic transmission with feedback.
- The technical level is trained with LMSE to restore keypoints under physical-channel distortion before complete-system fine-tuning.
- All SVC trainable parameters are subsequently fine-tuned end to end by minimizing the reconstruction loss.
- The basic SVC combines video synthesis with a three-level semantic wireless-transmission design and supports studying keypoint transmission instead of video transmission.
- ACK feedback is introduced as a further means of improving semantic-transmission performance.
B. Semantic HARQ with ACK Feedback for Video Conferencing
SVC-HARQ adds ACK-controlled incremental transmission and semantic error detection to adapt semantic video conferencing to changing channels. It uses quality or fluency assessment to decide whether further transmission is needed.
- Adaptive transmission: The framework makes SVC more adaptive under changing channels when the semantic detector is effective.Incremental transmission concentrates on fallible keypoints and uses separately trained parameters for the incremental process.
- HARQ framework: SVC-HARQ uses ACK feedback to trigger incremental transmission and retransmission when received semantic content fails the detector criterion.The first transmission is followed by incremental symbols, with retransmission if errors remain uncorrected.
- Semantic error detection: The semantic detector is central because it decides whether incremental transmission or retransmission is required.CRC is unsuitable because some subtle received-frame errors remain acceptable to conferees.
- Semantic error detection: A quality detector labels reconstructed frames as acceptable or unacceptable using a transfer-learned VGG-19-based classifier.The detector outputs a binary frame-quality indicator and is trained with cross-entropy.
- Semantic error detection: A fluency detector identifies sudden expression changes by measuring distances between keypoints in consecutive reconstructed frames.It labels frames using AKD, with threshold five, and is trained on reconstructed frames under different channels.
C. Adaptive Encoding with CSI Feedback
SVC-CSI uses channel-state feedback to reorder keypoint transmission toward stronger subchannels and explores quantized or full-resolution constellation modulation. Sorting channel conditions reduces feedback cost, while practical constellations constrain implementation complexity.
- CSI feedback: SVC-CSI adds a sorting module that arranges keypoint transmission according to decreasing subchannel gains.The sorted channel sequence is fed back rather than complete CSI values.
- Practical design: The sorted-feedback design reduces CSI feedback cost and simplifies the encoder-decoder design.SVC-CSI otherwise keeps the same training strategy as SVC.
- Constellation modulation: SVC-CSI supports full-resolution modulation with learned constellation points and quantized-resolution modulation with 16 possible locations.The quantized method combines two bits into each real symbol and introduces parameters α, β, and ρ.
- Constellation modulation: Full-resolution constellations are extremely complex for practical systems because of finite precision.The quantized-resolution alternative limits constellation points while retaining the SVC training strategy.
- CSI feedback: CSI feedback lets the network place important keypoints on high-SNR subchannels, protecting them at the technical encoder-decoder level.The method exploits differing subchannel conditions without requiring accurate CSI values for every subchannel.
IV. NUMERICAL RESULTS
The numerical evaluation compares semantic video conferencing with H264 and RS-based channel coding using transmission resources, image quality, perceptual, and keypoint-based metrics. Experiments use speaker video data and predefined IR-HARQ settings.
- Evaluation setup: The evaluation compares frameworks using required bits, perceptual loss, AKD, and SSIM.The study explicitly compares bit consumption with competing methods.
- Data and training: The training dataset contains 2000 videos and the testing dataset contains 100 videos from about 500 speakers.Videos are preprocessed to 256 × 256 and contain one speaker after filtering.
- Channel-coding setup: The IR-HARQ setup encodes 64 information symbols into 255, initially transmits 127, and reserves 128 for incremental redundancy.CRC detects errors after the first transmission.
- Metrics: AKD measures average distances between keypoints extracted from transmitted and received frames, representing facial motion and expression changes.A pretrained facial landmark detector supplies the keypoints.
- Metrics: SSIM evaluates structural similarity among image patches, while perceptual loss sums feature-level MSEs across layers of a pretrained network.SSIM is described as more robust than PSNR, and perceptual loss is used as a vision-network metric.
B. Performance of semantic coding
Semantic coding preserves facial content differently from H264 while using substantially fewer bits, and its robustness depends on the BER used during training. Under sufficiently high BER, SVC outperforms the conventional H264+RS comparison on the reported metrics.
- Content loss: H264 loses pixel information across the frame, whereas SVC primarily changes detailed facial expressions such as the mouth.The SVC information loss is usually not distinguishable as an independent image.
- BER robustness: When BER exceeds 0.027, SVC methods perform better than H264+RS on all three reported metrics.H264+RS remains unchanged while errors stay within its correction capability.
- BER robustness: SVC performance depends on training BER: the BER=0 model suits low BER, while the BER=0.05 model becomes better when BER exceeds 0.02.The training BER affects SVC similarly to selecting a channel-coding rate.
- Overall comparison: SVC saves resources because it transmits keypoints rather than compressed pixel information and can be advantageous under extremely high BER.The authors connect this resource advantage to high-resolution video conferencing.
C. Performance of SVC-HARQ
SVC-HARQ uses semantic error detection and incremental retransmission to preserve video quality across bit-error rates while adapting bit consumption. Its errors primarily affect facial expressions, and the framework uses fluency-related keypoint changes to trigger retransmission.
- Semantic and conventional errors: Semantic transmission errors mainly change facial expressions, whereas H264 errors can blur pixels and make the speaker unrecognizable.At BER=0.05, the semantic method has AKD about 9, while keypoints cannot be detected from the blurred H264+RS frame.
- Semantic error detection: The receiver detects keypoints again and uses detected AKD between adjacent frames as a measure related to video fluency.Detected AKD is generally below 0.05 without bit errors and increases with BER.
- Semantic error detection: The semantic detector evaluates received-frame fluency and can guarantee video quality without extra parity code.A VGG-based quality detector alone is insufficient because SVC errors are difficult to find independently and accepted ratios exceed 96%.
- SVC-HARQ operation: SVC-HARQ transmits 160 bits per frame and adds 160 incremental bits when the semantic detector finds errors.This contrasts with H264+RS-HARQ, which initially transmits 127 symbols per 64 information symbols and adds 128 symbols after CRC detection.
- SVC-HARQ performance: SVC-HARQ matches strong SVC performance across changing BERs while requiring 1/12 the bits per frame of H264+RS-HARQ.Its required bits per frame increase with BER, providing flexible bit consumption.
- Limitations: The semantic detector primarily protects video fluency, so an additional detector is needed to identify facial-expression errors.Detected-AKD error detection is not strict, producing nonsmooth AKD performance between BER 0.02 and 0.06.
D. Performance of SVC-CSI
SVC-CSI uses channel state information to place more keypoint information on stronger subchannels. It improves performance in matched channels but loses robustness when the testing channel differs from the training environment; combining CSI with HARQ mitigates this issue.
- Matched channels: Under matched channels, CSI feedback enhances performance at low BER and can outperform SVC when BER is below 0.14.SVC-CSI reaches the performance of SVC trained at BER=0 when BER=0 and exceeds SVC trained at BER=0.05 when BER<0.14.
- Mismatched channels: SVC-CSI performance decreases more sharply in mismatched channel environments and becomes worse than SVC trained at BER=0.05 when BER>0.04.The mismatch results from differences between testing and training channel environments, including statistical parameters such as delay spread.
- CSI-based allocation: SVC-CSI learns to allocate more information to subchannels with higher channel gains and lower noise power.The learned weights and symbol placement show that information is concentrated on better channel conditions.
- Constellation and power allocation: Full-resolution SVC-CSI adapts transmit power downward as channel conditions worsen, whereas quantized-resolution SVC-CSI increases power on worse subchannels.Quantized-resolution modulation uses the same modulation method across subchannels, with constellation points similar to 16-QAM.
- Constellation and power allocation: Quantized-resolution SVC-CSI uses a learned modulation that outperforms 16-QAM at low SNR but underperforms it when SNR≥8 dB.The learned modulation is therefore suitable for wired environments but cannot perfectly reconstruct frames at high SNR.
- CSI-HARQ: SVC-CSI-HARQ combines CSI and a CSI-free incremental transmission to improve robustness under mismatched channels.It is not worse than SVC-HARQ under mismatched conditions and performs better than SVC-HARQ when 0.02≤BER≤0.1.
V. CONCLUSIONS
The paper establishes semantic video conferencing by transmitting facial-expression keypoints rather than full video content, then adds semantic error detection, HARQ, and CSI feedback. The resulting framework reduces transmission resources, adapts bit consumption, and balances performance with robustness within its channel assumptions.
- SVC framework: SVC represents facial-expression motion using keypoint transmission, preserving resolution while losing detailed expressions.Conventional compression reduces resolution under limited bandwidth.
- Error model: SVC transmission errors change expressions, while conventional transmission errors directly destroy pixels.This distinction motivates semantic error detection for SVC-HARQ.
- SVC-HARQ: SVC-HARQ uses ACK feedback and semantic detection to adapt retransmission to received-frame errors.The framework can combine the performance of networks trained under different BERs and reach good performance flexibly.
- SVC-CSI: CSI feedback allocates more information to higher-gain subchannels and improves performance over SVC without CSI feedback.The approach sorts transmitted symbols or bits by subchannel SNR.
- Scope boundary: CSI feedback reduces robustness when the channel model used for testing differs from the training model.The paper identifies channel-model exploitation during training as the source of this robustness decrease.
- Overall conclusion: Combining CSI and ACK feedback balances performance, bit consumption, and robustness.This combination extends the framework beyond CSI-only allocation.