Source-linked AI summary
Ada-TokenCom: Rate-Adaptive Token Communications via Large-Model-Driven Token Compression and Generation
Zijun Zhang, Li Qiao, Mahdi Boloursaz Mashhadi, Zhen Gao, Mehdi Bennis, Kaibin Huang
TL;DR
Ada-TokenCom addresses limited digital token-level adaptation by combining AR compression, receiver-side generation, and dynamic wireless control. It transmits informative prefixes with arithmetic coding, generates tails at the receiver, and reports stronger semantic and perceptual results than baselines, while adding AR prediction latency.
Problem
Prior adaptive semantic communication methods mainly adjust continuous learned representations or preconfigured models, while joint digital token-level adaptation remains less explored.
Method
Ada-TokenCom arithmetic-codes AR-compressed head tokens, predicts untransmitted tail tokens with the same receiver-side AR model, and jointly selects prefix length and MCS.
Results
Ada-TokenCom outperforms digital and deep joint source-channel coding baselines, achieving higher CLIP similarity and up to about 60% lower LPIPS than SwinJSCC at comparable CBRs.
Takeaways & Limitations
The framework supports bandwidth-limited communication by trading receiver-side computation for reduced communication overhead.
Takeaways & Limitations
Additional latency from autoregressive token prediction remains a current limitation.
Abstract
from arXiv · showhide
Token Communications (TokenCom) has recently emerged as a new paradigm in which tokens serve as unified units for communication and computation, enabling efficient multimodal semantic and goal-oriented transmission. In this paper, we develop Ada-TokenCom, a rate-adaptive TokenCom framework based on large autoregressive models, which integrates next-token prediction with arithmetic coding to achieve ultra-low bitrate semantic communication at the token level. We propose a mixed reconstruction/generation scheme, where the transmitter encodes and transmits the highly informative tokens at the beginning of the token sequence leveraging a pre-trained autoregressive large model, while the receiver uses an identical model to predict the rest. Moreover, we design a Lyapunov-based algorithm to dynamically optimize both the source compression rate and the modulation and coding scheme, adapting to time-varying network conditions. Simulation results demonstrate that our proposed Ada-TokenCom framework outperforms both digital and deep joint source-channel coding-based semantic communication baselines.
I. INTRODUCTION
Token Communications uses structured token sequences to connect semantic communication with computation, while Ada-TokenCom targets adaptive digital transmission under changing wireless conditions. It builds on generative and token-level methods but focuses on jointly coordinating compression, receiver prediction, and link adaptation.
- Existing learned joint source-channel coding methods often specialize to particular modalities, tasks, datasets, or channel conditions, limiting generalization and digital-system integration.
- TokenCom treats tokens as basic units for semantic-level coding, transmission, inference, and generation across heterogeneous modalities.Tokenized representations may include discrete indices, latent visual codes, audio units, or continuous embeddings.
- Generative semantic communication uses diffusion models, foundation models, and multimodal large models to reconstruct missing content from compact semantic descriptions, prompts, tokens, or partial representations.
- Adaptive semantic transmission has mainly adjusted continuous learned representations or preconfigured semantic models rather than discrete token sequences.
- Ada-TokenCom coordinates AR token compression, receiver-side prediction, and wireless link adaptation under time-varying channels and average symbol-overhead constraints.
C. Contributions
Ada-TokenCom addresses the lack of a principled adaptive digital TokenCom system by combining AR token modeling, arithmetic coding, receiver-side generation, and Lyapunov-based cross-layer control. Its design uses token context and side information to support rate-quality adaptation.
- Existing TokenCom and generative semantic communication studies leave joint adaptation of AR compression, receiver prediction, MCS selection, and long-term overhead under time-varying channels insufficiently addressed.
- Ada-TokenCom partitions each sequence into arithmetic-coded head tokens and AR-predicted tail tokens, replacing part of explicit transmission with receiver-side computation.
- Decoder-only Transformer AR probability modeling is integrated with arithmetic coding for bandwidth-efficient digital token compression and transmission.
- The mixed reconstruction-and-generation mechanism provides an adaptive rate-quality trade-off by exploiting receiver-side generative priors.
- A Lyapunov-based algorithm jointly selects token truncation length and MCS under time-varying channels and long-term symbol-overhead constraints.
- Token sequences are formed by pretrained tokenizers, with each token index identifying a codebook codeword and side information supplying compact semantic context.
- Conditional self-information represents the minimum bits needed to encode a token, and its observed decrease across positions motivates transmitting informative prefixes.
B. Arithmetic Coding with Large AR Model-Based Priors
Ada-TokenCom uses AR-model conditional probabilities to drive arithmetic coding, recursively encoding a transmitted token prefix and reproducing the same distributions at the decoder for lossless prefix recovery.
- AR modeling and arithmetic coding operate sequentially on prefix-conditioned distributions, with more accurate predictions yielding shorter expected code lengths.
- The pretrained AR model produces predictive distributions over the vocabulary for each prefix position, beginning from an empty prefix.
- The encoder collects the position-wise distributions and recursively selects the subinterval corresponding to each actual token within the current arithmetic-coding interval.
- Encoding the retained prefix maps it to a final interval and binary fractional bitstream carrying the predictive conditional probability sequence.
- The decoder reproduces identical conditional distributions from side information and the decoded prefix, recovering transmitted tokens losslessly.
C. Mixed Reconstruction/Generation for Compression Rate Adaptation
Ada-TokenCom adapts bitrate by transmitting only informative head tokens and generating the remaining tail tokens with the receiver’s AR model. Prefix length controls the trade-off between communication rate and reconstruction fidelity.
- The transmitter arithmetic-codes informative head tokens while leaving tail tokens untransmitted for receiver-side AR prediction.
1) Reconstruction with AC-based Head Token Decoding:
The receiver decodes an informative transmitted token prefix using an identical autoregressive model and arithmetic coding, then generates the remaining tail tokens conditionally before detokenization.
- Head Token Decoding: The receiver reproduces the transmitter’s conditional distributions and sequentially decodes the transmitted prefix from the arithmetic-coded bitstream.It initializes the same coding interval and recovers tokens one by one.
- Head Token Decoding: Arithmetic decoding selects each token whose predictive subinterval contains the received code value, then recursively updates the current interval.
- Tail Token Prediction: After recovering the first n∗ tokens, the receiver autoregressively predicts the remaining N − n∗ tail tokens using the decoded prefix and shared side information.The large AR model supplies the token prediction function for this completion stage.
- Mixed Reconstruction: The first n∗ head tokens are recovered exactly through arithmetic coding, while the tail is completed through conditional autoregressive prediction.
- Rate Adaptation: Adjusting n∗ adapts the compression rate: shorter prefixes generate more tail tokens, whereas longer prefixes enable more faithful reconstruction.This combines exact transmitted-token recovery with generative completion while preserving semantic consistency.
III. CROSS-LAYER RATE ADAPTATION FOR WIRELESS TOKENCOM
Ada-TokenCom couples the retained token-prefix length with channel protection to adapt source compression and wireless transmission jointly under changing link conditions.
- Cross-Layer Design: The retained prefix length n∗ determines the source rate, while the selected MCS determines the channel protection level.
- Cross-Layer Design: Transmitting only the first n∗ tokens changes packet length, packet error rate, retransmission risk, and achievable semantic fidelity under limited channel resources.
- Online Configuration: Each slot jointly selects a prefix length n∗ and MCS m from candidate configurations, with packet length and error behavior evaluated for each choice.The packet length includes channel coding and protocol overhead such as CRC bits.
- PER Approximation: The framework approximates average packet-error behavior on quasi-static Rayleigh fading using waterfall thresholds characterized by MCS- and packet-length-dependent fitting coefficients.The coefficients are obtained offline by curve fitting.
C. Lyapunov-Based Joint Prefix-Length and MCS Adaptation
The adapter jointly chooses prefix length and MCS online, using Lyapunov optimization to balance semantic quality against a long-term channel-symbol budget as wireless conditions vary.
- Joint Adaptation: Prefix length n∗ controls both communication resource consumption and the number of tail tokens generated by conditional AR prediction.
- Design Trade-off: Reducing n∗ primarily lowers communication resource consumption without reducing receiver-side inference complexity, because the receiver still processes the full token sequence.
- Joint Adaptation: Longer prefixes can improve semantic fidelity in favorable channels, whereas shorter prefixes with stronger protection may suit poor channel states.
- Resource Accounting: Type-I ARQ retransmits failed packets in full, making expected transmission attempts 1/(1 − Pt(dt, γ̄t)) and increasing expected channel-symbol consumption.
- Optimization Objective: The long-term objective maximizes time-average semantic quality, using CLIP similarity, subject to an average channel-symbol budget Cth.PSNR and LPIPS are also identified as possible semantic-quality measures.
- Lyapunov Control: A virtual queue Zt tracks accumulated budget violation, so aggressive transmission decisions increase backlog and incur a larger penalty.
- Online Algorithm: In each slot, Algorithm 1 evaluates feasible prefix-length–MCS pairs, computes their transmission quantities and objective, selects the minimizer, and updates the virtual queue.
- Lyapunov Control: The drift-plus-penalty objective balances semantic quality and long-term channel-symbol consumption through adaptation parameter η.
D. Complexity Analysis of Online Adaptation
The online adapter has linear complexity in the number of feasible prefix-length–MCS configurations, while virtual-queue maintenance is constant-time.
- Complexity: The per-slot online adaptation complexity is O(|D|) = O(|N||M|), and the virtual queue update has complexity O(1).Candidate evaluations use constant-time operations, with prefix-length-to-bit-length mapping implemented by lookup.
IV. SIMULATION RESULTS
The simulations analyze token self-information and predictive entropy under different autoregressive models, showing that token uncertainty varies by position and becomes more predictable with sufficient context. These findings support transmitting informative prefix tokens while generating the remaining tail.
- Self-Information Analysis: Conditional self-information generally decreases along the generation order, indicating that earlier tokens contribute more information than later tokens.The profile is estimated from raster-scan autoregressive probabilities over discrete latent tokens.
- Spatial Token Structure: Self-information fluctuations predominantly occur at latent-grid discontinuities created by flattening the 2D VQGAN grid into a 1D sequence.The visualization uses a 16 × 16 VQGAN latent grid and LlamaGen-L.
- Entropy Analysis: Decoded examples are evaluated under specified head-token lengths to visualize how truncation changes reconstruction behavior.The simulations use VQGAN tokenization and LlamaGen architectures pretrained for class-conditioned image generation on ImageNet-related data.
- Self-Information Analysis: 11.42 bits/token for LlamaGen-B, 11.22 bits/token for LlamaGen-L, and 8.09 bits/token for Taming Transformers with class labels, compared with 14 bits/token for uniform binary token indices.The results indicate lower average self-information for the evaluated pretrained models when conditional class information is available.
- Entropy Analysis: Conditional entropy is lower in the center of the latent grid, where image objects are typically located, after class labels and preceding raster-scan tokens provide context.This indicates that later tokens can be predicted more reliably after sufficient prefix information has been observed.
C. Semantic Compression Performance
The semantic compression experiments compare Ada-TokenCom with learned, conventional, and token-based image compression methods across conditional and unconditional settings. Ada-TokenCom achieves very low bitrates while preserving semantic similarity and perceptual plausibility, with quality improving as more prefix tokens are transmitted.
- Rate and Quality Comparison: 0.0062–0.0468 bpp for Ada-TokenCom-cond, over 90% lower than Cheng2020-Attention at approximately 0.145–0.711 bpp.Ada-TokenCom-cond also uses substantially less rate than HiFiC at 0.688–1.003 bpp.
- Rate and Quality Comparison: 0.923–0.973 CLIP similarity is maintained by Ada-TokenCom-cond across its low-rate range.The framework also achieves favorable FID in the low-rate region, indicating semantic alignment and distribution-level realism.
- Rate and Quality Comparison: Ada-TokenCom reduces bitrate below vanilla TokenCom’s fixed 0.0547 bpp through autoregressive arithmetic coding and supports flexible rate adaptation via token truncation.Conventional codecs achieve higher PSNR, while Ada-TokenCom reports LPIPS of 0.140–0.595 at much lower rates.
- Reconstruction Under Truncation: Increasing the retained prefix length progressively improves visual quality and semantic consistency in both class-to-image and text-to-image reconstruction examples.The experiments visualize reconstruction at different truncation lengths for both conditioning scenarios.
- Reconstruction Under Truncation: Reducing the retained prefix length lowers both bpp and transmitted bytes while preserving the main semantics specified by the condition in the text-to-image setting.The text prompt contributes 252 bits of overhead, described as negligible relative to the image-token payload.
- Wireless Evaluation Inputs: CLIP similarity ranges from A = 0.843 at n∗ = 32 to A = 0.937 at n∗ = 256 for the evaluated class-conditioned ImageNet-V2 setup.The candidate prefix lengths are aligned with flattened VQGAN latent-grid boundaries.
1) End-to-End Comparison with Baseline Schemes:
Ada-TokenCom adapts token prefix length and MCS to channel conditions and long-term budgets, improving semantic and perceptual quality over SwinJSCC at low rates while sacrificing PSNR.
- End-to-End Reconstruction Comparison: 21%–28% CLIP similarity gains around CBR ≈10−2, reaching approximately 47% at lower CBR, outperform SwinJSCC across SNRs.Ada-TokenCom performs particularly strongly in the low-CBR regime.
- End-to-End Reconstruction Comparison: Up to about 60% lower LPIPS and about 77% lower FID demonstrate stronger perceptual and generative fidelity than SwinJSCC at comparable or low-to-medium CBR.SwinJSCC nevertheless achieves higher PSNR when pixel-level fidelity is the primary objective.
- Lyapunov-Based Adaptation: Ada-TokenCom selects prefix length and MCS jointly under a long-term symbol budget, adapting online without non-causal future-SNR knowledge.The Lyapunov drift-plus-penalty policy provides the semantic-rate adapter for this cross-layer decision.
- Lyapunov-Based Adaptation: η balances semantic quality against channel usage: small values under-utilize budget, whereas excessively large values can violate it.η ≈106 is selected as an aggressive yet near-feasible operating point across tested SNR regimes and budget levels.
- Lyapunov-Based Adaptation: Under a representative dynamic SNR process, the policy approaches the long-term oracle’s final-slot throughput and attains higher cumulative CLIP and perceptual reward under limited budgets.The perceptual reward is measured as (1 −LPIPS), with larger values indicating better perceptual quality.
E. Computational Complexity and Latency Analysis
Ada-TokenCom reduces transmitted bits through token truncation and semantic tail generation, but retains full-horizon sequential receiver inference, creating a latency–communication trade-off.
- Receiver Complexity: Token truncation reduces transmitted bits, but the receiver still performs N sequential AR evaluations for a sequence of length N.The first n∗ steps reconstruct the arithmetic-coded prefix, while the remaining N −n∗ steps generate the discarded tail.
- Receiver Complexity: Ada-TokenCom has the same order of decoding complexity as full token transmission with arithmetic coding, despite enabling semantic completion of discarded tokens.The method therefore exchanges communication savings for receiver-side computation.
- Encoder Efficiency: Full-token prefill removes the sequential transmitter loop for retained tokens and substantially reduces encoding latency.The implementation also uses torch.compile as a runtime optimization without changing probability distributions or the transmitted bitstream.
- Latency Trade-Off: The AR framework incurs higher decoding latency than lighter autoencoder baselines because it performs full-horizon sequential inference.This added computation is exchanged for improved compression efficiency and semantic reconstruction quality in the low-rate regime.
- Latency Trade-Off: 908.1 ms AR latency and 95.3 ms arithmetic-coding latency dominate decoder runtime for Ada-TokenCom.For a 256-token sequence, repeated prefill evaluations enforce exact encoder–decoder probability consistency.
- Scope and Limitations: The framework is suitable when receiver-side computation can be traded for reduced communication overhead, but additional AR prediction latency remains a limitation.Future work targets faster generation mechanisms and joint communication–generation optimization.