Source-linked AI summary

A Contemporary Survey on Semantic Communications:Theory of Mind, Generative AI, and Deep Joint Source-Channel Coding

Loc X. Nguyen, Avi Deb Raha, Pyae Sone Aung, Dusit Niyato, Zhu Han, Choong Seon Hong

arXiv:2502.16468v3cs.IT

TL;DR

Semantic communication lacks standardization across competing research directions, complicating consistent interpretation, objectives, and evaluation. This survey synthesizes Theory of Mind, Generative AI, and DJSCC approaches, their existing works, challenges, and opportunities, concluding that future systems may integrate multiple perspectives while addressing scalability and adaptability constraints.

  • Problem

    Semantic communication research lacks standardization across directions, producing inconsistencies in interpretation, objectives, and evaluation.

  • Method

    The survey provides an in-depth synthesis of Theory of Mind, Generative AI, and DJSCC approaches, including their theories, existing works, challenges, and future opportunities.

  • Results

    The survey concludes that semantic communication has a broad future scope, while its directions face challenges including storage, computation, scalability, and adaptability.

  • Takeaways & Limitations

    Future semantic communication may emerge from one direction or integrate multiple perspectives, with quantum technology identified as a potential research opportunity.

Abstract

from arXiv · show

Semantic communication is emerging as the next pillar in wireless communication technology due to its transformative capabilities in reducing communication overhead, enhancing robustness, and enabling intelligent information exchange. The most significant obstacle lies in the lack of standardization across various research directions, leading to inconsistencies in interpretation, objectives, and evaluation. In this survey, we provide an in-depth overview of three leading directions in semantic communication, namely Theory of Mind-based semantic communication, Generative AI-driven semantic communication, and Deep Joint Source-Channel Coding (DJSCC)-based semantic communication. These directions have been extensively studied and developed by research institutes worldwide, and their effectiveness continues to improve alongside advances in communication and computing technologies. The ToM-based semantic communication enables communication agents to interact intelligently, infer each other's intentions, and gradually form a shared understanding. The GAI-based semantic communication leverages generative models to create and interpret content beyond traditional compression, allowing flexible semantic encoding and decoding tailored to specific tasks. The DJSCC-based semantic communication direction integrates DL models to jointly optimize the source and channel coding processes for efficient semantic information transfer. Next, we present a detailed survey of existing works under each direction and open research problems in semantic communication. Furthermore, we identify and analyze critical challenges, such as scalability and adaptability, that currently hinder the deployment of semantic communication systems. Finally, we discuss potential research opportunities and future directions such as quantum computing to further enhance the capabilities of semantic communication.

I. INTRODUCTION

Semantic communication shifts the objective from reproducing transmitted bits to enabling intended interpretation, reducing redundant data and communication-resource use. The survey compares three machine-learning-based directions and their foundations, contributions, challenges, and future scope.

  • Motivation: Traditional communication reproduces transmitted messages, whereas semantic communication prioritizes whether receivers interpret messages as intended.Semantic systems can tolerate sequence differences when meaning is preserved.
  • Motivation: Semantic understanding identifies essential information and removes redundancy, reducing transmitted data while improving bandwidth and energy efficiency.The survey presents this capability as relevant to future 6G systems.
  • Survey scope: It explains how the directions eliminate redundant data through agent communication languages, generative models, and semantic-aware encoding.The three approaches respectively emphasize concise interaction, AI-based generation, and integrating semantics into encoding.
  • Challenges and outlook: The survey identifies knowledge-base storage, AI-model computation, and DJSCC network scalability as major challenges, alongside quantum technology as a future opportunity.It outlines future research tasks addressing these constraints.
  • Survey scope: The survey examines Theory of Mind, Generative AI, and DJSCC as three prominent semantic-communication directions built on machine-learning and deep-learning advances.These directions differ in their theoretical foundations and communication-system designs.
  • Theory of Mind: Theory of Mind-based communication uses interaction to model the receiver’s knowledge, behavior, and perspective during communication.The transmitter must consider the listener’s topic knowledge and may adapt through observed responses.

B. Generative AI-Based Semantic Communication

Generative AI-based semantic communication uses foundation models as communication agents to extract and generate task-relevant content with reduced transmission volume. It is related to ToM and DJSCC but differs in its use of pretrained knowledge and generative capabilities.

  • Generative AI approach: Generative AI-based semantic communication uses foundation models trained on vast datasets across tasks and application domains.These models serve as communication agents rather than task-specific models trained from scratch.
  • Generative AI approach: Foundation models can convert raw content into task-relevant descriptions, reducing transmitted data compared with transmitting complete source representations.The approach can increase user capacity and shorten communication times.
  • Relation to other directions: DJSCC jointly optimizes source and channel coding, embedding semantic information into transmission while improving robustness to channel noise.Its common autoencoder design uses learned encoders and decoders for reconstruction or downstream tasks.
  • Future direction: The field has not established a definitive direction, and future standardization may arise from one approach or from integrating multiple perspectives.The survey notes that experimentation, debate, and comparative performance will shape this development.
  • Relation to other directions: The survey presents the three directions as interrelated rather than mutually exclusive, with collaboration and minimal transmission cost as a shared principle.Their main differences concern transmitter and receiver design.
  • Relation to other directions: Generative AI equips transmitter and receiver with advanced pretrained models that can be fine-tuned for specific tasks and minimized transmitted signals.The survey characterizes this as an expanded form of ToM because the agents use broader accumulated knowledge.

A. Explanation for Theory of Mind

Theory of Mind-based semantic communication models interactions among agents, shared knowledge, reasoning, and channel conditions to improve semantic exchange. Research explores shared-language formation, richer knowledge representations, feedback, and multi-agent coordination.

  • Theory of Mind-based semantic communication lets agents accumulate shared knowledge, model one another, and form a common language for more efficient exchange.
  • Reasoning-based communication: Contextual reasoning allows agents to model listeners, remove redundant information, and improve semantic communication efficiency.
  • Knowledge representation: Graph-based knowledge structures represent entities, relations, and reasoning rules, enabling more accurate interpretation than predefined symbols.
  • Feedback and adaptation: A two-level feedback mechanism combines channel-quality and semantic feedback, while ToM and GNNs extract causal relations and adapt transmitter parameters.
  • Multi-agent communication: Multi-agent communication introduces language mismatch and semantic noise, requiring mechanisms such as semantic channel equalization and resource-aware partial transmission.

C. Challenges of the Direction

The Theory of Mind direction faces practical challenges from complex belief structures, user-dependent reasoning, and limited deployment evidence. Its broader development also depends on suitable knowledge bases, reasoning capabilities, and increasingly capable pretrained models.

  • Scalability: Hierarchical causal and reasoning sequences become difficult to scale, with curriculum learning already needed for systems using only three belief levels.
  • Adaptability: Human beliefs are more complex than the assumed hierarchy, while causal reasoning may suit one user but be wrong for another.
  • Deployment: Theory of Mind-based systems need practical scenarios and demonstrations of effectiveness to attract broader development.
  • Knowledge bases: Shared knowledge bases are essential because accumulated transmitter-receiver knowledge can significantly improve system performance.
  • Reasoning: Contextual and causal reasoning can reduce transmitted data while helping receivers construct semantically meaningful sentences.
  • Model development: Pretrained large language models are used as communication agents because training systems from scratch has limited capacity.

B. The existing research of Generative AI-based Semantic Communication

Generative AI-based semantic communication uses generative models and knowledge bases to encode, generate, and exchange task-relevant information across communication and application scenarios. Existing work spans resource management, computation offloading, latency-quality trade-offs, security, and metaverse or industrial applications.

  • Application scope: Generative AI supports semantic communication by controlling encoding and generation across resource allocation, network management, and data encoding applications.
  • System integration: Semantic communication and generative AI complement each other by handling immediate context and large multimodal data under bandwidth and computing constraints.
  • Resource sharing: In mixed reality, users transmit generated content and extracted semantic information to nearby users, reducing repetitive local computation.
  • Semantic resource allocation: YOLO-based semantic communication discards irrelevant image information and allocates more power to important regions for object detection and digital-twin data exchange.
  • Latency and quality: Generative AI models can serve as workload-adjustable transceivers by balancing transmission latency, computing latency, image quality, channel quality, and service requirements.
  • Metaverse and industrial applications: Metaverse and industrial-factory systems apply generative AI-based semantic communication to high-quality transmission, security, and knowledge-based rendering.

3) Collaborations from Edge and Server:

Edge and server collaboration addresses limited device resources in generative semantic communication, while reconstruction-focused methods improve fidelity beyond text-only transmission. These approaches span visual, video, audio, and text modalities.

  • Edge and Server Collaboration: Edge-device and cloud collaboration supports generative semantic communication when local devices cannot complete inference, update models, or reach target devices.Fine-tuning on local devices and edge servers can reduce privacy exposure during data collection.
  • Edge and Server Collaboration: Incentive-based knowledge markets encourage IoT devices to share knowledge graphs through utility-driven incentives and fair pricing.
  • Image Reconstruction: Text-only transmission can preserve task-oriented meaning and reduce signal size, but generative reconstruction may produce an image different from the original.This limitation motivates methods that improve reconstruction ability.
  • Image Reconstruction: Visual prompts, structural features, and diffusion-derived visual information are used to improve reconstruction while controlling communication cost.The approaches include contour-like prompts, edge maps, colors, textures, and other visual features.
  • Video, Audio, and Text: Generative semantic communication extends beyond images to video and audio, using text transcripts or captions to support reconstruction under modality-specific conditions.Video work reconstructs talking-head video with audio from transmitted text, while audio work addresses noise and missing speech segments.
  • Video, Audio, and Text: A semantic-importance-aware priority strategy allocates more power to important frames, with the proposed semantic loss and BERT outperforming ChatGPT and random outputs in simulations.

6) Generative AI-based Semantic Communication in Vehicular Network:

Generative AI-based semantic communication is applied to vehicular networks to address reliability, latency, bandwidth, and safety demands. Its broader capabilities include controllable generation, network management, interpretation, denoising, and reasoning, alongside deployment and security concerns.

  • Vehicular Applications: Vehicular semantic communication targets exceptionally high reliability and ultra-low latency, motivating generative-AI approaches for connected transportation.
  • Vehicular Applications: Multimodal encoding reduces transmission length and communication latency in vehicular networks, supporting traffic-safety applications such as accident information transmission.
  • Vehicular Applications: For roadside imagery, MSAM extracts meaningful semantic information and GAN-based processing addresses bandwidth consumption and latency in intelligent transportation systems.
  • Challenges: Deployment remains constrained by generative-model computation and device energy limits, while interoperability gaps and vulnerability to maliciously fabricated content create additional concerns.
  • Capabilities: Generative AI supports network management, adaptive strategies, high-quality content creation, cross-modal interpretation, denoising, and reasoning in semantic communication.

V. DIRECTION III: DEEP JOINT SOURCE-CHANNEL CODING BASED SEMANTIC COMMUNICATION

D-JSCC-based semantic communication combines semantic transmission with joint source-channel coding to overcome limitations of separate coding under practical channel conditions. The section motivates neural-network-based joint optimization through the cliff effect, resource constraints, and changing multimedia requirements.

  • Motivation: Semantic communication transmits essential meaning rather than exact bit-level representations, allowing receivers to interpret messages without requiring identical transmitted and received bit sequences.
  • Limitations of Separate Coding: Separate source and channel coding can suffer a cliff effect because small increases in channel noise may suddenly collapse communication quality in practical environments.
  • Joint Source-Channel Coding: JSCC jointly optimizes source compression and channel encoding, accounting for source and channel characteristics to improve performance under real-world conditions.
  • Joint Source-Channel Coding: JSCC can adapt to application, channel, and QoS requirements while maintaining communication quality under challenging conditions through enhanced error resilience.
  • Transition to D-JSCC: Traditional JSCC struggles with high-dimensional multimedia data and rapidly changing heterogeneous channels, motivating deep-learning-based JSCC systems.

C. Deep Joint Source-Channel Coding

D-JSCC integrates compression and channel encoding in end-to-end neural systems for robust semantic transmission across multimedia and wireless settings. Research addresses graceful degradation, adaptive rates, realistic channels, hardware constraints, semantic noise, and multi-user operation.

  • Framework and Applications: D-JSCC jointly compresses and encodes source data while capturing essential semantic information across image, audio, speech, video, text, and multimodal transmission.
  • Image Transmission: For wireless image transmission, D-JSCC provides graceful quality degradation as channel conditions worsen and can outperform conventional digital schemes under severe impairments.
  • Practical Wireless Channels: Practical D-JSCC research models channel effects including noise, fading, interference, multipath propagation, and unknown channel state information.
  • Adaptive Transmission: Adaptive rate control improves bandwidth efficiency, particularly at high SNR or for less informative images, without substantial quality loss relative to fixed-rate models.
  • Hardware Constraints: DeepJSCC-Q uses quantization to map encoder outputs onto finite constellation points, with fixed or learned constellations and KL regularization during training.
  • Semantic and Multi-User Communication: Semantic communication systems reduce task-relevant bit rates compared with JPEG and JPEG2000, while multi-user methods such as NOMASC target spectral efficiency and transmission rates.
  • Semantic and Multi-User Communication: Conditional encoders and trustworthy decoders address user-specific semantic extraction and catastrophic forgetting in multi-user systems.

2) Text Communication:

Text semantic communication shifts transmission from bit-level correctness toward preserving meaning, using learned semantic representations and channel-aware designs to improve efficiency and reliability. The surveyed works address variable-length coding, resource allocation, fading channels, and deployment constraints, while fixed codeword lengths remain a limitation.

  • Text Communication: DeepSC uses a Transformer to extract and encode meaningful text information, jointly optimizing semantic and channel coding for noisy transmission.The system targets semantic-level accuracy rather than bit-level accuracy and seeks to reduce semantic errors.
  • Text Communication: CSI-aided training, refined least-squares estimation, pruning, quantization, and finite-bit constellations address fading, pilot overhead, model complexity, and hardware costs.These techniques target robust transmission and deployment on resource-constrained devices.
  • Text Communication: Semantic spectral efficiency measures text communication efficiency using channel assignment and transmitted semantic symbols for resource allocation.The associated optimization maximizes overall semantic spectral efficiency.
  • Text Communication: Existing text communication works largely assume fixed codeword lengths, limiting flexibility across sentence lengths and channel conditions.This limitation motivates variable-length approaches such as SC-RS-HARQ.
  • Text Communication: SC-RS-HARQ combines semantic coding with Reed-Solomon coding and HARQ, adapting code lengths to sentence length and channel conditions.This variable-length design addresses the inefficiency of fixed-length coding for varying sentences.

3) Audio and Speech Transmission:

Audio and speech semantic communication combines semantic extraction, joint source-channel coding, generative restoration, and environmental adaptation to transmit or reconstruct meaningful speech efficiently. The surveyed systems target noise, missing data, perceptual quality, computational complexity, and changing channel conditions.

  • Audio and Speech Transmission: DeepSC-S integrates semantic and channel coding to address speech transmission’s emphasis on bit accuracy and poor performance in dynamic environments.It is part of a body of D-JSCC work reported to outperform traditional systems under low SNR.
  • Audio and Speech Transmission: SemAudio uses Transformer-XL to extract essential information from streaming audio while jointly designing streaming semantic and channel coding for reliable recovery.Its design addresses semantic extraction, recovery, and channel distortion or attenuation in real-time audio transmission.
  • Audio and Speech Transmission: A perceptually motivated low-complexity system reduces transmitted speech data by 60% while improving transmission efficiency and received-speech quality.A multiresolution joint loss aligns reconstruction with human auditory perception.
  • Audio and Speech Transmission: DSST combines nonlinear semantic transforms, entropy-guided rate allocation, and SNR adaptation to save bandwidth while maintaining high audio quality.Its SNR adaptation enables one model to operate across diverse channel conditions.
  • Audio and Speech Transmission: A generative audio semantic communication framework treats degraded transmission as an inverse problem and uses conditional diffusion to reconstruct noisy or missing audio.Automatic denoising-parameter computation adapts reconstruction to channel conditions without prior channel knowledge.

4) Video Streaming:

Video semantic communication transmits task-relevant or compact semantic representations instead of complete video data, using adaptive coding, generative reconstruction, and feature selection. The surveyed methods improve efficiency and robustness but face computational, modality, and channel-condition constraints.

  • Video Streaming: Deep JSCC video systems use adaptive entropy models and variable-length coding to control bandwidth allocation across frames or frame regions.These methods exploit temporal information and adapt feature transmission to available resources.
  • Video Streaming: A semantic encoder-decoder combines bi-optical flow extraction, noise attention, feature choice, and feature fusion to preserve important inter-frame information under channel noise.Noise attention incorporates SNR into encoding and decoding, while feature choice removes low-semantic-information features.
  • Video Streaming: Transmitting shared and individual feature maps for frame groups reduces representation dimensionality and improves communication efficiency.The approach extracts common information across a group of frames rather than treating every frame independently.
  • Video Streaming: The OAR mechanism transmits the first frame of each GoP plus object-attribute-relation graphs, achieving a CBR value of 1/300 and outperforming H.265 at lower bit-rates.A generative receiver uses reference-frame appearance and object motion information to synthesize new frames.
  • Video Streaming: Multimodal video systems improve packet-loss recovery and audio-visual alignment but can incur high computational complexity and limited modality or dataset scope.Reported boundaries include reduced real-time performance, unclear adaptation to additional modalities, and evaluation restricted to the AVE subset of Audioset.

5) Multi-modal Data Transmission:

Multimodal semantic communication exploits relationships among modalities for task-oriented transmission, reconstruction, retrieval, translation, and question answering. The surveyed works use cross-modal encoding, fusion, alignment, and conventional channel coding, while highlighting compatibility and generalization challenges.

  • Multi-modal Data Transmission: Multimodal semantic communication supports both task-oriented and data-reconstruction approaches for transmitting multiple modalities across users.The surveyed literature responds to growing demand for multimodal transmission.
  • Multi-modal Data Transmission: Early multimodal systems separately decode image and text before using memory, attention, and composition networks to capture cross-modal correlations and answer questions.The setup pairs an image transmitter with a text transmitter whose questions concern the transmitted image.
  • Multi-modal Data Transmission: Later systems extend multimodal communication to image retrieval and machine translation using Transformer semantic encoders, decoders, and information fusion modules.Fusion links keywords with corresponding image regions for visual question answering.
  • Multi-modal Data Transmission: Cross-modal encoders, knowledge graphs, alignment losses, and amendment networks model shared semantics and dynamically adjust auxiliary information across modalities.These designs address interpretation and coordination between modalities during encoding and decoding.
  • Multi-modal Data Transmission: A hybrid framework combines multimodal semantic coding with conventional channel coding to address incompatibility with modern digital communication and reduce redesign when tasks change.The surveyed problem includes both end-to-end analog compatibility and the need to retrain or redesign systems for different tasks.

D. Challenges within the Directions

Deep JSCC-based semantic communication has achieved strong performance, but heterogeneous transmitter and receiver models make optimization difficult in realistic multi-user settings.

  • Deep JSCC-based semantic communication outperforms traditional wireless communication in performance metrics and compression ratios.
  • Most existing work assumes one transmitter and one receiver, while multi-transmitter and multi-receiver settings remain limited.
  • Assuming identical models and computational capacity across devices is unrealistic because practical transmitters and receivers may use different semantic-communication models.
  • Heterogeneous deep-learning models complicate parameter optimization and can cause catastrophic forgetting that degrades performance on earlier tasks.

E. Summary and Lesson Learned

Deep JSCC improves communication efficiency and robustness by adapting encoding to wireless conditions and data modalities, while related works extend semantic communication to practical networking scenarios.

  • Summary and Lesson Learned: Deep JSCC research has progressed from low-energy designs toward high-performance architectures for next-generation communication requirements.
  • Summary and Lesson Learned: Integrating wireless conditions and data semantics into encoding avoids the cliff effect at low SNR and improves performance as channel conditions improve.
  • Summary and Lesson Learned: Modality-specific encoding reduces transmitted data while achieving better performance than conventional wireless communication systems.
  • Other Promising Semantic Communication Works: Semantic features extracted from traffic signs can compress data for connected autonomous vehicular networks, reducing bandwidth use and supporting faster decisions.
  • Other Promising Semantic Communication Works: Related systems apply semantic communication to RIS-assisted links, integrated sensing, robust mmWave beamforming, and task-dependent radio-resource slicing.
  • Summary and Lesson Learned: Each semantic communication direction still faces unresolved challenges and requires further development.

1) Theory of Mind-based Semantic Communication:

The survey presents semantic communication as a fragmented field organized around three theories and highlights knowledge, scalability, generative-AI, DJSCC, and quantum-computing challenges that remain open.

  • Theory of Mind-based Semantic Communication: Communication agents need knowledge bases that update beliefs and facts, remove outdated information, and use low-dimensional representations to support reasoning.
  • Theory of Mind-based Semantic Communication: Scaling communication networks requires addressing knowledge-base management, reasoning ability, and increasingly multimodal data interpretation beyond the current textual focus.
  • Generative AI-based Semantic Communication: Generative-AI semantic communication remains early-stage because devices need energy-efficient models despite limited resources and the high energy and hardware demands of generative AI.
  • Generative AI-based Semantic Communication: Reliable deployment also requires control and explainability of generative-AI outputs for trustworthy, ethical, and standardized communication systems.
  • Deep JSCC-based Semantic Communication: DJSCC has achieved strong performance, but heterogeneous deep-learning models and support for many tasks remain important development challenges.
  • Quantum-based Semantic Communication: Quantum semantic communication has been proposed using modules for embedding, semantic encoding, anonymous teleportation, decoding, and detection, but practical implementation remains infeasible at this stage.
  • Conclusion and Discussion: The survey organizes semantic communication around Theory of Mind, AI-generated content, and joint source-channel coding because no standardized protocol or unified framework has been established.
  • Conclusion and Discussion: The paper aims to provide an overall view of semantic communication and identify what remains to be improved, while the eventual relationship among the three directions remains uncertain.
Loading 2502.16468v3…