Source-linked AI summary
Personalized Saliency in Task-Oriented Semantic Communications: Image Transmission and Performance Analysis
Jiawen Kang, Hongyang Du, Zonghang Li, Zehui Xiong, Shiyao Ma, Dusit Niyato, Yuan Li
TL;DR
The paper addresses inefficient UAV image retrieval, missing personalization in semantic encoding, and limited mathematical understanding of fading effects in task-oriented semantic communication. It proposes a triple-based scene-graph framework, personalized attention-weighted encoding, fading analysis, and game-based multi-user resource allocation. Numerical evaluations report improved personalization, interference performance, and communication quality, including a 64% communication-cost reduction.
Problem
UAV semantic communication lacks efficient image retrieval, personalized encoding for user interests, and sufficient analysis of wireless fading effects on semantic triplet delivery.
Method
The paper combines triple-based scene-graph retrieval, personalized attention-based triplet weighting, mathematical fading-channel analysis, and game-theory-based multi-user resource allocation.
Results
The proposed framework and schemes improve personalization and anti-interference performance, while PERSF-SEMCOM reduces communication cost by 64% in the reported image-transmission example.
Takeaways & Limitations
Personalized saliency and resource-aware semantic transmission support more efficient UAV image-sensing services for users with different interests.
Abstract
from arXiv · showhide
Semantic communication, as a promising technology, has emerged to break through the Shannon limit, which is envisioned as the key enabler and fundamental paradigm for future 6G networks and applications, e.g., smart healthcare. In this paper, we focus on UAV image-sensing-driven task-oriented semantic communications scenarios. The majority of existing work has focused on designing advanced algorithms for high-performance semantic communication. However, the challenges, such as energy-hungry and efficiency-limited image retrieval manner, and semantic encoding without considering user personality, have not been explored yet. These challenges have hindered the widespread adoption of semantic communication. To address the above challenges, at the semantic level, we first design an energy-efficient task-oriented semantic communication framework with a triple-based {\color{black}scene graph} for image information. We then design a new personalized semantic encoder based on user interests to meet the requirements of personalized saliency. Moreover, at the communication level, we study the effects of dynamic wireless fading channels on semantic transmission mathematically and thus design an optimal multi-user resource allocation scheme by using game theory. Numerical results based on real-world datasets clearly indicate that the proposed framework and schemes significantly enhance the personalization and anti-interference performance of semantic communication, and are also efficient to improve the communication quality of semantic communication services.
I. INTRODUCTION
The paper targets UAV image-sensing-driven task-oriented semantic communication by addressing inefficient image retrieval, non-personalized semantic encoding, and limited analysis of wireless fading. It proposes a scene-graph framework, personalized encoding, mathematical channel analysis, and game-based multi-user resource allocation.
- I. INTRODUCTION: Traditional UAV communication transmits all captured images, wasting communication resources and UAV energy on images users do not need.Keyword-based retrieval can still suffer packet drops and retransmissions under limited energy or poor wireless channels.
- I. INTRODUCTION: Existing semantic encoding methods do not account for individual user preferences, risking loss of information important to particular users during fading transmission.The paper motivates personalized coding weights and user-specific resource allocation.
- I. INTRODUCTION: The proposed framework models image information as a triple-based scene graph to retrieve selected images matching users’ query text.This replaces transmission of all images with semantic matching of user queries to image information.
- I. INTRODUCTION: A personalized semantic encoder assigns higher weights to triplets more relevant to each user, supporting differential transmission of important information.The weighting mechanism is based on users’ subjective interests.
- I. INTRODUCTION: The paper mathematically analyzes semantic triplet transmission under generalized fading channels and derives a game-theory-based multi-user resource allocation scheme.The scheme uses a retrieval-task utility function to improve utilization of limited UAV resources.
II. RELATED WORK
Prior work covers semantic encoding, task-oriented communication, privacy, and fading-channel performance, but task-oriented systems generally ignore differences among users. The proposed system combines personalized saliency, semantic triplet retrieval, and power allocation for UAV image services.
- II. RELATED WORK: Semantic communication transmits extracted semantic information rather than all original data, reducing unnecessary communication overhead.The receiver reconstructs the original information from the transmitted semantic representation.
- II. RELATED WORK: Prior task-oriented communication systems adapt transmitted semantics to goals but ignore individual differences in users’ needs.This motivates incorporating personalized saliency into the proposed framework.
- II. RELATED WORK: Wireless semantic communication performance is negatively affected by multipath fading, while prior studies examined selected channel models such as Rayleigh and Rician fading.The paper positions its mathematical fading analysis against this prior performance-analysis literature.
- II. RELATED WORK: The proposed system predicts user saliency, extracts image triplets, fuses objective and subjective attention, and uses the result for personalized semantic communication.This pipeline connects customized user information with UAV-captured imagery.
- II. RELATED WORK: For each user, the UAV extracts image triplets and transmits them using TDMA, while power is allocated across users and individual triplets under an energy constraint.The two optimization problems separately address inter-user allocation and within-user triplet allocation.
C. SINR Analysis
The section models a multi-user UAV downlink in which beamforming and transmit powers determine each user's SINR under interference and noise. Users are positioned relative to a fixed-altitude UAV, with channel quality affected by distance and path loss.
- System Model: The UAV serves K users from NT antennas while transmitting semantic features from aerial images.Each user has an individual preference, and the UAV uses its antenna array to obtain array gains and improve channel quality.
- Geometry: User positions are represented by uk, the UAV position by uu, and their separation determines the UAV-to-user link distance.The kth user's horizontal coordinate is uk=(xk,yk,0), while the UAV hovers at fixed altitude zu.
- Signal Model: Linear beamforming multiplies user k's data symbol by beamformer wk, while interfering paths contribute interference power at the receiver.The model includes NIk interfering paths, average interfering transmit power, and interference symbols for each user.
- SINR: The kth user's SINR depends on desired transmit power, channel gain, interference, and noise.The section introduces Pk as the kth user's transmit power and σ2 as the noise term in the SINR expression.
- Beamforming: Maximum ratio transmission provides the optimal beamforming vector for the modeled user channel.The beamformer is expressed using the channel vector hk.
D. Channel Model
The channel model uses Fisher-Snedecor F fading for UAV links and Rayleigh fading for interference, then derives distributions for channel gains and SINR. Because exact multi-antenna expressions are analytically difficult, the summed desired-channel gain is approximated by a single F distribution.
- Desired Link: The UAV-to-user link follows Fisher-Snedecor F composite fading, combining Nakagami-m small-scale variation with inverse-Nakagami-m shadowing.Measurements at 5.8 GHz are cited as supporting this model in both line-of-sight and non-line-of-sight scenarios.
- Desired Link: The desired beamformed gain is a sum of NT Fisher-Snedecor F random variables, which is approximated by a single Fisher-Snedecor F distribution.The approximation avoids the less interpretable multivariate Fox’s H-function representation.
- Distributional Model: The approximated gain Z is modeled as F(mfk, msk, z̄k), whose PDF and CDF are used in subsequent analysis.The notation identifies the multipath fading parameter, shadowing parameter, and average gain.
- Interference: Interference signals are modeled with Rayleigh fading, and the large-scale interference fading is incorporated into the mean channel value.The sum of NIk independent Rayleigh-fading signals has a Nakagami-m amplitude with m=NIk.
- SINR Distribution: Theorem 1 provides closed-form PDF and CDF expressions for the SINR using the desired-channel and interference distributions.The derivation uses the modeled channel gain, interference statistics, and transmit-power terms.
E. Approximation Analysis
The section develops an accurate approximation for a difficult integral appearing in the SINR CDF, then derives high-transmit-power asymptotics and verifies them through outage probability comparisons.
- Motivation: Closed-form PDF and CDF expressions are difficult to interpret directly, motivating an approximate CDF and high-SNDR analysis.The paper states that numerical analysis is used to verify the resulting expressions.
- Accurate Approximation: The integral IA lacks an established closed solution, so the paper derives an accurate approximation for it.The authors characterize IA as absent from the cited mathematical integral references and difficult to solve exactly.
- Accurate Approximation: Lemma 1 gives an accurate approximation of IA when ρ is small.The approximation is then substituted into the SINR CDF derivation.
- Asymptotic Analysis: Theorem 3 approximates the SINR CDF in the high-transmit-power regime.This asymptotic expression is developed after the approximate CDF is obtained.
- Verification and Insights: The accurate approximation nearly matches the closed-form expression, while the high-SINR approximation is close when Pk exceeds 25 dBW.The outage probability comparison also shows that larger mfk produces a faster decrease in outage probability as transmit power increases.
A. Bit Error Probability
The section formulates bit error probability for semantic triplet transmission under modulation-specific conditional error behavior. It connects triplet encoding and transmission conditions to the probability that semantic information is dropped.
- Bit Error Probability: Under different modulation formats, the kth user's bit error probability is expressed using modulation-specific parameters λ1 and λ2.The conditional bit error probability is represented through the incomplete-gamma expression in the model.
- Bit Error Probability: Theorem 4 derives the bit error probability Ek for user k.The theorem provides the analytical expression used for subsequent triplet-drop calculations.
- Triplet Drop Probability: The triplet drop probability Pk is calculated from the kth user's bit error probability Ek and the derived expression.The drop probability is evaluated with the help of the preceding bit-error result.
- Application Context: In UAV aerial photography, transmitting every aerial image to every user is impractical in fading wireless environments.The application instead motivates selective delivery based on users' image needs.
A. OA-SemCom: A Fully Objective Approach
PERSF-SEMCOM detects image triplets, combines objective visual attention with user-specific saliency, prioritizes triplets, and allocates transmission power according to each user’s preferences.
- Triplet Detection: Triplet Detection extracts subject-relation-object triplets, attention heatmaps, and subject/object bounding boxes from each captured image.These outputs provide the semantic representation and spatial information used by later personalization and priority modules.
- Personalized Saliency Fusion: PERSF-SEMCOM fuses RelTR’s objective attention with a personalized saliency heatmap to estimate user-specific triplet priorities.The personalized saliency predictor uses user identity and the captured image; attention maps are normalized and combined by weighted sum.
- Triplet Priority Estimation: The fused subject and object heatmaps are cropped by entity boxes, and their maxima are multiplied to calculate each triplet’s priority.This priority is then used to distinguish key triplets for individual users.
- Power Allocation and Retrieval: The power-allocation module assigns more transmission power to higher-priority triplets, reducing their likelihood of being dropped over the wireless channel.The main algorithm uses RCGA for allocation, while the receiver matches received triplets against the user’s query through accurate or fuzzy search.
- Power Allocation and Retrieval: At the receiver, Triplet Search computes a query-match score and downloads the current image when the score exceeds a user-specified threshold.TS-AM supports exact triplet-format queries, whereas TS-FM supports free-text queries.
VI. NUMERICAL RESULTS
The numerical evaluation uses pretrained scene-graph and saliency models, real-world visual data, channel settings from prior work, and Naive-SemCom and OA-SemCom as benchmarks.
- Experimental Platform: The platform uses Ubuntu 18.04, an Intel Xeon E5-2678 CPU, and four GeForce RTX 2080 Ti GPUs.RelTR is pretrained on Visual Genome, which contains 108k images, 150 object categories, and 50 relationship categories.
- Models and Data: RelTR serves as the Triplet Detection model, while a pretrained saliency-prediction model supports the personalized saliency module.RelTR’s reported top-50 scene-graph recall is 25.2, with a corresponding mean value of 8.52.
- Models and Data: Unless otherwise specified, small-scale fading and system parameters such as transmission power and distance follow common settings from prior literature.The study also references the STREET dataset and channel-configuration parameters for three users.
- Baselines: Naive-SemCom allocates equal power without semantic-triplet importance, whereas OA-SemCom quantifies triplet priorities for scheduling.These two approaches provide the performance-comparison baselines.
B. Results and Analysis
Experiments show that personalized saliency, priority-based allocation, and channel-aware analysis improve utility, reduce optimality gaps and communication overhead, and reveal sensitivity to fading, interference, and distance.
- Effectiveness of PERSF-SEMCOM: The optimal utility typically occurs at α between 0.1 and 0.2, where subjective attention contributes 80%-90% of fused attention.PERSF-SEMCOM outperforms Naive-SemCom in most tested cases, while α = 1.0 degenerates to OA-SemCom and performs worse than Naive-SemCom.
- Transmit Power: At P = 3kW, PERSF-SEMCOM reduces the optimality gap by 54%, 57%, and 37% for the three users, respectively.OA-SemCom instead widens the gaps by 26%, 57%, and 86%.
- Channel Conditions: When multipath and shadow fading weaken, utility increases; utility responds faster to mfk than to an equivalent increase in msk.The analysis attributes the stronger sensitivity to BER increases caused by multipath effects rather than shadowing.
- Channel Conditions: When interference exceeds 10 W, every additional 5 m of transmission distance causes an approximately 58% utility decrease.The paper recommends reducing transmission distance by adjusting the UAV trajectory under substantial interference.
- Communication Overhead: PERSF-SEMCOM reduces transferred image data from 224.8MB to 81.29MB, a 64% communication-cost reduction, after sending 873 triplets with 10.24KB overhead.The selective process sends 64 images to specified subscribers instead of transmitting all captured images.
- Overall Findings: The conclusion reports that the framework and schemes achieve personalized semantic communication and significantly enhance UAV resource utilization on real-world datasets.The approach combines a triple-based scene graph, personalized attention-based encoding, fading analysis, and game-based multi-user allocation.
APPENDIX A PROOF OF LEMMA 1
The appendix derives the closed-form expression in Lemma 1 by substituting intermediate expressions, evaluating an integral, and applying the bivariate Meijer’s G-function definition.
- Derivation: The proof substitutes equation (7) and a cited equation into (A-1), then rewrites the exponential function using a Mellin-Barnes integral.These substitutions transform the expression into a form suitable for evaluating the integral term I1.
- Conclusion: Substituting I1 into (A-2) and applying the bivariate Meijer’s G-function definition yields equation (9) and completes the proof.The special-function representation provides the final closed-form result.
B. Proof of CDF
The proof derives the CDF by expressing intermediate integrals, applying hyper-geometric and Meijer’s G-function identities, and substituting the resulting forms into earlier expressions.
- The derivation begins from the CDF definition and combines equations (9) and (A-8).
- Intermediate integral I2 is expressed and solved before being substituted into (A-9).
- The integral expression of the hyper-geometric function is used to obtain (10) and complete the proof.
- I3 is solved using cited integral identities and then substituted into (B-3).
- The expression for IA is rewritten as (12) using the definition of Meijer’s G-function.
APPENDIX C PROOF OF THEOREM 2
Theorem 2 is proved by deriving the CDF of γ_k and evaluating E_k through successive substitutions and special-function identities.
- The proof starts from the definition of the CDF and reduces the integration in (C-3) using Lemma 1.
- E_k is first expressed using the definition of the Gamma function.
- Expression (C-4) is substituted into (D-1), after which integral I4 is solved using a cited identity.
- Combining I4 with (D-2) produces E_k as (16), completing the proof.