Source-linked AI summary
SGPDFuse: Semantically-Guided Physics-Disentanglement General Multi-Modal Image Fusion
Haozhen Wei, Chengjun Jiang, Yutong Guo, Xinrui Ju, Xingyuan Li, Xiang Chen, Jinyuan Liu
TL;DR
Multimodal image fusion must preserve scene content while suppressing sensor-specific degradation, but existing aggregation methods do not separate these factors. SGPDFuse uses a foundation-model-based Semantic-Physical Parametric Bridge and semantic alignment to perform physics-disentangled fusion, achieving state-of-the-art performance across three benchmark task families with one architecture.
Problem
Existing fusion methods aggregate entangled multimodal signals without distinguishing invariant scene content from transient environmental variation, while prior physical approaches require task-specific supervision or measurements.
Method
SGPDFuse predicts physical variation confidence maps from DINOv3 features, disentangles intrinsic components, and supervises fusion with cosine semantic alignment and Gram-matrix texture regularization.
Results
SGPDFuse achieves state-of-the-art performance across infrared-visible, multi-focus, and multi-exposure benchmarks using a single architecture.
Takeaways & Limitations
The unified framework supports physics-guided fusion across distinct tasks while preserving semantic content, physical texture, and downstream recognition performance.
Takeaways & Limitations
The factorization assumes that variation and intrinsic components predominantly occupy distinct frequency bands, although this spectral alignment is not absolute.
Abstract
from arXiv · showhide
Multimodal image fusion (MMIF) aims to integrate complementary sensor data into a single representation that preserves intrinsic scene reality while eliminating environmental interferences. Most existing approaches rely on blind feature aggregation, which excels at signal accumulation but fails to distinguish essential content from physical degradations. We propose SGPDFuse, which bridges this gap by mapping inputs into a physics-disentangled structural representation via a Semantic-Physical Parametric Bridge (SPPB) built on pretrained vision foundation models, utilizing the Intrinsic-Variation principle to decouple invariant scene attributes from transient environmental factors. To guide this decomposition, we introduce a Semantic Alignment mechanism: we explicitly anchor the fused representation to salient semantic features in the same foundation model feature space via cosine similarity to preserve critical targets, while enforcing physical texture fidelity through Gram-matrix regularization to strictly eliminate unnatural artifacts. Extensive experiments demonstrate that SGPDFuse achieves state-of-the-art performance across infrared-visible, multi-focus, and multi-exposure benchmarks using a single architecture.
1 Introduction
SGPDFuse reframes multimodal image fusion as physics disentanglement, separating invariant scene content from environmental variation rather than blindly aggregating signals. It uses semantic foundation-model features to guide this decomposition and semantic-physical supervision across multiple fusion tasks.
- Motivation: Existing fusion methods conflate intrinsic scene properties with environmental variations, producing artifacts such as ghost targets, exposure distortions, and texture loss.The affected factors include temperature, reflectance, and sharp texture versus reflections, illumination changes, and optical blur.
- Proposed framework: SGPDFuse uses a Semantic-Physical Parametric Bridge built on LoRA-adapted DINOv3 features to predict variation confidence maps for disentanglement.The maps drive decomposition of both base and detail features across the input modalities.
- Proposed framework: Semantic Alignment preserves salient content through cosine similarity while Gram-matrix regularization enforces physical texture fidelity and suppresses unnatural artifacts.The two forms of supervision operate in the DINOv3 feature space and jointly guide the fused representation.
- Scope and evaluation: The Reconstruction Alignment loss illustrates how the fused output is aligned with source representations in the proposed fusion framework.Figure 2 is identified as an illustration of this loss for multimodal image fusion.
- Scope and evaluation: The framework applies one architecture across infrared-visible, multi-focus, and multi-exposure fusion tasks.Its stated objective is unified physics-guided decomposition across these distinct settings.
2 Related Work
Prior fusion methods specialize in task-specific feature aggregation or physical modeling, but they generally do not estimate physical variation parameters from entangled multimodal observations without explicit measurements. SGPDFuse addresses this gap using invariant semantic features from vision foundation models to support unified decomposition.
- Task-specific fusion: Task-specific fusion methods aggregate raw-source features, which can preserve environmental interference as ghost targets, defocused texture loss, or exposure halos.These artifacts arise because intrinsic scene content and extrinsic factors are treated as equally valid fusion targets.
- Physics-informed methods: Physical-modeling approaches use domain-specific formulations such as thermal emission, reflectance-illumination, or defocus models, but remain confined to individual tasks.The cited approaches require specialized supervision including emissivity measurements, paired exposures, or depth ground truth.
- Open gap: Estimating physical variation parameters from entangled multimodal observations without explicit physical measurements remains unaddressed in prior work.This is identified as a critical unresolved challenge in the related-work discussion.
- Proposed perspective: SGPDFuse proposes that DINO semantic features encode material and structural properties invariant to physical variations, enabling unified Intrinsic-Variation decomposition.The foundation-model perspective is presented as a bridge between semantic representation and physical disentanglement.
3 Intrinsic-Variation Structure in Multi-Modal Imaging
Multi-modal observations mix stable scene properties with transient environmental or sensor-induced variation, making explicit disentanglement necessary for artifact-free fusion. SGPDFuse uses scale-aligned factorization to separate intrinsic content from variation before fusing.
- Observed modalities combine stable physical scene properties with transient, sensor-specific interference that existing fusion methods leave entangled.
- Intrinsic components represent modality-invariant scene properties, whereas variation components represent environmental or sensor-induced factors that fusion should suppress.
- Blind aggregation retains a weighted residual of variation components, so artifact-free fusion requires explicit intrinsic estimation and variation suppression.
- Infrared-visible, multi-exposure, and multi-focus imaging instantiate the intrinsic–variation structure through thermal reflection, illumination, and blur-related factors.
- Intrinsic and variation components tend to occupy different frequency bands, motivating independent disentanglement across base and detail scales despite no absolute frequency partition.
- The resulting pipeline decomposes sources into base and detail, estimates W, disentangles Φ and Ψ at each scale, and fuses only intrinsic components.
4 Methodology
SGPDFuse disentangles intrinsic scene content from environmental variation before fusion. It uses semantic features to estimate physical variation, retains intrinsic components, and aligns the fused output with source semantics and textures.
- Framework overview: SGPDFuse uses a shared dual-branch encoder and SPPB to estimate variation confidence, disentangle intrinsic Φ from variation Ψ, and fuse only Φ.The pipeline processes base and detail features before physics-guided aggregation.
- Semantic-physical bridge: DINOv3 semantic tokens are injected into the detail branch and separately used by SPPB to estimate task-specific physical parameters.The parameters include emissivity, exposure optimality, and in-focus probability.
- Intrinsic-variation disentanglement: The complementary assignment sets W_A=W and W_B=1−W, while residual correctors refine cases where both modalities contain distinct intrinsic information.Variation components are discarded before fusion at both base and detail scales.
- Physics-guided fusion: Base features use W-biased squeeze-and-excitation weighting, whereas detail features use ℓ1 activity weighting to select locally sharper content.This asymmetric design follows the differing frequency-domain properties of intrinsic content.
- Qualitative results: Qualitative comparisons report superior visual quality and enhanced detail preservation across diverse multimodal fusion tasks.Figure 4 uses magnified views to compare SGPDFuse with other state-of-the-art methods.
- Reconstruction alignment: Reconstruction alignment anchors the fused output to DINOv3 source features through cosine content alignment and Gram-matrix style regularization.The loss uses cosine similarity and Gram matrices in the shared semantic feature space.
5.1 Experimental Setup
The evaluation covers infrared-visible, multi-exposure, and multi-focus fusion using established benchmarks, full-resolution testing, and complementary fusion and downstream metrics.
- Datasets: Experiments evaluate infrared-visible, multi-exposure, and multi-focus fusion across M3FD, RoadScene, TNO, MEFB, and MFIF.Downstream infrared-visible evaluation uses M3FD for object detection and FMB for semantic segmentation.
- Implementation: Test images are evaluated at full resolution without post-processing, using a shared encoder with 64 feature channels and four NAFNet blocks.The implementation also uses DINOv3 ViT-S/16 with LoRA adaptation.
- Metrics: Fusion quality is measured with CC, PSNR, TE, and MS-SSIM, while downstream tasks use mAP50:95 for detection and mIoU for segmentation.MS-SSIM is reported as SSIM in Table 1.
5.2 Experiments on Multi-Task
Across infrared-visible, multi-focus, and multi-exposure benchmarks, SGPDFuse produces sharper, more faithful fusion results and ranks best or second-best across reported metrics.
- Qualitative Comparison: SGPDFuse recovers sharper vehicle outlines and headlight structures on M3FD while retaining fine-grained branch and ground textures on TNO.Competing methods exhibit blurring or halo artifacts on TNO due to unresolved thermal emissivity variation.
- Qualitative Comparison: For multi-focus fusion, SGPDFuse achieves the clearest focus transitions, while competing methods show ringing or local blur.
- Qualitative Comparison: For multi-exposure fusion, SGPDFuse jointly recovers over-exposed and underexposed details with balanced luminance free of tone-mapping artifacts.
- Quantitative Comparison: SGPDFuse achieves the best or second-best results across all datasets and metrics, with MS-SSIM gains of +0.07, +0.09, +0.05, and +0.02 over respective runners-up.The reported gains correspond to M3FD, T&R, MFIF, and MEFB, respectively.
- Quantitative Comparison: Consistent TE and PSNR gains span three fusion tasks under distinct physical priors, supporting the architecture’s reported generalizability.
5.3 Downstream Application on IVIF
Downstream evaluations show that SGPDFuse’s fused images support strong object detection and semantic segmentation performance on IVIF benchmarks.
- Downstream Evaluation: SGPDFuse achieves the highest mAP50:95 on M3FD and the best mIoU on FMB among competing methods.
- Downstream Evaluation: Detection gains on thermally distinctive Car, Motorcycle, and Trunk categories indicate more faithful preservation of object signatures.
- Downstream Evaluation: Improved mIoU on FMB confirms preservation of semantic boundaries needed for recognition.
5.4 Ablation Studies
Ablation studies show that physics-grounded weighting, intrinsic–variation disentanglement, LoRA adaptation, and semantic alignment each contribute to SGPDFuse’s performance.
- SPPB Ablation: Replacing the physics-grounded confidence map with uniform weighting causes consistent degradation across all metrics.The uniform variant uses W ≡0.5 and fails to distinguish intrinsic from variation content.
- SPPB Ablation: Removing intrinsic–variation disentanglement causes further performance drops, showing that W-weighted projection suppresses modality-specific artifacts before fusion.
- SPPB Ablation: Freezing DINOv3 without LoRA adaptation hurts performance, indicating that task-aware fine-tuning is needed for material-sensitive emissivity estimation.
- Semantic Alignment Ablation: Removing LRecA degrades structural fidelity, while ablating Lcontent or Lstyle causes semantic drift or texture incoherence, respectively.The results support content and style alignment as complementary components of fusion training.
6 Conclusion
SGPDFuse replaces blind feature aggregation with physics-disentangled representations that separate invariant scene attributes from transient environmental factors. Semantic and texture alignment then supervise fusion in DINOv3 feature space, with experiments showing consistent state-of-the-art performance in one unified architecture.
- Conclusion: SGPDFuse uses a Semantic-Physical Parametric Bridge to separate invariant scene attributes from transient environmental factors under the Intrinsic-Variation principle.
- Conclusion: Physics-guided fusion operates on intrinsic components, while cosine similarity and Gram-matrix regularization provide semantic and texture supervision in DINOv3 feature space.
- Conclusion: Extensive experiments demonstrate consistent state-of-the-art performance within a single unified architecture.