Source-linked AI summary

X-ModalNet: A Semi-Supervised Deep Cross-Modal Network for Classification of Remote Sensing Data

Danfeng Hong, Naoto Yokoya, Gui-Song Xia, Jocelyn Chanussot, Xiao Xiang Zhu

arXiv:2006.13806v2cs.CV

TL;DR

Remote sensing classification must transfer useful information between modalities because HSI is information-rich but locally limited, while MSI and SAR are broadly available yet less discriminative and sparsely labeled. X-ModalNet performs semi-supervised cross-modal learning with adversarial, interactive, and label-propagation components. It improves over state-of-the-art methods on HSI-MSI and HSI-SAR datasets, while full use of unlabeled samples incurs sharply increasing storage, transmission, and computation costs.

  • Problem

    Limited HSI coverage, poorer MSI or SAR feature discrimination, noisy observations, and scarce annotations create a cross-modality classification challenge in remote sensing.

  • Method

    X-ModalNet transfers HSI information to MSI or SAR classification using self-adversarial and interactive-learning modules plus iterative label propagation on high-level feature graphs.

  • Results

    X-ModalNet outperforms state-of-the-art methods, with at least 6% higher Pixel Acc. and mIoU than CorrNet on homogeneous datasets and about 9% higher Pixel Acc. and 10% higher mIoU on HSI-SAR.

  • Takeaways & Limitations

    The framework demonstrates semi-supervised cross-modality learning that transfers limited HSI knowledge to large-scale MSI or SAR classification.

  • Takeaways & Limitations

    Using all unlabeled samples improves classification but causes exponentially increasing storage, transmission, and computation costs, so the experiments use the test set as the unlabeled set.

Abstract

from arXiv · show

This paper addresses the problem of semi-supervised transfer learning with limited cross-modality data in remote sensing. A large amount of multi-modal earth observation images, such as multispectral imagery (MSI) or synthetic aperture radar (SAR) data, are openly available on a global scale, enabling parsing global urban scenes through remote sensing imagery. However, their ability in identifying materials (pixel-wise classification) remains limited, due to the noisy collection environment and poor discriminative information as well as limited number of well-annotated training images. To this end, we propose a novel cross-modal deep-learning framework, called X-ModalNet, with three well-designed modules: self-adversarial module, interactive learning module, and label propagation module, by learning to transfer more discriminative information from a small-scale hyperspectral image (HSI) into the classification task using a large-scale MSI or SAR data. Significantly, X-ModalNet generalizes well, owing to propagating labels on an updatable graph constructed by high-level features on the top of the network, yielding semi-supervised cross-modality learning. We evaluate X-ModalNet on two multi-modal remote sensing datasets (HSI-MSI and HSI-SAR) and achieve a significant improvement in comparison with several state-of-the-art methods.

1. Introduction

Remote sensing offers large-scale multimodal imagery for urban understanding, but limited labels, noisy observations, modality differences, and incomplete high-information coverage make pixel-wise classification difficult. X-ModalNet addresses this cross-modality setting by transferring HSI information to large-scale MSI or SAR prediction through semi-supervised multimodal learning.

  • Motivation and Objective: Remote sensing imagery supports large-scale urban scene understanding and pixel-wise semantic classification across applications.Operational satellites provide globally available SAR and MSI data, while classification assigns a semantic category to each studied-scene pixel.
  • Motivation and Objective: HSI provides rich spectral information but limited coverage, whereas MSI and SAR are widely available yet have poorer feature representation, creating a cross-modality learning problem.The paper asks whether limited HSI can improve classification of large-scale MSI or SAR data.
  • Motivation and Objective: Limited labels, noisy annotations, environmental variation, and sensor noise constrain reliable remote sensing classification.The paper identifies costly labeling and variability from illumination, topology, atmospheric effects, and instrument configurations as central challenges.
  • Method Overview and Contributions: X-ModalNet uses a hybrid backbone with CNN processing for MSI or SAR and DNN processing for HSI to combine spatial and spectral information.The architecture is designed to use MSI or SAR spatial resolution together with HSI spectral resolution.
  • Method Overview and Contributions: The framework combines self-adversarial and interactive-learning modules with iterative label propagation to improve robust multimodal representations and exploit unlabeled samples.Label propagation progressively updates pseudo-labels on an updatable graph, while the two plug-and-play modules target robustness and discriminative cross-modal fusion.
  • Method Overview and Contributions: X-ModalNet is evaluated on two cross-modal datasets, HSI-MSI and HSI-SAR, with extensive ablation analysis.The second dataset uses collected and processed Sentinel-1 SAR data.

2. Related Work

Related work spans deep learning for scene parsing, multimodal remote sensing analysis, and semi-supervised learning. Existing approaches provide useful fusion or representation strategies but remain limited by labeling costs, rough-grained inputs, or linearized modeling for heterogeneous data.

  • Deep Learning for Scene Parsing: Deep CNN-based scene-parsing methods have advanced rapidly, but satellite and aerial imagery remain less investigated than street-view imagery.Large urban areas require highly diverse training samples because remote sensing offers a nearly horizontal field of vision.
  • Remote Sensing Image Classification: HSI classification work has used DNNs, spatial-spectral CNNs, and cascaded RNNs to model spectral and spatial information.These approaches focus on predicting category labels or parsing HSI scenes.
  • Multimodal Representation Learning: Multimodal representation learning seeks a joint discriminative space, but linearized methods remain limited in data representation and fusion, especially for heterogeneous data.Other approaches couple modality-specific subnetworks using similarity, correlation, or sequentiality constraints.
  • Multimodal Remote Sensing Analysis: Remote sensing multimodal networks have combined optical, OpenStreetMap, Lidar, or multiple image streams for semantic mapping and material segmentation.Many such methods target rough-grained scene parsing and can struggle in complex urban scenes because of relatively poor feature representation.
  • Semi-Supervised Remote Sensing Learning: Semi-supervised remote sensing studies use unlabeled samples to address expensive labeling, including regression-based multitask learning, manifold alignment, and factor analysis.The paper positions deep-learning-based semi-supervised multimodal learning as less investigated in this context.

3. The Proposed X-ModalNet

X-ModalNet is a three-stream semi-supervised multimodal network for remote-sensing classification, combining CNN pathways for MSI or SAR with a DNN pathway for HSI. Its self-adversarial, interactive-learning, and label-propagation modules target robustness, compact cross-modal representations, and improved use of unlabeled samples.

  • Network Architecture: X-ModalNet uses a hybrid three-stream architecture, with CNN streams for MSI or SAR and a DNN stream for HSI.The design exploits MSI/SAR spatial information and HSI spectral information within a semi-supervised multimodal framework.
  • Self-Adversarial Module: The self-adversarial module splits feature inputs into two streams and concatenates original and adversarial representations for end-to-end robust feature learning.Its adversarial features are learned within X-ModalNet rather than generated by an independently trained GAN.
  • Interactive Learning Module: Interactive learning copies weights between modalities and fuses the resulting representations through addition to reduce cross-modal gaps without additional computational cost.The strategy is designed to produce smoother, more compact multimodal feature blending.
  • Label Propagation Module: The label-propagation module initializes pseudo-labels for unlabeled samples, trains the network, and updates labels using high-level features and graph-based propagation.The process repeats until pseudo-labels stabilize; experiments found that three to four repetitions are usually sufficient for convergence.
  • Objective Function: The overall objective jointly combines labeled-sample cross-entropy, pseudolabel loss, reconstruction loss, and adversarial loss.Reconstruction is applied to each modality and to unlabeled samples, while adversarial loss acts on the self-adversarial module.

4.1. Data Description

The evaluation uses two cross-modal remote-sensing datasets, with sparse labels and hyperspectral data available only during training. The HSI-MSI data include simulated spectral and spatial resolutions, while the HSI-SAR data pair EnMap HSI with Sentinel-1 SAR imagery.

  • X-ModalNet is evaluated on HSI-MSI and HSI-SAR datasets using corresponding training and test ground-truth maps.Figure 4 presents false-color imagery and label maps for both datasets.
  • HSI is assumed to be available during training but absent during testing, defining the cross-modality learning setting.
  • Homogeneous HSI-MSI Dataset: The HSI-MSI scene contains 349×1905 pixels and 144 bands at 10 m GSD, aligned with eight-band MSI at 2.5 m GSD.MSI is generated through spectral degradation using Sentinel-2 response functions.
  • Homogeneous HSI-MSI Dataset: Spatial simulation produces HSI at 10 m GSD by applying an isotropic Gaussian point spread function and upsampling to the MSI size.
  • Heterogeneous HSI-SAR Dataset: The HSI-SAR data combine a 797×220-pixel, 244-band EnMap HSI scene at 30 m GSD with Sentinel-1 SAR imagery at 13 m GSD.The SAR image has four polarimetric bands.

4.2. Implementation Details

The model is trained with standard deep-learning regularization and optimization procedures, using validation-based hyperparameter selection. Semi-supervised training incorporates labeled and unlabeled target-modality samples, with the test set selected as unlabeled data for fair comparison.

  • Hyperparameters are selected by grid search on a validation set, while Adam optimization uses a poly learning-rate policy.The implementation is built with TensorFlow.
  • Batch normalization and dropout are applied before activation functions to facilitate training and reduce overfitting.
  • Training lasts 150 epochs for HSI-MSI and 200 epochs for HSI-SAR, with minibatches of 300.
  • Labeled and unlabeled MSI or SAR samples share network parameters during optimization.
  • The test set is used as the unlabeled set for all semi-supervised compared methods to ensure a fair comparison.The authors report that similarly sized independent unlabeled samples produce similar classification results.

4.3. Comparison with State-of-the-art

X-ModalNet is compared with linear, autoencoder-based, correlation-based, and cross-weighted multimodal baselines on HSI-MSI and HSI-SAR datasets. It achieves the strongest reported performance, with especially large gains on heterogeneous HSI-SAR data.

  • The baselines include a linear SVM, CCA, unimodal DAE, bimodal DAE, bimodal SDAE, MDL-CW, Corr-AE, and CorrNet.These methods represent direct classification, shared-subspace, autoencoder, correlation, and cross-weighted fusion strategies.
  • Results on the Homogeneous Datasets: On HSI-MSI, bimodal SDAE reaches around 79% accuracy after improving over earlier multimodal baselines.Table 4 reports Pixel Acc. and mIoU as evaluation measures.
  • Results on the Homogeneous Datasets: At least 6% higher Pixel Acc. and mIoU than CorrNet are reported for X-ModalNet on the HSI-MSI comparison.CorrNet is identified as the second-best method in this comparison.
  • Results on the Heterogeneous Datasets: On HSI-SAR, X-ModalNet improves over CorrNet by about 9% Pixel Acc. and 10% mIoU.The reported trend favors methods using hyperspectral information over Baseline and Unimodal DAE.
  • Results on the Heterogeneous Datasets: CCA performs worse on heterogeneous HSI-SAR data than on the homogeneous setting, while MDL-CW exceeds most compared methods and nearly 20% over baseline is reported.The passage attributes this pattern to heterogeneous-data representation and interactive weight learning.

4.4. Visual Comparison

Visual comparisons examine highlighted regions in HSI-MSI and HSI-SAR classification maps, alongside t-SNE feature visualizations and module ablations. X-ModalNet produces more effective or realistic material parsing, while multimodal methods yield smoother maps.

  • On the Houston2013 HSI-MSI region, X-ModalNet identifies materials more effectively, particularly the Commercial class.Multimodal-input methods also produce smoother parsing results than single-modality methods.
  • On the EnMap HSI-SAR region, X-ModalNet produces a more competitive and realistic parsing result, especially for Soil and Plants.
  • Ablation analysis compares X-ModalNet module combinations using Pixel Acc. on both datasets and examines batch normalization and dropout.
  • The visual analysis is complemented by t-SNE visualizations of multimodal features learned with different module combinations.

4.5. Ablation Studies

The ablation study shows that progressively adding X-ModalNet components improves feature discrimination and classification performance, while dropout and batch normalization support generalization.

  • Ablation results: Progressively adding X-ModalNet components generates more discriminative features in the top encoder layer.The learned latent-space features become more discriminative as modules are added step by step.
  • Ablation results: Without proposed modules, Pixel Acc. is about 83.14% for HSI-MSI and 64.44% for HSI-SAR.
  • Ablation results: Adding IL improves classification accuracy by around 2% ∼3% over the configuration without proposed modules.
  • Ablation results: Adding LP after IL increases Pixel Acc. by 1.5% for HSI-MSI and 2% for HSI-SAR.
  • Ablation results: Removing dropout degrades generalization, while removing BN sharply reduces classification accuracy.The paper attributes the BN-related decline to low-efficiency gradient propagation that harms network learning.

4.6. Robustness to Noises

The robustness experiment evaluates X-ModalNet under Gaussian white-noise attacks across multiple signal-to-noise ratios, comparing models with and without the SA module.

  • Noise attack setup: Gaussian white noise is added at SNRs from 10 dB to 40 dB in 10 dB intervals to simulate corrupted inputs.
  • Noise attack setup: Figure 8 compares Pixel Acc. before and after enabling the SA module on both datasets.

5. Conclusion

The conclusion presents X-ModalNet as a cross-modal classification framework that transfers HSI knowledge to large-scale MSI or SAR prediction. It combines interactive learning, self-adversarial training, and iterative label propagation to improve representation, noise resistance, and performance.

  • Conclusion: X-ModalNet transfers knowledge from locally collected HSI to large-scale MSI or SAR for cross-modal classification.
  • Conclusion: The framework combines IL for discriminative representation, SA for noise resistance, and iterative LP for further performance improvement.
  • Conclusion: The paper identifies incorporating the physical mechanism of spectral imaging into network learning as future work.
Loading 2006.13806v2…