Source-linked AI summary
Rotated Multi-Scale Interaction Network for Referring Remote Sensing Image Segmentation
Sihan Liu, Yiwei Ma, Xiaoqing Zhang, Haowei Wang, Jiayi Ji, Xiaoshuai Sun, Rongrong Ji
TL;DR
Referring remote-sensing segmentation is challenged by the spatial-scale and orientation diversity of aerial imagery and the limited scale and scope of existing datasets. The paper introduces RRSIS-D and RMSIN, whose validation on the new dataset is reported to show superior performance.
Problem
Aerial imagery’s spatial-scale and orientation diversity challenges conventional RIS methods, and existing RRSIS datasets are limited in scale and scope.
Method
RMSIN combines within-scale IIM, cross-scale CIM, and decoder-based ARC, and the paper introduces the 17,402-triplet RRSIS-D benchmark.
Results
RMSIN’s validation on RRSIS-D is reported to show superior performance and establish a new benchmark for future research.
Takeaways & Limitations
RRSIS-D provides a large, varied resource for rigorous evaluation of RRSIS methods.
Abstract
from arXiv · showhide
Referring Remote Sensing Image Segmentation (RRSIS) is a new challenge that combines computer vision and natural language processing, delineating specific regions in aerial images as described by textual queries. Traditional Referring Image Segmentation (RIS) approaches have been impeded by the complex spatial scales and orientations found in aerial imagery, leading to suboptimal segmentation results. To address these challenges, we introduce the Rotated Multi-Scale Interaction Network (RMSIN), an innovative approach designed for the unique demands of RRSIS. RMSIN incorporates an Intra-scale Interaction Module (IIM) to effectively address the fine-grained detail required at multiple scales and a Cross-scale Interaction Module (CIM) for integrating these details coherently across the network. Furthermore, RMSIN employs an Adaptive Rotated Convolution (ARC) to account for the diverse orientations of objects, a novel contribution that significantly enhances segmentation accuracy. To assess the efficacy of RMSIN, we have curated an expansive dataset comprising 17,402 image-caption-mask triplets, which is unparalleled in terms of scale and variety. This dataset not only presents the model with a wide range of spatial and rotational scenarios but also establishes a stringent benchmark for the RRSIS task, ensuring a rigorous evaluation of performance. Our experimental evaluations demonstrate the exceptional performance of RMSIN, surpassing existing state-of-the-art models by a significant margin. All datasets and code are made available at https://github.com/Lsan2401/RMSIN.
1. Introduction
RRSIS challenges conventional referring segmentation with aerial imagery’s scale and orientation diversity, while existing datasets and methods are limited. The paper introduces RRSIS-D and RMSIN, combining within- and cross-scale interaction with rotated convolution.
- 1. Introduction: RRSIS-D contains 17,402 image-caption-mask triplets and is described as three times the size of its predecessor, with higher-resolution imagery and broader geographic diversity.Its masks are generated semi-automatically with SAM and refined to improve fidelity to aerial imagery.
- 1. Introduction: Conventional RIS methods struggle with aerial imagery’s broad spatial-scale variation and multiple object orientations, while existing RRSIS datasets are limited in scale and scope.The paper cites land-use categorization, climate-impact studies, and urban-infrastructure management as application areas for RRSIS.
- 1. Introduction: RMSIN combines IIM for within-scale detail, CIM for cross-scale feature fusion, and decoder-based ARC to address aerial-object orientation variation.The paper presents the modules as targeting the scale and orientation challenges of RRSIS.
2. Related work
Prior referring segmentation methods use varied visual-language fusion strategies, but the paper describes them as limited on aerial imagery. It positions RRSIS-D and RMSIN as a response to gaps in dataset complexity and model performance.
- 2. Related work: RIS methods include recurrent refinement, dynamic filters, Transformer cross-modal decoders, and language-aware visual encoding, but the paper reports limited performance in remote sensing.The paper also notes that extreme semantic differences between natural and aerial images remain a challenge.
- 2. Related work: Earlier remote-sensing referring segmentation work introduced a dataset and deep–shallow feature interactions, while the paper says that model faces more complex datasets poorly.The paper compares Yuan et al. ’s model on RRSIS-D.
- 2. Related work: The paper introduces RRSIS-D as a more extensive, intricate dataset and RMSIN as a new model, evaluating the earlier model on the new dataset.
3. RRSIS-D
RRSIS-D uses SAM-assisted mask generation followed by human refinement and curation. Its 17,402 images cover varied categories and mask scales, including very small targets and objects exceeding 400,000 pixels.
- 3. RRSIS-D: The annotation pipeline generates masks from bounding-box prompts with SAM, then manually refines problematic masks and curates the data.SAM accuracy may vary on partial aerial images because of the domain gap between aerial and natural images; annotations are converted to RefCOCO format.
- 3. RRSIS-D: 17,402 images are paired with masks and referring expressions; images are 800 × 800, with 20 semantic categories and 7 attributes.Airplane is the largest category at 15.6% of the dataset.
- 3. RRSIS-D: Mask coverage varies substantially: many targets occupy only a small image fraction, while some objects exceed 400,000 pixels.Figure 4 plots image mask-coverage percentage against total mask count.
- 3. RRSIS-D: The wide range of mask sizes and numerous small targets makes prediction challenging across the dataset.
4. RMSIN
RMSIN combines a multi-stage vision-language encoder with an oriented-aware decoder to predict segmentation masks. Its IIM and CIM handle within-scale and cross-scale feature interactions, while ARC incorporates orientation information into decoding.
- 4.1. Overview: RMSIN transforms image and language inputs into fused multi-scale features through CSIE, then uses an ARC-based decoder to predict a segmentation mask.CSIE contains IIM at each stage and a CIM; the decoder performs parallel inference on features from multiple stages.
- 4.2.1 Intra-scale Interaction Module: IIM extracts information within each scale and facilitates vision-language interaction, using varied receptive-field branches and a visual gate to complement local image details.The module also uses a cross-modal alignment branch to fuse modalities.
- 4.2.2 Cross-scale Interaction Module: CIM combines features from IIM stages for multi-scale interaction, with local-relationship compensation and a scale-aware gate to preserve local details and regulate cross-scale outputs.Its multi-scale attention uses depth-wise convolutions with varied kernel sizes and strides to resize features.
- 4.3. Oriented-aware Decoder: The oriented-aware decoder uses features from CSIE stages to generate masks, addressing the varied orientations of aerial objects with ARC rather than static horizontal kernels.The decoder applies ARC to half of its convolution layers.
- 4.3.1 Adaptive Rotated Convolution: ARC predicts candidate angles and weights from input features, re-parameterizes convolution kernels by rotary resampling, and combines filtered features into orientation-aware features.The kernel weights are interpolated after transforming their sampling coordinates according to the predicted angles.
- 4.3.1 Adaptive Rotated Convolution: For mask prediction, the cited description identifies Seg(·) as a nonlinear block comprising a 3 × 3 convolution layer, a batch normalization layer, and a ReLU, within an overall top-down process.
5. Experiments
RMSIN outperforms LAVT on RRSIS-D, with ablations and visualizations examining its scale-interaction and rotation-aware components. The experiments also report improved predictions across object scales, noisy backgrounds, and orientations.
- 5.1. Implementation Details: Experiments use a Swin Transformer visual backbone and base BERT language backbone, with training for 40 epochs using AdamW.The setup uses four RTX 2080 GPUs, batch size 8, and evaluates oIoU, mIoU, and P@X.
- 5.3. Ablation study: IIM improves precision at lower IoU thresholds, CIM further refines predictions across IoU levels, and their combination achieves the highest performance, with reported margins of 3.5% to 4.5%.The combined gains are particularly noted for P@0.5, P@0.7, and mIoU.
- 5.2. Comparison with state-of-the-art RIS methods: 3.64% and 3.16% mIoU gains over LAVT on validation and test, respectively, accompany RMSIN's gains of over 3.0% in P@0.5, P@0.6, and P@0.7 for small or rotated objects.The comparison uses the RRSIS-D validation and test subsets.
- 5.3. Ablation study: Adding the complete CIM design gives the most substantial enhancement, with a reported metric increase over 4.14%, and is associated with preserving local details and extracting multi-scale information.Table 4 cumulatively reintroduces design components from vanilla self-attention.
- 5.3. Ablation study: The Oriented-aware Decoder, which concatenates features and uses ARC to extract angular information, surpasses the alternative decoder designs across all metrics.The comparisons include summation in place of cross-stage concatenation and static convolutions in place of ARC.
- 5.3. Ablation study: Replacing all three decoder layers with ARC follows a consistent upward performance trend, while increasing predicted angles from 1 to 4 boosts performance by approximately 1%.The angle-count experiment is conducted with L=3.
- 5.4.1 Quantitative Results: Qualitatively, RMSIN identifies targets across large and tiny scales, noisy backgrounds, and varied angles, while LAVT shows missing parts and shifted predicted masks.Figure 6 compares RMSIN with LAVT across these scenarios.
- 5.4.2 Visualization of Features from Encoder: Feature visualizations indicate that CSIE provides more accurate deeper-layer semantics and ARC supplies a spatial prior important for rotated-object segmentation.The visualizations compare outcomes as modules are progressively added.
6. Conclusion
The conclusion presents RMSIN as a method for handling the scales and orientations in RRSIS, alongside RRSIS-D as a large evaluation resource. Validation on that dataset reports superior performance and positions it as a benchmark for future research.
- RMSIN integrates IIM and CIM for spatial-scale variation and ARC for object orientations in aerial imagery.
- RRSIS-D contains 17,402 image-caption-mask triplets and is presented as a resource of substantial scale and variety.
- Validation on RRSIS-D underscores RMSIN's superior performance and establishes the dataset as a benchmark for future RRSIS research.