Source-linked AI summary

ORSIm Detector: A Novel Object Detection Framework in Optical Remote Sensing Imagery Using Spatial-Frequency Channel Features

Xin Wu, Danfeng Hong, Jiaojiao Tian, Jocelyn Chanussot, Wei Li, Ran Tao

arXiv:1901.07925v2cs.CV

TL;DR

Object detection in optical remote sensing imagery is limited by incomplete representations of rotation, translation, and scale variation, especially under restricted labeled data. ORSIm detector combines spatial-frequency channel features, feature refinement, fast image-pyramid scaling, and boosting-based classification. Experiments on two datasets report better and more robust performance than previous state-of-the-art methods.

  • Problem

    Existing detection approaches inadequately represent complex deformations across rotation, translation, and scale, while richer learned representations require substantial labeled training data.

  • Method

    ORSIm detector integrates spatial-frequency channel features, learning-based feature refinement, fast image-domain pyramid scaling, and AdaBoost within a VJ-style detection framework.

  • Results

    Experiments on two optical remote sensing datasets indicate that ORSIm detector performs better and is more robust to various deformations than previous state-of-the-art methods.

  • Takeaways & Limitations

    The framework provides a more general detection approach for remote sensing imagery with rotation-invariant features and multiscale processing.

Abstract

from arXiv · show

With the rapid development of spaceborne imaging techniques, object detection in optical remote sensing imagery has drawn much attention in recent decades. While many advanced works have been developed with powerful learning algorithms, the incomplete feature representation still cannot meet the demand for effectively and efficiently handling image deformations, particularly objective scaling and rotation. To this end, we propose a novel object detection framework, called optical remote sensing imagery detector (ORSIm detector), integrating diverse channel features extraction, feature learning, fast image pyramid matching, and boosting strategy. ORSIm detector adopts a novel spatial-frequency channel feature (SFCF) by jointly considering the rotation-invariant channel features constructed in frequency domain and the original spatial channel features (e.g., color channel, gradient magnitude). Subsequently, we refine SFCF using learning-based strategy in order to obtain the high-level or semantically meaningful features. In the test phase, we achieve a fast and coarsely-scaled channel computation by mathematically estimating a scaling factor in the image domain. Extensive experimental results conducted on the two different airborne datasets are performed to demonstrate the superiority and effectiveness in comparison with previous state-of-the-art methods.

I. INTRODUCTION

Optical remote sensing detection remains challenged by incomplete feature representations for rotation, translation, and scale changes. ORSIm detector addresses these issues by combining spatial-frequency features, feature refinement, fast pyramid scaling, and boosting-based detection.

  • I. INTRODUCTION: Existing methods often fail to represent object features completely across deformation types and densely sampled scales.The desired feature space should be robust to shifts and rotations while capturing substantial patterns with a coarse image pyramid.
  • I. INTRODUCTION: Remote sensing imagery introduces complex rotation behavior, while rotation augmentation remains limited by preset angles and large labeled-data requirements.Fractional rotation angles can therefore remain difficult for learning-based methods to address adaptively.
  • I. INTRODUCTION: ORSIm detector extends the VJ framework with spatial-frequency channel features, feature refinement, fast image-pyramid estimation, and AdaBoost learning.Its SFCF combines rotation-invariant frequency-domain features with spatial features such as color and gradient magnitude.
  • I. INTRODUCTION: SFCF jointly models rotation and shift invariance to address complex deformation behavior in remote sensing imagery.The frequency-domain representation is complemented by shift-invariant spatial channel features.
  • I. INTRODUCTION: The framework estimates image-domain scaling factors to generate a fast image pyramid for multiscale object detection.This enables coarsely sampled pyramid processing without sacrificing the stated detection objective.

II. RELATED WORK

Related work spans channel-feature extraction, rotation-invariant descriptors, multiresolution pyramids, and boosting for remote-sensing detection. ORSIm addresses limitations in feature diversity and scale computation within this broader context.

  • A. Channel Features: Channel features transform input images into spatially discriminative representations and have been applied to geospatial object detection.
  • A. Channel Features: FourierHOG models rotation-invariant descriptors in a continuous frequency domain using Fourier-based convolutional transformations.
  • A. Channel Features: ORSIm extends FourierHOG by combining spatial and frequency channel features, addressing FourierHOG’s limited feature diversity.
  • B. Feature Channel Scaling: Traditional multiscale methods construct a feature pyramid at every sampled scale, making computation time-consuming.
  • B. Feature Channel Scaling: Figure 4 contrasts training feature-point extraction with feature-channel scaling.
  • B. Feature Channel Scaling: Feature channel scaling estimates a scale factor to compute finely sampled pyramids at a fraction of the cost.

C. Boosting Decision Tree

The detector uses boosting within the ORSIm pipeline, combining extracted spatial and frequency features into region-based representations before pooling and model output.

  • C. Boosting Decision Tree: Boosting methods iteratively select weak learners to address hard examples from previous rounds.
  • C. Boosting Decision Tree: Algorithm 1 takes training images and parameters as input and outputs a detector and detection results.
  • C. Boosting Decision Tree: The pipeline extracts pixel-wise spatial channel features from Eqs. (2–3).
  • C. Boosting Decision Tree: It extracts pixel-wise frequency channel features and computes region-based SFCF representations.
  • C. Boosting Decision Tree: A pooling-like operation then produces aggregate channel features for subsequent processing.

18 Step 4: Test Phase with Feature Channel Scaling

The test-phase pipeline uses spatial-frequency channels, learned feature representations, and scale estimation to process feature pyramids efficiently. Its SFCF combines spatial channels with rotation-invariant frequency channels and region-based aggregation.

  • Test Phase: The test procedure obtains feature pyramids at different scales and feeds them into the learned detector.
  • Feature Extraction: ORSIm extracts spatial-frequency channels, refines them through subspace learning or ACF, and feeds them into a boosting decision tree.
  • Feature Extraction: SFCF combines RGB, gradient-magnitude, and rotation-invariant channels from spatial and frequency domains.
  • Frequency Features: Fourier-domain magnitude channels provide rotation-invariant features, while equal-and-opposite Fourier bases remove rotation information from phase.
  • Frequency Features: Relative phase information is represented by coupling convolutional features from two neighboring kernel radii.
  • Region-Based Features: Region-based channel features use triangular convolution kernels of different sizes in spatial and frequency domains to capture semantic context.

B. Feature Learning or Refine

ORSIm refines the feature cube along spatial and channel dimensions using subspace learning or aggregation-based pooling. These strategies target computational efficiency, representation quality, and adaptive support regions.

  • Feature Learning or Refine: Feature learning or refinement acts along spatial and channel directions to reduce the gap between spatial and frequency domains.
  • Feature Learning or Refine: PCA-based subspace learning reduces computational and storage costs while improving feature representation to some extent.
  • Feature Learning or Refine: ACF pooling dynamically adjusts support-region sizes while maintaining structural consistency with the overall image.

C. Training Phase with Ensemble Classifier Learning

The training phase uses softcascade boosting with depth-3 decision trees to combine weak classifiers and improve discrimination between intra- and inter-samples.

  • Softcascade boosting with depth-3 decision trees is applied to improve discrimination between intra- and inter-samples.
  • Boosting combines many weak learners into a stronger classifier by minimizing training errors while controlling test errors.

D. Test Phase with Feature Channel Scaling

The test phase avoids the heavy cost of finely sampled image pyramids by estimating channel scaling factors and reusing them across pyramid images.

  • The framework replaces computationally expensive finely sampled image pyramids with a fast model that estimates feature-channel scaling factors.
  • For input image I and resampling factor s, λ is estimated during training so channel features across pyramid images can be computed quickly.R(I, s) denotes the image resampled by s.
  • Channel images at different scales are obtained through linear or nonlinear transformations in spatial and frequency domains.

IV. EXPERIMENTAL RESULTS AND ANALYSIS

The experiments evaluate ORSIm on satellite-car and NWPU VHR-airplane datasets using train-test splits and cross-validation, with PR curves visualizing comparative performance.

  • Two airborne datasets evaluate ORSIm: satellite car targets and NWPU VHR-airplane targets, using 60% training samples and 40% testing samples.
  • Figure 7 presents PR curves comparing ORSIm with state-of-the-art approaches.
  • Five-fold cross-validation reports average results across folds to assess generalization with limited labeled remote-sensing data.
  • The satellite dataset contains 1319 manually labeled cars from 30 low-resolution images with varying illumination and building shadows.
  • The NWPU VHR-airplane dataset includes 650 positive airplane images and 150 negative images acquired from Google Earth and Vaihingen data.

B. Experimental Setup

The experimental setup combines diverse spatial and frequency channel features, multiscale sampling, AdaBoost classification, and several detection-performance criteria.

  • SFCF extraction: SFCF combines spatial color and gradient-magnitude channels with rotation-invariant frequency channels.Candidate color spaces include RGB, LUV, and HSV.
  • Feature pyramid: The framework samples objects at multiple scales and applies a fast feature pyramid to accelerate hard-negative mining and testing.The four stated scales are s=1, 2, 4, and 8.
  • Classifier setting: AdaBoost trains the classifier through weighted majority voting while increasing weak learners from 32 to 2048 to reduce over-fitting.
  • Evaluation criteria: Performance is assessed with PR curves, AP, AR, and AF, using a 50% bounding-box overlap threshold for true positives.
  • Classifier selection: Three classifiers—linear SVM, RF, and AdaBoost—are compared across HOG, ACF, FourierHOG, and SFCF descriptors.
  • Classifier selection: AdaBoost adaptively reweights weak classifiers, whereas RF weights sub-classifiers equally, making AdaBoost more suitable for the current dataset according to the authors.

2) Overview of performance comparison:

ORSIm detector is evaluated against multiple object-detection methods on two remote-sensing datasets using precision, recall, average F-measure, and runtime comparisons. It achieves the highest detection precision, while fast feature pyramids improve speed and two-step NMS addresses overlap errors.

  • ORSIm detector largely outperforms the investigated methods on both datasets, with dramatically higher precision attributed to SFCF and AdaBoost.The comparison includes eight methods and reports AP, AR, AF, and mean running times.
  • Methods using fast feature pyramids enable faster detection than methods without them.
  • The proposed detector achieves the highest detection precision despite being slower than ACF and YOLO2.
  • Visual errors include roofs misidentified as cars and missed transport cars, possibly due to limited training samples, class imbalance, and weak visible edges.
  • Two-step NMS fixes false detections caused by confusing airplane tails or similar colored regions.The improved results are shown in Fig. 9(b).

D. Sensitivity Analysis

Sensitivity analysis examines parameter configurations and image resolution across the two datasets. Performance favors LUV-based channel combinations, remains relatively stable across proper parameter ranges, and begins degrading near a 0.5 sampling rate.

  • Towards Parameter Setting: LUV color space performs better on both datasets, especially when combined with color channels and gradient magnitude.A similar trend remains after adding rotation-invariant feature channels.
  • Towards Parameter Setting: Radial profiles, Fourier orders, and sampling-window size are relatively insensitive within a proper range.The selected Fourier order is m = 4 for both datasets.
  • Towards Parameter Setting: Eight scales per octave produces the best result in the tested pyramid-factor configurations.Increasing the number of weak classifiers raises precision but also increases computational cost; the selected value is 2048.
  • Image Resolution: Detection performance may begin to degenerate around a 0.5 image sampling rate and gradually decrease afterward.
  • Overall conclusion: ORSIm detector is reported to perform better and remain more robust to various deformations than previous state-of-the-art methods.
Loading 1901.07925v2…