Source-linked AI summary

Geometry-Driven Opti-Acoustic Co-Registration and View-Invariant Reflectivity Mapping for Side-Scan Sonar

Taqi Hamoda, Nuno Gracias

arXiv:2608.23479v1cs.CV

TL;DR

The paper addresses the difficulty of aligning optical and SSS imagery when acoustic artifacts and viewpoint changes create a physical domain gap. It uses a 3D geometry- and physics-guided pipeline with SfM, FBR correction, and reflectivity isolation to produce deterministic co-registration. The resulting datasets are described as highly accurate and suitable for self-supervised benthic mapping without manual acoustic annotation.

  • Problem

    Acoustic scattering, speckle noise, geometric distortions, and viewpoint dependence make optical-acoustic alignment difficult for existing feature-based methods.

  • Method

    The framework uses a dense SfM seafloor mesh, dynamic First Bottom Return correction, and inverse Lambertian reflectivity mapping with dual-Gaussian weighting.

  • Results

    The pipeline achieves deterministic pixel-level co-registration and isolates view-invariant seabed reflectivity from geometric and slant-dependent acoustic effects.

  • Takeaways & Limitations

    The automated pipeline generates strictly co-registered opti-acoustic datasets without manual acoustic annotation for future self-supervised benthic habitat mapping.

Abstract

from arXiv · show

Side-Scan Sonar (SSS) is a primary modality for large-scale underwater mapping, yet automated perception and cross-modal alignment are severely bottlenecked by acoustic complexities such as speckle noise, shadows, and extreme viewpoint dependencies. Traditional handcrafted descriptors and modern deep learning matchers fail to bridge the physical domain gap between optical and acoustic imagery without 3D geometric constraints. To overcome these limitations, we propose a novel geometry-driven framework for pixel-level opti-acoustic co-registration and view-invariant reflectivity mapping. Our method utilizes Structure-from-Motion (SfM) to reconstruct a dense 3D seafloor mesh, acting as a geometric anchor between the visual and acoustic domains. We introduce a First Bottom Return (FBR) extraction algorithm to dynamically correct non-linear altitude drift caused by uncalibrated SfM reconstruction. Furthermore, we apply an inverse Lambertian model and a dual-Gaussian weighting function to isolate the intrinsic seabed reflectivity, effectively neutralizing slant-range propagation loss and geometric view-dependence. By deterministically associating these isolated acoustic properties with optical pixels, our pipeline generates highly accurate, strictly co-registered multi-modal datasets. This automated, physics-guided approach eliminates the need for manual annotation and paves the way for advanced self-supervised learning in benthic habitat mapping.

I. INTRODUCTION

SSS supports long-range underwater mapping but produces imagery shaped by acoustic acquisition, propagation loss, shadows, speckle, and viewpoint-dependent backscatter. These effects complicate consistent interpretation and cross-modal perception.

  • SSS enables large-scale underwater mapping because acoustic energy travels hundreds of meters in turbid or light-deprived waters.Optical sensors are typically limited to operational ranges under 10 meters because electromagnetic radiation is rapidly absorbed and scattered.
  • Side-looking transducers emit fan-shaped pulses, and echo timing is converted into range using a constant sound speed of 1500 m/s.The resulting measurements form a continuous two-dimensional backscatter intensity map.
  • The Lambertian SSS model expresses recorded intensity as intrinsic seabed reflectivity multiplied by incidence-angle and acoustic-propagation terms.The formulation links I(x, y) to ρ(x, y), cos(θ(x, y)), and L(x, y).
  • Real SSS imagery is distorted by non-uniform transducer power, acoustic shadows, multiplicative speckle noise, and viewpoint-dependent intensity patterns.Different survey directions can produce substantially different intensity patterns and shadow geometries over the same seafloor structure.

A. State of the Art in Feature Extraction and Matching

Optical feature extractors and matchers struggle with acoustic noise and cross-modal physics, while sparse and dense approaches make different coverage, robustness, and computational trade-offs. Existing constraints also depend on features that are invariant across both modalities.

  • Handcrafted detectors such as SIFT, SURF, and AKAZE lose repeatability because acoustic speckle and view-dependent backscatter disrupt local intensity gradients.
  • Deep detectors and Transformer-based matchers improve optical perception but degrade on acoustic wave distortion and stripe noise.The cited examples include SuperPoint, LightGlue, and LoFTR.
  • Sparse Methods: Sparse methods preserve relative accuracy under noise but produce too few correspondences for precise dense mapping.
  • Dense Methods: Dense methods provide broad spatial coverage but remain sensitive to low-contrast acoustic artifacts, with costly computation and low inlier rates.LoFTR is described as relying on pre-trained optical weights and 2D homographic approximations that disregard three-dimensional seafloor topography.
  • Symmetric epipolar constraints require features invariant to both modalities’ underlying physics, a requirement current optical models do not fulfill.The constraint is expressed using a fundamental matrix F ∈ R3×3.

B. Contributions

The paper introduces a physics-informed, geometry-driven framework for opti-acoustic co-registration and reflectivity mapping. Its contributions combine 3D geometric anchoring, dynamic altitude correction, physics-guided reflectivity isolation, and automated dataset generation.

  • A reconstructed 3D optical geometry provides an anchor for deterministic, pixel-level alignment between visual and acoustic data.
  • First Bottom Return extraction dynamically corrects non-linear altitude drift and scale artifacts in uncalibrated SfM reconstructions.
  • An inverse Lambertian model with dual-Gaussian weighting separates intrinsic seabed reflectivity from propagation loss and geometric view-dependence.
  • The scalable pipeline generates co-registered opti-acoustic datasets without manual acoustic annotation, supporting future self-supervised benthic habitat mapping.

II. METHODOLOGY

The methodology establishes pixel-level correspondence between SSS and optical imagery by reconstructing a shared 3D spatial representation. This representation bridges differences in sensor geometry, altitude, and physical modality.

  • The proposed geometry-driven pipeline targets robust pixel-level correspondence between side-scan sonar and optical imagery.
  • The framework uses reconstructed spatial structure to connect otherwise disparate acoustic and optical measurements.
  • A shared 3D spatial representation bridges the disparate geometries, altitudes, and physical modalities of the sensors.

A. Geometric Reconstruction and Calibration

The geometric reconstruction stage uses dense, high-overlap optical surveying and IMU-constrained SfM to build a 3D seafloor mesh. Uncalibrated sensor modeling introduces non-linear altitude offsets through focal-length and radial-distortion effects.

  • Dense, high-overlap lawnmower-survey imagery is processed with Vehicle IMU-constrained SfM to produce a precise 3D seafloor mesh.
  • Joint focal-length optimization can create artificial zooming, causing spatial errors to propagate non-linearly away from the optical center.
  • Unconstrained radial distortion can minimize local reprojection error while failing to model water-column distortions, amplifying altitude offsets.

B. Altitude Offset Correction via First Bottom Return

The altitude-correction stage extracts the First Bottom Return from sonar waterfall imagery and tracks it with a weighted cost function. FBR-derived altitude then dynamically corrects reconstructed geometry using nearby pings.

  • Sonar channels are median-filtered and enhanced with CLAHE before a greedy algorithm tracks the First Bottom Return over a 100-bin sliding window.
  • The tracking cost combines gradient magnitude, distance from nadir, and distance from preceding FBR locations with weights 0.4, 0.3, and 0.3.
  • FBR-derived ground-truth altitude dynamically corrects 3D geometry through an adaptive temporal window containing the 50 closest pings.

C. Physics-Guided Reflectivity Isolation

The pipeline removes acoustic view-dependence and slant-range loss to estimate intrinsic seabed reflectivity, then projects it onto the 3D mesh for co-registered mapping.

  • Inverse Lambertian correction divides aligned SSS intensities by incidence-angle effects and normalizes along-track to remove slant-dependent propagation loss.Shadowed mesh vertices are discarded before valid vertices are mapped to specific SSS bins.
  • A dual-Gaussian weighting function projects relative reflectivity values onto the 3D mesh as a weighted average.The weighting neutralizes sonar emission physics and angular sensitivity.
  • The resulting reflectivity isolates view-invariant material properties from transient acoustic artifacts and supports accurate co-registered data.Figure 3 visualizes relative acoustic reflectivity using OpenCV’s Rainbow colormap.

III. RESULTS AND DISCUSSION

Experiments across rocky and seagrass regions show that local FBR correction aligns SSS imagery with geometric reprojections, while reflectivity mapping removes view-dependent and slant-dependent effects.

  • Evaluation covered diverse benthic environments, specifically rocky and seagrass regions.
  • Local temporal FBR correction precisely aligned raw SSS waterfalls with geometric reprojections and vastly outperformed global offset adjustment.The correction compensated for non-linear scaling artifacts introduced by uncalibrated SfM.
  • Inverse Lambertian processing isolated view-invariant seabed reflectivity while removing geometric view-dependence and slant-dependent propagation loss.Figure 4 presents estimated reflectivity maps for mountainous and seagrass meshes, where brighter areas indicate higher acoustic reflectivity.
  • Isolated reflectivity values accurately represented intrinsic seabed properties and enabled synthesis of SSS intensities for novel-view acoustic simulation.

C. Opti-Acoustic Co-Registration Results

The framework establishes deterministic pixel-level alignment between optical and acoustic data by using dense 3D geometry as a spatial bridge.

  • Localized acoustic bins were deterministically associated with the 3D SfM point cloud to achieve accurate pixel-level co-registration.
  • The resulting alignment was demonstrated across mountainous and seagrass test regions.
  • A dense 3D SfM mesh served as a structural anchor for deterministic mapping between visual and acoustic domains.The geometry-driven, physics-informed approach addresses scattering, speckle noise, and geometric distortions.

A. Future Work

Future work targets more rigorous evaluation, improved acoustic modeling, and downstream self-supervised learning using automatically generated co-registered data.

  • Future work will pursue rigorous quantitative evaluation and framework refinement through advanced modeling, synthetic benchmarking, and self-supervised learning integration.
  • Replacing the simplified Lambertian approximation with advanced scattering models and bathymetric priors may improve fidelity on complex, heterogeneous seabeds.
  • Synthetic data generated from isolated reflectivity values will provide a rigorous benchmark for novel-view acoustic simulation.
  • Automatically generated, perfectly co-registered datasets are planned for training self-supervised learning architectures for automated benthic habitat mapping.This direction is intended to bypass manual acoustic annotation.
Loading 2608.23479v1…