Source-linked AI summary

SeaDronesSee: A Maritime Benchmark for Detecting Humans in Open Water

Leon Amadeus Varga, Benjamin Kiefer, Martin Messmer, Andreas Zell

arXiv:2105.01922v2cs.CV

TL;DR

SeaDronesSee addresses the lack of large-scale maritime UAV data and environmental metadata needed for detecting and tracking people in open water. It introduces a benchmark with detection and tracking tasks, evaluates state-of-the-art systems, and shows that metadata can improve detector accuracy.

  • Problem

    Large-scale maritime datasets and detailed environmental metadata are lacking, although robust vision systems are needed to detect, localize, and track people in open water.

  • Method

    The paper introduces SeaDronesSee, a large-scale open-water dataset with precise frame-level metadata, three detection and tracking tasks, baseline evaluations, and a central evaluation server.

  • Results

    3.1 APavg: 5×Angle@3 outperforms its ResNet-50-FPN baseline by 3.1 APavg, while extensive experiments establish state-of-the-art detector and tracker baselines.

  • Takeaways & Limitations

    The benchmark provides a foundation for maritime UAV computer vision and aims to accelerate metadata-aware detection and tracking research.

  • Takeaways & Limitations

    The evaluated systems are not capable of running in real time on embedded hardware, an important use case for UAV-based search and rescue missions.

Abstract

from arXiv · show

Unmanned Aerial Vehicles (UAVs) are of crucial importance in search and rescue missions in maritime environments due to their flexible and fast operation capabilities. Modern computer vision algorithms are of great interest in aiding such missions. However, they are dependent on large amounts of real-case training data from UAVs, which is only available for traffic scenarios on land. Moreover, current object detection and tracking data sets only provide limited environmental information or none at all, neglecting a valuable source of information. Therefore, this paper introduces a large-scaled visual object detection and tracking benchmark (SeaDronesSee) aiming to bridge the gap from land-based vision systems to sea-based ones. We collect and annotate over 54,000 frames with 400,000 instances captured from various altitudes and viewing angles ranging from 5 to 260 meters and 0 to 90 degrees while providing the respective meta information for altitude, viewing angle and other meta data. We evaluate multiple state-of-the-art computer vision algorithms on this newly established benchmark serving as baselines. We provide an evaluation server where researchers can upload their prediction and compare their results on a central leaderboard

1. Introduction

SeaDronesSee addresses the shortage of large-scale UAV imagery for maritime search and rescue by introducing an annotated open-water benchmark with environmental metadata, multispectral imagery, and evaluation infrastructure.

  • Motivation: Maritime SAR vision systems need robust detection and tracking because people are small relative to search areas and viewing conditions vary.UAV imagery can include changing altitudes and viewing angles, while maritime scenes add reflections, shadows, waves, and sea foam.
  • Motivation: Large-scale maritime UAV datasets are scarce, while existing land-based or satellite datasets lack the resolution or perspective needed for SAR missions.Common UAV datasets focus on traffic, whereas many maritime datasets use satellite radar imagery intended for ship detection.
  • Dataset: SeaDronesSee provides high-resolution open-water RGB imagery with bounding-box annotations for swimmers, floaters, life jackets, people on boats, and boats.The imagery was captured at resolutions from 3840×2160 px to 5456×3632 px.
  • Dataset: The dataset includes precise frame-level metadata such as altitude, camera angle, speed, and time to support metadata-aware vision systems.The authors identify limited metadata in existing UAV datasets as an impediment to developing multimodal methods.
  • Dataset: SeaDronesSee additionally provides multispectral imagery containing Near Infrared at 842 nm and Red Edge at 717 nm.These channels are intended to support detectors using nonvisible light spectra for human detection in maritime settings.
  • Benchmark: The benchmark covers object detection, single-object tracking, and multi-object tracking, with state-of-the-art experiments, baselines, and an evaluation server for leaderboard comparisons.Training and validation annotations are released, while test predictions are evaluated through the server.

2. Related Work

Existing aerial datasets predominantly cover land-based traffic, while maritime datasets largely emphasize satellite-based ship detection and provide limited environmental metadata. SeaDronesSee addresses this gap with annotated maritime imagery designed for object detection and tracking.

  • UAV datasets such as VisDrone and UAVDT primarily depict traffic scenarios and support object detection or tracking.
  • Table 1 compares annotated aerial datasets by object content, altitude and angle information, and other metadata.Some altitude and angle values are estimates, and satellite datasets lack available altitude ranges.
  • Maritime datasets commonly use satellite-based synthetic aperture radar and focus on ships rather than people in open water.These datasets lack the resolution needed for search-and-rescue person detection and provide only top-down views.
  • UAVDT provides coarse altitude, viewing-angle, and lighting metadata, whereas other datasets vary in precision and coverage.Mid-Air provides precise metadata but lacks annotated objects, while Bozcan excludes camera-angle information.
  • Existing tracking datasets often describe clip-level attributes rather than frame-by-frame environmental conditions.Examples include aspect-ratio change, background clutter, and fast motion.

3. Data Set Generation

SeaDronesSee was generated from multi-UAV maritime recordings with varied cameras, multispectral sensing, precise frame metadata, manual expert-checked labels, and controlled dataset splits. Test annotations are withheld for server-based evaluation.

  • Footage was collected over several days from open-water subjects using quadcopters and a fixed-wing UAV under a planned flight schedule.More than 20 test subjects participated, with safety procedures including boat transport and vertical separation.
  • Multiple cameras mounted on four UAV platforms were used to reduce camera bias, with video recorded at 30 fps.For object detection, at most three frames per second were extracted to reduce redundant frames.
  • Multispectral recordings contain RGB, RedEdge, and near-infrared channels across five wavelengths.The RedEdge-MX captured top-down imagery at 1 fps, and bands were merged using the manufacturer’s development kit.
  • Every frame includes metadata aligned to video using nearest-neighbor timestamp matching.Metadata is logged at 10 hertz from onboard clock, barometer, IMU, GPS, and gimbal sensors.
  • Metadata values remain within the error thresholds of the different sensors, but extended sensor-error analysis is outside the paper’s scope.
  • Annotations cover swimmers, floaters, life jackets, people on boats, and boats, with expert checks after manual labeling.The floater class cannot be reliably inferred solely from swimmer location and life-jacket annotations.
  • Training, validation, and testing subsets were balanced by altitude and viewing-angle statistics, using a 4/7, 1/7, and 2/7 split.Video clips were split similarly, while test annotations were withheld for evaluation-server scoring.
  • Objects appear blurry in the example crops because they are small and captured from high altitudes despite originating from high-resolution images.

4. Data Set Tasks

SeaDronesSee supports object detection, single-object tracking, and multi-object tracking in challenging maritime UAV imagery. The tasks span varied cameras, classes, altitudes, viewing angles, and RGB or multispectral data.

  • The benchmark defines three tasks: object detection, single-object tracking, and multi-object tracking.These tasks isolate finding and following people or other objects in maritime UAV footage.
  • Maritime UAV imagery contains reflections, shadows, wave or foam occlusions, and particularly small objects across large search areas.
  • Object Detection: Object detection uses 5,630 images split into 2,975 training, 859 validation, and 1,796 testing images.RGB channels are used for the object detection task, including RGB channels from multispectral images.
  • Object Detection: The object distribution is slightly skewed toward boats because safety precautions required boats near the subjects.Swimmers with life jackets are the next most common objects and may be easier to detect because their jackets often have contrasting colors.
  • Object Detection: Figure 3 encodes training-image distributions by camera type and object distributions by class.
  • Object Detection: Figure 4 encodes image distributions by capture altitude and viewing angle, with every altitude and angle interval sufficiently represented.Roughly 50% of images were recorded below 50 m, while higher-altitude images more often use a 90° downward view.
  • Single-Object Tracking: Single-object tracking provides 208 short clips totaling 393,295 frames and excludes long-term tracking.The sequences are split into 58 training, 70 validation, and 80 testing clips.
  • Multi-Object Tracking: Multi-object tracking provides 22 clips with 54,105 frames and 403,192 annotated instances.MOT-Swimmer tracks floaters and swimmers, while MOT-All-Objects-In-Water also includes boats.

5. Evaluations

The evaluation establishes baseline performance for object detection and tracking on SeaDronesSee, including robustness across altitude and viewing-angle domains and experiments using metadata.

  • Evaluation setup: The benchmark evaluates object detection, single-object tracking, and multi-object tracking with established task-specific protocols and baseline models.Detection uses COCO-style AP and AR metrics; tracking evaluates success, precision, and MOT measures.
  • Object Detection: Large Faster R-CNN with a ResNeXt-101 64-4d backbone performs best overall, while faster models perform worse but can support real-time inference.CenterNet-Hourglass104 closely follows the best detector, while CenterNet-ResNet18 and EfficientDet-D0 are faster alternatives.
  • Object Detection: 54.7 AP50 for ResNeXt-101-FPN exceeds Hourglass104’s 50.3 overall, but Hourglass104 is stronger in the medium-angle domain.The comparison illustrates that aggregate performance does not fully characterize domain-specific robustness.
  • Single-Object Tracking: PrDiMP- and DiMP-family trackers outperform Atom in success and precision, but SeaDronesSee performance is similar to or worse than on UAV123.DiMP18 and Atom run at approximately 27.1 fps on an Nvidia RTX 2080 Ti, yet the faster trackers are not real-time on embedded hardware.
  • Multi-Object Tracking: Tracktor++ performs better than FairMOT in both multi-object tracking tasks, possibly because it uses a Faster R-CNN with a ResNet50 backbone.FairMOT uses CenterNet with DLA34 or ResNet34 backbones in the reported experiments.
  • Meta-Data-Aware Object Detector: 3.1 APavg50 is the gain of 5×Angle@3 over its ResNet-50-FPN baseline at the same inference speed.The gains reach +9.2 APavg50 for acute angles and +6.4 APavg50 for medium angles, which are underrepresented domains.

6. Conclusions

SeaDronesSee is an introductory maritime UAV benchmark for detecting and tracking humans in open water. It combines three evaluation challenges with environmental metadata and multispectral imagery to support future metadata-aware and multimodal systems.

  • Conclusions: SeaDronesSee provides a large-scale benchmark for detecting and tracking humans in open water using UAV imagery.The work describes it as the first large-scale data set of this kind.
  • Conclusions: The benchmark supplies full environmental information for every frame, enabling research on metadata-aware and multimodal object detection and tracking.The stated metadata scope creates opportunities in an area the paper characterizes as previously restricted.
  • Conclusions: Three challenges—object detection, single-object tracking, and multi-object tracking—are supported through an evaluation server.The server is intended to facilitate benchmark comparisons.
  • Conclusions: Multispectral imagery includes wavelengths that can distinguish objects from the water background in maritime scenarios.The paper presents these images as promising for detecting humans in open water.
Loading 2105.01922v2…