Source-linked AI summary

SPEED+: Next-Generation Dataset for Spacecraft Pose Estimation across Domain Gap

Tae Ha Park, Marcus Märtens, Gurvan Lecuyer, Dario Izzo, Simone D'Amico

arXiv:2110.03101v2cs.CVcs.LG

TL;DR

Spaceborne pose-estimation research lacks large, accurately labeled imagery that represents the visual and illumination variability of target space environments. SPEED+ addresses this gap with labeled synthetic training data and diverse TRON HIL test imagery, and uses the dataset to compare model robustness across the synthetic-to-HIL domain gap. The paper also identifies current scope boundaries in target coverage and facility accuracy.

  • Problem

    Large-scale accurately labeled target imagery is impractical to acquire in space, while existing datasets rely heavily on synthetic images that do not represent spaceborne visual and illumination variability.

  • Method

    SPEED+ combines large-scale labeled synthetic imagery with two unlabeled TRON HIL imagery domains and applies existing CNN, domain-adaptation, and randomization methods for robustness studies.

  • Results

    SPEED+ provides a dataset and benchmark for characterizing the domain gap and comparing robustness of spaceborne pose-estimation models trained on synthetic images.

  • Takeaways & Limitations

    SPEED+ enables on-ground robustness evaluation using high-fidelity HIL imagery as a surrogate domain for assessing models intended for spaceborne navigation.

  • Takeaways & Limitations

    The released dataset contains images of only one known target, and facility calibration can produce centimeter-level position errors beyond close range.

Abstract

from arXiv · show

Autonomous vision-based spaceborne navigation is an enabling technology for future on-orbit servicing and space logistics missions. While computer vision in general has benefited from Machine Learning (ML), training and validating spaceborne ML models are extremely challenging due to the impracticality of acquiring a large-scale labeled dataset of images of the intended target in the space environment. Existing datasets, such as Spacecraft PosE Estimation Dataset (SPEED), have so far mostly relied on synthetic images for both training and validation, which are easy to mass-produce but fail to resemble the visual features and illumination variability inherent to the target spaceborne images. In order to bridge the gap between the current practices and the intended applications in future space missions, this paper introduces SPEED+: the next generation spacecraft pose estimation dataset with specific emphasis on domain gap. In addition to 60,000 synthetic images for training, SPEED+ includes 9,531 hardware-in-the-loop images of a spacecraft mockup model captured from the Testbed for Rendezvous and Optical Navigation (TRON) facility. TRON is a first-of-a-kind robotic testbed capable of capturing an arbitrary number of target images with accurate and maximally diverse pose labels and high-fidelity spaceborne illumination conditions. SPEED+ is used in the second international Satellite Pose Estimation Challenge co-hosted by SLAB and the Advanced Concepts Team of the European Space Agency to evaluate and compare the robustness of spaceborne ML models trained on synthetic images.

1. INTRODUCTION

Spaceborne pose-estimation models lack large, accurately labeled, representative imagery for training and validation. SPEED+ addresses this gap with diverse HIL imagery, synthetic training data, and robustness studies focused on the synthetic-to-space domain gap.

  • Motivation: Large-scale labeled spaceborne imagery is impractical to acquire, limiting CNN training for autonomous navigation around noncooperative targets.Accurate pose determination is needed for safe docking and capture, while monocular cameras are favored because of spacecraft power and computational constraints.
  • Motivation: Existing spaceborne images lack sufficient quantity, pose diversity, environmental variability, and accurate labels for comprehensive robustness evaluation.Prior evaluations were therefore often qualitative or based on hand-labeled annotations.
  • Approach: Hardware-in-the-loop robotic testbeds provide an alternative by generating many target images with accurate pose labels in a domain different from synthetic training imagery.The approach physically recreates aspects of the space environment using a target mockup and robotic testbed.
  • Prior benchmark: SPEED, the prior publicly available multi-source benchmark, contained 15,300 Tango images, including 15,000 synthetic and 300 simulated TRON images.The restricted pose and illumination configurations of the simulated images limited comprehensive robustness analysis across domains.
  • Contribution: SPEED+ combines large-scale labeled synthetic imagery for training with two distinct unlabeled HIL domains for testing, emphasizing the synthetic-to-HIL domain gap.The dataset was introduced for robustness studies and made available through the second Satellite Pose Estimation Competition.
  • Contribution: SPEED+ test images cover full orientation and distances up to 10 m under realistic Earth-albedo and direct-sunlight conditions.The dataset supports comprehensive evaluation of robustness across high-fidelity simulated environmental settings and is used in SPEC2021.

2. RELATED WORK

Prior spacecraft-vision datasets and testbeds provide useful imagery or maneuver simulation but do not fully combine diverse, accurately labeled, large-scale data with spaceborne domain-gap evaluation. SPEED+ is positioned against terrestrial domain-gap benchmarks whose visual challenges differ from those of spaceborne imagery.

  • Spacecraft datasets: Existing spacecraft datasets include synthetic and spaceborne imagery, but some lack spaceborne labels or provide only bounding-box and segmentation annotations.Examples include URSO and a dataset of random satellites with approximately 3,000 images.
  • Spaceborne testbeds: Air-bearing navigation testbeds simulate maneuver commands, but their large-scale annotated-dataset capabilities had not been demonstrated.A separate GRASL effort generated 100 HIL images, but its static tripod restricted configurable pose distribution.
  • Domain-gap benchmarks: Terrestrial domain-gap datasets support classification, detection, and segmentation, but their imagery does not contain the distinctive visual challenges of spaceborne scenes.The paper contrasts synthetic street imagery and real driving datasets with the conditions encountered in spaceborne navigation.

3. SPEED+ OVERVIEW

SPEED+ combines a labeled synthetic training domain with two unlabeled TRON HIL testing domains that differ in illumination. The dataset is designed for robustness comparisons across the synthetic-to-HIL domain gap and emphasizes realistic space-environment image formation.

  • Dataset composition: SPEED+ contains 59,960 labeled synthetic images split 80:20 for training and validation, plus lightbox and sunlamp HIL domains generated by TRON.The lightbox uses diffuser-based albedo simulation, while the sunlamp represents direct high-intensity illumination; HIL domains are mainly intended for testing.
  • Dataset composition: Compared with SPEED’s approximately 300 HIL images, SPEED+ provides a much larger quantity of HIL imagery for robustness evaluation.The comparison concerns the predecessor’s simulated images and SPEED+’s two HIL domains.
  • Availability and use: SPEED+ is publicly released through SPEC2021 so researchers can develop and compare robust pose-estimation models using labeled synthetic and unlabeled HIL data.The dataset is intended for comparison by the aerospace community and others.
  • Realism: The HIL imagery uses a CAD-based spacecraft mockup with emulated critical features, including antennae, brackets, external details, and printed solar cells.The mockup details are manufactured with stated tolerances and materials intended to reproduce relevant spacecraft appearance.
  • Realism: TRON’s calibrated light boxes and sun lamp are designed to reproduce space-environment illumination, including Earth-albedo radiance and direct sunlight.The light boxes are calibrated to a maximum radiance of 14 W/m2sr corresponding to the stated Earth-albedo and solar-irradiance conditions.

4. GENERATING SPEED+ IMAGES

SPEED+ HIL imagery is collected in TRON using controlled robot motion, diverse spaceborne illumination, and pose labels spanning broad position and orientation spaces. Post-processing masks facility components, preserves relevant sunlamp effects, and removes samples with labeling or masking problems.

  • Facility and data collection: TRON combines two 6-DOF robot arms, independent Vicon tracking, and KUKA telemetry to collect accurately labeled spacecraft imagery.One arm holds the Tango mockup while the other carries the camera along a ceiling-mounted rail.
  • Facility and data collection: 10 Earth-albedo lightboxes and a three-position metal-halide sun lamp recreate varied spaceborne illumination conditions.The lightboxes and lamp positions are configured to illuminate the half-scale Tango spacecraft mockup from different directions.
  • Pose and illumination sampling: 10,000 prescribed poses are generated under operational constraints, with target orientations sampled uniformly in SO(3).The target remains at a constant facility location, the camera boresight points toward it, and labels are partitioned across seven lightbox and three sunlamp configurations.
  • Pose and illumination sampling: HIL position and orientation labels are well sampled across 3D Cartesian and SO(3) spaces, while sunlamp incidence covers all target-facing angles.This broad sampling improves coverage relative to the restricted pose distribution of earlier SPEED real images.
  • Image post-processing: Lightbox images replace masked facility backgrounds with random starfields or Earth imagery, whereas sunlamp processing retains challenging flare and glow effects.Sunlamp glow can obscure target boundaries and act as occlusion noise, so applying the lightbox masking procedure unchanged would remove relevant visual effects.
  • Image post-processing: Post-processing leaves 6,740 lightbox and 2,791 sunlamp images after rejecting misaligned masks, unrecognizable reflections, and glow-processing failures.Facility artifacts such as infrared markers, mounting holes, and some sunlamp glow shapes remain in the released HIL images.

5. EXPERIMENTS

The experiments characterize the synthetic-to-HIL domain gap using baseline pose-estimation CNNs, domain-bridging methods, and oracle comparisons. Synthetic-only models degrade on HIL images, especially under sunlamp illumination, while straightforward adaptation and randomization leave substantial challenges.

  • Baseline performance: The study compares KRN, SPN, and HigherHRNet on synthetic validation, lightbox HIL, and sunlamp HIL images.KRN and SPN are SPEED baselines, while HigherHRNet represents heatmap-based pose estimation.
  • Experimental setup: Ground-truth bounding boxes replace object detection during preprocessing to simplify evaluation of the 6D pose-estimation models.The study assumes perfect cropping boxes derived from pose labels.
  • Domain bridging: Unsupervised domain adaptation and domain randomization represent distinct approaches to reducing discrepancies between synthetic training images and HIL test images.The study applies DANN and style augmentation to KRN.
  • Baseline performance: Synthetic-only testing reveals a significant domain gap, with especially severe degradation on sunlamp images affected by high contrast and camera overexposure.Lightbox performance also suffers from texture differences between computer graphics and the physical spacecraft model.
  • Baseline performance: HigherHRNet exhibits a much smaller performance drop than KRN and SPN, suggesting greater robustness from heatmap prediction than attitude classification or direct keypoint regression.The passage presents this as an indication rather than a definitive causal result.
  • Domain bridging: Oracle results support the accuracy and learnability of HIL pose labels, whereas DANN and style augmentation show that bridging the synthetic-to-HIL gap remains difficult.The authors state that more sophisticated algorithms and hyperparameter tuning are needed beyond straightforward application of existing approaches.
  • Flight-image comparison: KRN score distributions on SPEED+ HIL and PRISMA flight images show similar trends, supporting HIL images as surrogates for robustness analysis against the synthetic domain.The comparison follows five synthetic-only training sessions with different random seeds.

6. DISCUSSIONS

The discussion identifies current scope and evaluation boundaries for SPEED+. The released dataset contains one known target, and SPEC2021 relaxes an operational constraint by making unlabeled target images publicly available.

  • Dataset limitations: The current SPEED+ release contains images of a single known target, although future versions are planned to include more targets.The TRON facility can generate many labeled HIL images for a mockup with limited human intervention.
  • Dataset limitations: A planned KUKA robot-arm update is intended to improve pose-label accuracy and reduce human intervention during sample rejection.The update is described as pending rather than completed.
  • Operational scope: SPEC2021 makes unlabeled lightbox and sunlamp images publicly available even though real servicers typically lack target images before rendezvous.Participants may use or ignore these images when training robust models.
  • Study scope: The baseline studies characterize SPEED+ and its suitability for robustness research rather than identify the best CNN architecture or training algorithm.A broader analysis of factors contributing to robust models is deferred until after the competition.

7. CONCLUSIONS

SPEED+ is presented as a dataset for evaluating spacecraft pose-estimation and navigation models across a synthetic-to-spaceborne-like domain gap. The authors frame it as a step toward ML-supported autonomous proximity operations.

  • Dataset contribution: SPEED+ focuses on the domain gap between synthetic training images and test images captured by a robotic testbed under spaceborne-like illumination.The dataset is characterized using existing pose-estimation CNNs and established domain-adaptation and randomization algorithms.
  • Evaluation purpose: The dataset supports robustness comparisons for ML models through the international Satellite Pose Estimation Competition.The competition is identified as the application context for evaluating different models.
  • Broader significance: The authors position SPEED+ as an important step toward autonomous proximity operations such as refueling defunct space assets and active debris removal.This consequence is framed within future mission concepts rather than demonstrated operational capability.

A. ACCESSIBILITY

SPEED+ is distributed through public competition and repository infrastructure. The dataset is accessible through ESA’s challenge platform and Stanford’s Digital Repository under a specified Creative Commons license.

  • Competition access: SPEC2021 is hosted on ESA’s Kelvins platform, which provides the dataset, structure description, leaderboard, and public discussion page.The platform also hosted the first SPEC and other space-related challenges.
  • Dataset access: The dataset is publicly released through the Stanford Digital Repository with a unique DOI and an additional Zenodo link via SPEC2021.These are the stated distribution channels.
  • Dataset access: SPEED+ is licensed under Creative Commons Attribution-NonCommercial-ShareAlike 4.0.The license is specified as CC BY-NC-SA 4.0.

B. CAMERA MODEL

The SPEED+ HIL camera uses a calibrated Point Grey Grasshopper 3 with a Xenoplan 1.4/17mm lens. Its radial and tangential distortion parameters follow OpenCV conventions and are provided in camera.json.

  • The HIL images use a Point Grey Grasshopper 3 camera with a Xenoplan 1.4/17mm lens.The camera was calibrated before use, with calibrated parameters reported in Table 5.
  • Radial and tangential distortion parameters follow OpenCV conventions and are provided in camera.json.

C. TRAINING DETAILS

The baseline studies use PyTorch and GPU training, with method-specific optimization, architecture, input, and fallback settings. The pipeline also retains characteristic HIL artifacts while rejecting unusable sunlamp samples.

  • Training setup: All baseline methods are implemented with PyTorch v1.8.0 and trained on an NVIDIA GeForce RTX 2080 Ti 12GB GPU.
  • KRN: KRN uses AdamW with β1 = 0.9 and β2 = 0.999, training synthetics for 50 epochs with an initial learning rate of 0.001 and 0.95 decay.
  • SPN: SPN uses the first five layers of ImageNet-pretrained AlexNet and trains its backbone and Branches 2 and 3 simultaneously.
  • Keypoint pipeline: The keypoint pipeline resizes inputs to 480 × 320, selects maximum-intensity heatmap pixels, and applies EPnP for pose estimation.
  • HIL image processing: Retained HIL artifacts include infrared markers, mounting holes, and unnatural surface glow, while extreme glow, overexposure, and failed post-processing lead to sunlamp-sample rejection.

Images of Rejected Samples

The paper documents rejected sunlamp samples and validated HIL labels through projected spacecraft wireframes. The remaining supplied passages are author biographies rather than dataset evidence.

  • Images of Rejected Samples: Sunlamp samples are rejected when extreme surface glow, camera overexposure, or failed post-processing makes them unusable.
  • Images of Rejected Samples: Projected Tango wireframes align with HIL mockups, validating the accuracy of labels estimated from the TRON facility.
Loading 2110.03101v2…