Source-linked AI summary
One year in a forest: Analyzing the challenges of autonomous navigation in subarctic environments
Matěj Boxan, Nicolas Lauzon, Veronica Vannini, Mathis Turgeon-Roy, François Pomerleau
TL;DR
Autonomous navigation in subarctic forests lacks evaluation under substantial seasonal variation, despite unreliable GNSS and limited infrastructure. The paper conducts a year-long field study of nine methods over 64 km, finding significant degradation in state-of-the-art techniques and greater fragility in complex SLAM systems than in a proprioceptive baseline.
Problem
Subarctic navigation is difficult because GNSS and cloud computing are unreliable, while established exteroceptive methods are mainly evaluated in structured or weakly seasonal environments.
Method
The paper evaluates nine odometry, localization, and mapping methods through a year-long mobile-robot deployment covering 64 km of seasonal subarctic data.
Results
State-of-the-art methods degrade significantly in subarctic conditions, while simple proprioceptive odometry shows strong robustness and complex SLAM provides limited accuracy gains with greater fragility.
Takeaways & Limitations
Field reports should include proprioceptive baselines, climate classifications, and explicit failure rates when evaluating navigation outside laboratory conditions.
Abstract
from arXiv · showhide
Subarctic regions have the potential to see increased deployment of autonomous robots in applications including forestry, mining, and environmental monitoring. In these conditions, an autonomous system's reliance on GNSS or cloud computing is precarious due to dense tree canopies and atmospheric attenuation, necessitating onboard sensing and data processing. However, established exteroceptive modalities, including cameras, lidars, and radars, are typically evaluated in structured urban settings or in environments that lack significant seasonal variations. To address this, we present a field report on a year-long deployment of a mobile robot in a subarctic boreal forest. We evaluate 64 km of data using nine odometry, localization, and mapping methods and assess their performance across seasonal changes. The performed experiments suggest that the environment changes significantly hinder the performance of state-of-the-art techniques, which show increased fragility when subject to conditions characterized by self-similar scenes or tall snowbanks. Additionally, complex Simultaneous Localization and Mapping (SLAM) algorithms offer limited accuracy gains over a proprioceptive baseline while significantly increasing system fragility. Furthermore, by correlating the position drift with features and confidence weight distribution, we show that visual-based SLAM methods are particularly affected by the seasonal changes. Additionally, we investigate the task of cross-season localization in a prior map. While lidar-based methods successfully completed localization runs between seasons, radar and visual methods are prone to failure due to a few matching features between runs, even within the same season. Finally, we detail the challenges and lessons learned from this year-long trial, including a multi-season Teach and Repeat (T&R) evaluation using both radar and lidar-based pipelines.
I. INTRODUCTION
The paper motivates autonomous navigation in subarctic boreal forests, where seasonal conditions constrain GNSS, communications, sensing, and onboard operation. It introduces a year-long analysis of nine navigation techniques across changing environmental conditions.
- Subarctic forests create opportunities for autonomous systems in mining, forestry, and inspection despite harsh environmental conditions.
- GNSS is precarious because low-horizon satellites and snow-laden canopies obstruct line of sight.
- Cloud and cellular dependence is unreliable because atmospheric attenuation and limited networks require edge computing.
- The study focuses on humid subarctic boreal forests with major seasonal changes from snow accumulation and temperature swings.The reported summer-to-winter temperature difference exceeds 60 °C, with up to six meters of annual snowfall.
- The field campaign analyzes nine lidar-, radar-, and visual-based odometry, localization, and mapping techniques over a year.The contributions include assessing seasonal effects, SLAM fragility, and cross-season localization.
- Seasonal conditions challenge cameras, lidars, and radars through illumination changes, snow-covered objects, soft substrates, and evolving snowbanks.Cameras face saturation or underexposure, while radar matching is affected by changing surfaces and vehicle attitude.
B. LOCALIZATION AND MAPPING IN MULTI-SEASON
Prior multi-season localization research is concentrated in milder or structured environments, while subarctic landscapes undergo stronger structural changes. This paper extends evaluation to visual, lidar, and radar place recognition across a full year in a Dfc climate zone.
- Only 8 % of long-term localization and mapping methods explicitly address seasonal transitions.
- Existing multi-season datasets commonly capture urban changes from dynamic objects or construction rather than broad environmental restructuring.
- Seasonal localization challenges vary by region and are closely linked to climate zones.The paper uses Köppen-Geiger classifications to contextualize seasonal changes across deployments.
- Prior work covers seasonal datasets in humid subtropical, Mediterranean, oceanic, and continental climates, including agricultural and urban settings.
- Place recognition is important for loop closure and re-localization, but prior studies often focus on mild transformations or single-sensor modalities.
- This paper reports visual, lidar, and radar place-recognition performance across the full year in a subarctic Dfc climate zone.
III. METHODOLOGY
The methodology frames autonomous-navigation evaluation as a system-wide field problem rather than only a controlled parameter comparison. It introduces time-series concepts for disturbances, stressors, robustness, fragility, and resilience while acknowledging limits on definitive classification.
- Field evaluation is emphasized because localization failures outdoors increase maintenance and operational overhead, especially in remote locations.
- The proposed methodology evaluates system behavior in uncontrolled environments where disturbances and stresses influence localization errors.
- A. TAXONOMY OF TIME SERIES EVALUATION: The paper distinguishes short meaningful disturbances from longer-duration stressors in long-term deployments.
- A. TAXONOMY OF TIME SERIES EVALUATION: Fragile systems fail shortly after disturbances or under stress, whereas robust systems maintain performance with sublinear error growth over time.
- A. TAXONOMY OF TIME SERIES EVALUATION: The taxonomy is intended to identify qualitative trends and behavioral patterns rather than assign definitive quantitative categories.Compounded stressors and disturbances make variables difficult to isolate in off-road environments.
B. TAXONOMY OF LOCALIZATION AND MAPPING METHODS ACROSS SEASONS
The paper frames localization and mapping methods as a complexity–fragility trade-off in multi-season outdoor environments, then defines metrics for evaluating drift, trajectory error, and failures under field conditions.
- Evaluation perspective: The evaluation compares algorithmic complexity with system fragility rather than selecting a single optimal method, emphasizing trends across seasonal perturbations.
- Method taxonomy: Proprioceptive odometry uses wheel encoders and IMUs, remaining invariant to appearance changes but vulnerable to short disturbances such as wheel slippage.
- Method taxonomy: Exteroceptive odometry adds camera, lidar, or radar measurements to correct proprioceptive drift, but increases fragility risk because sensing is affected by environmental disturbances.
- Method taxonomy: Loop closure and pose-graph optimization globally propagate constraints to correct drift, while prior-map localization anchors pose but can degrade when seasonal or structural changes remove valid correspondences.
- Trajectory metrics: ATE measures global consistency but can exaggerate early rotational errors and depends heavily on initial alignment, whereas RTE measures normalized local drift over sliding windows.
- Trajectory metrics: RTDE compares estimated and ground-truth traveled distances without orientation, but can disguise fragile systems as robust by ignoring rotational errors.
- Trajectory metrics: SARTE replaces RTE when ground truth lacks orientations by sequentially aligning windows using the preceding window, but its accuracy depends on ground-truth quality and window geometry.
- Failure metrics: A run is classified as failed when its estimated trajectory ends before 0.95 tGT, and LFR measures the portion of a prior-map localization trajectory lost after failure.
IV. EXPERIMENTAL SETUP
The experimental setup draws on a 12-month field campaign with more than 64 km of synchronized sensor data collected across diverse subarctic driving conditions.
- Field campaign: The campaign recorded more than 64 km of sensor data over 12 months under on-road, off-pavement, and off-trail conditions in a subarctic Dfc climate zone.
A. DATA RECORDING
The year-long deployment used a tracked half-tonne UGV to collect repeated data across six trajectories and three seasons in a subarctic boreal forest.
- The platform was a half-tonne Clearpath Warthog UGV equipped with stereo and monocular cameras, lidar, radar, GNSS receivers, and an IMU.
- The robot was manually controlled for most runs, while a minority used lidar Teach and Repeat with waypoints from a prior deployment.
- Six repeated trajectories covered on-road, off-road, off-trail, forest-road, quarry, and irregular-terrain routes.Distances ranged from 300 m to 2200 m, with terrain including steep elevation changes, dense vegetation, heavy snow, trees, and boulders.
- All trajectories included at least three deployments in winter, summer, and autumn.
B. EVALUATED METHODS
The study evaluates diverse odometry, localization, and mapping methods across lidar, radar, camera, and proprioceptive modalities, with localization modes varying by method.
- Nine odometry and localization algorithms were selected to represent lidar, radar, and camera modalities and different operating principles.The set included both classic and deep-learning-based visual feature extraction approaches.
- Localization along a prior path was unavailable for proprioceptive odometry, DROID-SLAM, KISS-SLAM, and Navtech-Radar-SLAM for stated map-representation or implementation reasons.
- The study asks how methods perform under subarctic stress, how modalities respond to seasonal changes, and how methods localize in an existing map over long-term seasonal change.
- Nominal pose accuracy is evaluated on six diverse trajectories over twelve months in subarctic conditions.
A. SUBARCTIC CONDITIONS AS A STRESSOR ON SLAM
Across six trajectories, subarctic conditions degraded many exteroceptive methods and exposed fragility in loop closure and SLAM systems, while the proprioceptive baseline remained comparatively strong.
- All methods were evaluated offline, with sensor frames processed sequentially during postprocessing to avoid data-loss and resource issues.
- 10 %pt was the average degradation of exteroceptive odometry relative to the Proprioceptive baseline.
- Only KISS-ICP achieved a mean error lower than 1 %, while visual methods generally transferred poorly from benchmark conditions.cuVSLAM* averaged 8.7 % error, and DROID-SLAM* exceeded 13 % mean SARTE.
- Loop closing and PGO improved state estimation for only two of six evaluated methods.The authors attribute much of the improvement to PGO smoothing rather than detected loop closures.
- ORB-SLAM3 reduced mean degradation from 22 %pt to 0.5 %pt after PGO.This was the largest improvement among evaluated PGO methods and followed 42 detected loop closures across 60 runs.
- Boreal-forest self-similarity and sensor occlusion challenged place recognition, leaving loop-closure performance consistently low across seasons.
B. SEASONS AS A STRESSOR ON SENSING
Seasonal changes affected modalities differently: visual methods were sensitive, whereas lidar and radar localization and mapping showed comparatively limited inter-season variation.
- The proprioceptive median error was below 0.9 % in summer and increased to 1.8 % in autumn and winter.Winter also produced a larger interquartile range.
- WILN maintained a median error below 0.6 % across summer, autumn, and winter.The authors relate this robustness to dense lidar point clouds capturing nearby structure and stable tree canopies.
- cuVSLAM and DROID-SLAM worsened significantly from summer to autumn and winter, while ORB-SLAM3 performed similarly in summer and autumn.
- Radar showed a less pronounced seasonal trend, with summer medians for RT&R and Navtech-Radar-SLAM within winter distributions’ interquartile ranges.
- The authors conclude that ranging-sensor localization and mapping is robust to seasonal changes, whereas visual-based methods are sensitive.
C. SEASONS AS A STRESSOR ON VISUAL FEATURE
Seasonal changes reduce visual feature and confidence-weight support, increasing translational drift, especially over flat winter snow. Cross-season localization is also difficult for visual and radar methods, while lidar methods are more robust.
- C. SEASONS AS A STRESSOR ON VISUAL FEATURE: Winter images contain fewer visual features and lower lower-image DROID-SLAM confidence, conditions associated with higher translational error.The analysis focuses on ORB, cuVSLAM, and DROID-SLAM feature or weight distributions.
- C. SEASONS AS A STRESSOR ON VISUAL FEATURE: More than 175 detected features per image is associated with RTDE below 1% in predominantly summer data.
- C. SEASONS AS A STRESSOR ON VISUAL FEATURE: DROID-SLAM error exceeds 30% in winter when lower-image confidence weights sum to under 100, but drops to 5% when the sum exceeds 200.
- C. SEASONS AS A STRESSOR ON VISUAL FEATURE: Flat, featureless snow is identified as the primary source of translational drift for the evaluated visual methods, whereas plowed roads reduce the seasonal effect.Vehicle tracks can provide strong local features, but these features are season-dependent.
- PRIOR MAP: Cross-season localization remains challenging: prior-map methods show high failure rates, while cuVSLAM initialized successfully only in its three control sequences.The figure reports LFR and SARTE across mapping and active-localization seasons; light cells indicate better performance.
VI. CHALLENGES AND LESSONS LEARNED
The field campaign section summarizes the year-long deployment, its remaining operational challenges, a multi-season Teach and Repeat trial, and evaluation at scale.
- VI. CHALLENGES AND LESSONS LEARNED: The campaign comprised ten individual field days and five longer deployments, each spanning a full work week, for 35 total field days.
- VI. CHALLENGES AND LESSONS LEARNED: The authors report remaining challenges for subarctic robot operations and a multi-season Teach and Repeat field trial executed in March 2025.
A. ADDITIONAL PERTURBATIONS
Beyond algorithmic stress, the deployment encountered collisions, snow immobilization, thermal variability, attitude changes, and sensor-specific Teach and Repeat failures.
- A. ADDITIONAL PERTURBATIONS: The robot experienced 18 small collisions and 62 partial snow immobilizations during off-trail operation.Most collisions involved boulders or small trees, while most immobilizations occurred uphill during a January deployment with fresh snow.
- A. ADDITIONAL PERTURBATIONS: Snow properties affected mobility: fresh cohesive November snow provided traction after compaction, whereas wet April snow caused wheel slip and sticking.Static equipment also accumulated snow and required regular cleaning.
- A. ADDITIONAL PERTURBATIONS: Recorded temperatures ranged from −34 ◦C in winter to 25 ◦C in summer, including a 26 ◦C change on May 28th.
- VARIATIONS: Radar localization assumes similar UGV roll and pitch between teach and repeat runs because radar is a two-dimensional sensor.Snowbanks can alter vehicle attitude and produce ground strikes or feature-poor scan regions.
- VARIATIONS: ICP-based lidar localization failed near snowbanks up to 3 m high but recovered after filtering returns below 1.5 m above the sensor.The proposed filtering solution may not transfer to lidar sensors with narrower vertical fields of view.
- VARIATIONS: Snow-covered ground degraded controller motion-model predictions, with sharp turns requiring more energy and track rotational speed because of reduced traction.
C. TRAJECTORY EVALUATION AT SCALE
The paper argues that large-scale off-road evaluation needs environmental characterization and failure reporting alongside standard error metrics. Standard metrics can conceal catastrophic failures, while ground-truth orientation is also difficult to obtain reliably.
- C. TRAJECTORY EVALUATION AT SCALE: Köppen climate classification is proposed to contextualize seasonal conditions across field datasets and public benchmarks.The authors connect this context to comparing results and discussing method adaptability.
- C. TRAJECTORY EVALUATION AT SCALE: Standard ATE or RTE metrics may hide catastrophic failures in meteorologically challenging or feature-poor environments.Transitions between high estimation error and complete failure remain difficult to represent with a common metric.
- C. TRAJECTORY EVALUATION AT SCALE: Independent ground-truth trajectories are difficult to acquire because GNSS-derived orientation is highly sensitive to elevation noise.The deployed UGV used three receivers, but point-to-point orientation estimates remained sensitive to GNSS measurement noise.
- C. TRAJECTORY EVALUATION AT SCALE: The proposed methodology evaluates nine radar-, lidar-, and visual-based methods on over 64 km recorded across twelve months of seasonal change.
- C. TRAJECTORY EVALUATION AT SCALE: The taxonomy showed that performance degradation is difficult to attribute because environmental stressors and unmodeled implementation quality both influence measurements.The authors recommend proprioceptive baselines, climate classification, and explicit failure-rate reporting in future field reports.
- C. TRAJECTORY EVALUATION AT SCALE: Robust long-term autonomy across seasonal off-road subarctic conditions still requires further investigation, including improved place-recognition and error metrics.
APPENDIX
Qualitative results compare lidar-based odometry and pose-graph-optimized trajectories across seasonal deployments, highlighting winter snow-bank effects and PGO’s scale-estimation improvement.
- WILN diverges from the Ground Truth trajectory during the winter Yellow deployment because of snow banks.
- Figure 14 and Figure 15 present lidar-based odometry and odometry with Pose Graph Optimization, respectively.
- PGO improves scale estimation from KI.