Source-linked AI summary
Learning Image-based Tree Crown Segmentation from Enhanced Lidar-based Pseudo-labels
Julius Pesonen, Stefan Rua, Josef Taher, Niko Koivumäki, Xiaowei Yu, Eija Honkavaara
TL;DR
Separating individual tree crowns in aerial imagery is difficult, while reliable image-based models usually require costly manual annotations. This paper trains RGB and multispectral models from ALS-derived pseudo-labels enhanced by SAM 2, and reports that the resulting method outperforms available out-of-the-box alternatives on the Espoonlahti dataset. Its main limitation is imperfect transferability and overlapping predictions in some deployment settings.
Problem
Individual crown separation is challenging, and image-based learning methods require training annotations that are difficult to create across locations and tree species.
Method
The method trains RGB and multispectral image models using CHM-based ALS segments whose masks are enhanced with SAM 2.
Results
The method outperforms available out-of-the-box alternatives on the Espoonlahti dataset using automatically generated pseudo-labels rather than fully annotated training data.
Takeaways & Limitations
SAM 2-enhanced CHM annotations provide relatively high-quality, location-specific training annotations for optical image models without manual annotation cost.
Takeaways & Limitations
The models have limited transferability to new datasets and can produce overlapping masks that NMS fails to remove when one mask is much larger.
Abstract
from arXiv · showhide
Mapping individual tree crowns is essential for tasks such as maintaining urban tree inventories and monitoring forest health, which help us understand and care for our environment. However, automatically separating the crowns from each other in aerial imagery is challenging due to factors such as the texture and partial tree crown overlaps. In this study, we present a method to train deep learning models that segment and separate individual trees from RGB and multispectral images, using pseudo-labels derived from aerial laser scanning (ALS) data. Our study shows that the ALS-derived pseudo-labels can be enhanced using a zero-shot instance segmentation model, Segment Anything Model 2 (SAM 2). Our method offers a way to obtain domain-specific training annotations for optical image-based models without any manual annotation cost, leading to segmentation models which outperform any available models which have been targeted for general domain deployment on the same task.
1. Introduction
Individual tree crown mapping supports urban and forest management, but separating overlapping crowns in aerial imagery requires substantial, adaptable training data. The study proposes training RGB and multispectral image models from ALS-derived pseudo-labels enhanced with SAM 2.
- Tree crown maps support urban inventories, forest health monitoring, carbon-storage assessment, pest protection, and fire-risk assessment.
- Aerial data covers larger areas faster than field measurements while providing higher spatial resolution than satellite imagery.
- Separating individual trees from backgrounds and one another remains difficult because imagery varies by location, sensor, weather, and acquisition time.
- Learning-based methods generally depend on large amounts of human-annotated data, which is impractical to create for all locations and species.
- The study trains RGB and multispectral image models with CHM-based ALS segments and uses SAM 2 to enhance the pseudo-labels.
2. Related work
Tree crown segmentation research uses lidar, optical imagery, or both, but optical models commonly require manual supervision and lidar data can be costly or unavailable. The study addresses an unexplored combination: SAM-enhanced lidar segments for training optical image models.
- Tree crowns are detected with lidar, optical sensors, or both, while optical methods with the strongest performance typically require hand-crafted labels.
- Surveyed prior results span F1 scores of 0.82 to 0.98, precision of 0.85 to 0.91, recall of 0.88 to 0.91, and mIoU of 0.49 to 0.94.
- Lidar-derived segments can combine lidar’s structural information with optical imagery, but high-density lidar is costly and low-resolution lidar produces noisy labels.
- Prior work had not explored using SAM-enhanced lidar segments to train optical image-based models for individual tree crown segmentation.
3. Materials and methods
The study combines high-resolution multispectral imagery with ALS point clouds to generate coarse tree segments, filter outdated false positives, enhance labels with SAM 2, and train a pseudo-supervised Mask R-CNN. Evaluation uses a manually annotated test area.
- 3.1. Data: The dataset combines a 5 cm orthophoto covering approximately 2 km2 with multispectral ALS point clouds acquired in 2016.The orthophoto was captured in July 2023, while ALS channels used 1550 nm, 1064 nm, and 532 nm wavelengths.
- 3.1. Data: Testing used one 4471 by 4893 pixel area with 362 manually segmented trees; approximately 24 000 trees remained for training and validation.
- 3.2. Coarse segments: Individual crown pseudo-labels came from a 0.5 m CHM, smoothed and segmented with treetop markers using marker-controlled watershed.
- 3.2. Coarse segments: Because ALS preceded imagery by seven years, segments in areas replaced by buildings were filtered using an average NDVI threshold of 0.2.
- 3.3. SAM 2 enhancement: SAM 2 converted blocky CHM segments into more precise masks by using each coarse segment’s bounding box as an image query.
- 3.4. Pseudo-supervised model: An improved Mask R-CNN with an ImageNet-pretrained ResNet-50-FPN backbone predicted tree masks from RGB or multispectral imagery using coarse and SAM 2-enhanced labels.
4. Experiments and results
The proposed pseudo-supervised image models outperform available out-of-the-box alternatives, while SAM 2 label enhancement improves downstream performance. Results also expose limitations involving transferability, image edges, and overlapping predictions.
- Comparative results: The proposed models outperform all out-of-the-box alternatives on the Espoonlahti dataset without fully annotated training data.Pseudo-label generation can be fully automated where ALS and aerial imagery are available.
- Overall interpretation: The results indicate that automated label collection and zero-shot enhancement can provide sufficient data for training more broadly generalisable tree-crown models.The reported metrics fall within ranges observed for models trained with fully supervised data, although the best model still has room for improvement.
- Comparative results: The proposed method produces more balanced predictions than conservative zero-shot models and the supervised U-net's false-positive, oversized segments.Detectree2 and SAM 3 have low recall, whereas the supervised U-net produces more false positives and overly large segments.
- Ablation: SAM 2 label enhancement improves downstream model performance across input modalities, while the optimal channel configuration remains unclear.Near-infrared channels appear useful, but models combining them with other channels do not consistently capture their full value.
- Failure points: The method's practical weaknesses include limited transferability, degraded edge predictions, and overlapping masks that NMS may not filter.These issues should be considered when deploying the models or developing new models from the study.
- Failure points: Overlapping masks remain problematic when one predicted mask is substantially larger than another because IoU-threshold NMS may not capture the overlap.An additional filtering step based on mask intersection and area could address this case.
- Failure points: The final RGB image-based method fails on some public tree-crown images, likely because their spatial resolution and colours differ from the study dataset.Heavier augmentation is suggested as a possible mitigation but was not explored further.
- Failure points: Predictions deteriorate near image edges because training labels covered only the image middle; larger-area inference should use overlapping patches.The recommended overlap is 1/4 of the image size in each dimension.
5. Conclusions
The method uses CHM-generated annotations enhanced by SAM 2 to train optical image-based tree crown models without manual annotation. However, broader validation and data diversity remain necessary for generalisable real-world performance.
- CHM-generated annotations can train image-based tree crown segmentation models, while SAM 2 preprocessing clearly enhances those labels.
- The approach provides relatively high-quality training annotations for optical models without manual annotation cost.
- The study lacks extensive data from varied locations, limiting evaluation and training of more generalisable models; metric gains may not perfectly reflect real-world performance.
- More varied geographic areas and time periods are needed to develop models that perform better across scenarios.
Funding sources
The study received funding from the Research Council of Finland and the European Union’s NextGenerationEU instrument through the Research Council of Finland.
- The Research Council of Finland funded projects on autonomous drone-based hyperspectral forest vegetation analysis and forest ecosystem change measurement.
- The European Union’s NextGenerationEU instrument supported the Multirisk Project through a Research Council of Finland grant.
- The listed Research Council of Finland decision and grant numbers are 357380, 346382, and 353263.
Appendix
The appendix documents supplementary visualizations and Detectree2 checkpoint metrics, including NDVI-based segment filtering and comparisons across models, input modalities, and training labels.
- Figure A.1 shows the NDVI distribution of initial tree crown segments and marks the cut-off threshold with a black dashed vertical line.
- Table A.1 reports test-set metrics for all Detectree2 model checkpoints.
- Figure A.2 visualizes outputs from Grounded SAM and DeepForest models.
- Figure A.3 compares outputs from models using different input modalities and coarse or pseudo-label training, with All denoting RGB, red edge, and NIR inputs.