Source-linked AI summary
Tree species mapping in Denmark: A comparison of spectral-temporal features with geospatial foundation model embeddings
Alkiviadis Koukos, Spyros Kondylatos, Thomas Nord-Larsen, Lotte Nyborg, Christian Tøttrup, Kenneth Grogan
TL;DR
The paper addresses the need for national, fine-scale tree-species information and tests whether EO foundation-model embeddings can complement or replace engineered Sentinel features. It compares these representations with canopy height across classifiers and stand types, finding that STF maximizes accuracy while TESSERA is especially competitive with limited training data. The resulting best model produces an open-access 10 m national tree-species map of Denmark.
Problem
Large-scale tree-species information is often lacking, while mixed stands and sparse field inventories complicate realistic national mapping.
Method
The study compares Sentinel-1/2 spectral-temporal features with TESSERA and AlphaEarth embeddings, combines both with canopy height, and evaluates RF, XGBoost, and MLP models for pure and mixed stands.
Results
79.9% overall accuracy was achieved by the area-adjusted validation of the resulting national 10 m tree-species map.
Takeaways & Limitations
Dense spectral-temporal representations remain strongest for maximizing accuracy, while TESSERA offers a viable alternative with greater robustness when training data are limited.
Takeaways & Limitations
Optical satellite observations and their derived embeddings cannot fully characterize subcanopy composition or fine-scale canopy heterogeneity, and the study labels only dominant species.
Abstract
from arXiv · showhide
We map tree species across Denmark using National Forest Inventory plots and EO data, while evaluating the potential of foundation models for large-scale forest characterization. We compare two alternative input representations for tree species classification: (i) manually engineered spectral-temporal features (STF) derived from multi-temporal Sentinel-1 and Sentinel-2 observations, and (ii) embeddings generated by the EO FMs TESSERA and AlphaEarth. Both representations are complemented with canopy height information. Random forest, XGBoost, and Multi-Layer Perceptron (MLP) classifiers are evaluated for all input representations, with separate assessments for pure and mixed forest stands. The STF-based MLP achieves the highest classification performance, yielding macro F1 scores of 0.843 and 0.653 for pure and mixed stands, respectively. The MLP trained on TESSERA embeddings delivers competitive performance for pure stands, achieving results within 1.1 percentage points of the best-performing model. TESSERA consistently outperforms STF-based models when fewer than approximately 25% of training plots are available, demonstrating a substantial advantage under limited training data. Multi-year observations systematically improve classification accuracy relative to single-year inputs, while ablation experiments reveal the complementary contributions of Sentinel-1 backscatter, spectral indices, and canopy height data. The best-performing model is subsequently applied at the national scale to generate a 10 m tree species map of Denmark. Area-adjusted validation indicates an overall map accuracy of 79.9%. The resulting map, released as an open-access product, is the first high-resolution national tree species map of Denmark and provides a valuable resource for forest monitoring, ecological research, and land management applications.
1. Introduction
The study addresses the need for fine-scale, national tree-species information by comparing engineered Sentinel features with EO foundation-model embeddings for Danish mapping. It also evaluates performance under realistic pure and mixed stand conditions and varying training-data availability.
- Motivation: Large-scale complementary information on tree species remains limited despite extensive knowledge of tree distribution and forest cover.Such information supports biodiversity, carbon, water-cycle, economic, recreational, and cultural assessments.
- Motivation: Fine-scale species information is increasingly needed because diverse forest management reduces the suitability of conventional stand-level inventory and mapping approaches.Forest owners and managers require spatially explicit information for planning and management.
- Research gap: Existing foundation-model evaluations rarely address fine-grained ecological tasks or compare against detailed multi-temporal Sentinel representations.This leaves their value for tree-species mapping relative to engineered features uncertain.
- Approach: The framework compares Sentinel-1 and Sentinel-2 spectral-temporal features with TESSERA and AlphaEarth embeddings, combining both with canopy height and testing RF, XGBoost, and MLP classifiers.The comparison is conducted for national-scale Danish tree-species classification.
- Evaluation design: Mixed-species stands are evaluated separately because their structural and spectral complexity makes classification more difficult than in pure stands.This provides a more operationally realistic assessment than evaluations restricted to pure stands.
- Contributions: The study contributes a systematic comparison across training-data availability, feature sources, temporal inputs, and information components, alongside an open-access 10 m national map.The analysis examines spectral indices, Sentinel-1 data, canopy height, and multi-year observations.
2. Data
The study uses Denmark’s systematic NFI network as reference data for national mapping, combining field measurements of species and structure with geospatially distributed sampling. Plot labels and selection criteria distinguish pure and mixed forest conditions for model development and validation.
- Study area: Denmark’s fragmented forests span diverse forest types and management conditions, creating a challenging context for satellite-based species classification.Forests are often interspersed with agriculture, urban areas, and infrastructure.
- NFI sampling design: The NFI samples Denmark on a 2 × 2 km grid, with four plots arranged at the corners of a 200 × 200 m square in each cell.One-fifth of plots are measured annually during a five-year cycle.
- NFI sampling design: One-third of NFI clusters are permanent and remeasured each cycle, whereas two-thirds are temporary and randomly relocated within their grid cells.Permanent clusters support tracking of forest change over time.
- NFI sampling design: The NFI contains around 43,000 plots, but only clusters with at least one forest plot are measured during a given cycle.Field plots are identified using recent aerial imagery and the FAO forest definition.
- Reference data: The study uses 8,663 plots after selecting 2014–2022 measurements, removing rare classes, and excluding plots affected by post-measurement disturbances.Labels target the dominant tree species in the upper canopy layer.
- Plot selection and labels: Pure plots require domfrac ≥0.90 and standfrac ≥0.80, while mixed plots require domfrac × standfrac ≥0.50.The resulting dataset contains 2,683 pure plots and 3,098 mixed plots.
2.3. Additional forest layers
Additional forest layers and processed Sentinel-2 observations provide the spatial framework and spectral-temporal inputs for national tree species classification.
- Forest extent: The 2022 Digital Forest Map defines the mapped forest extent and is rasterized to the study’s 10 m analysis grid.It covers approximately 653,250 ha in polygon form and reports approximately 98% accuracy.
- Validation layers: NST stand-level species data support visual assessment but omit many privately owned and commercial forests.The dataset remains useful for checking whether broad spatial patterns are captured.
- Sentinel-2 processing: Sentinel-2 imagery from 2020–2022 is atmospherically and radiometrically processed before forming a national 10 m data cube.The cube supports querying, interpolation, and spectral-index calculations.
- Spectral features: Classification uses Sentinel-2 bands and indices selected to capture canopy phenology, moisture content, and chlorophyll-related properties.The provided equation fragments include formulas for NDVI and GEMI-related features.
- Temporal features: Whittaker-Eilers smoothing produces 10-day interpolated series, supplemented by annual seasonal summary statistics.The summaries include maximum, mean, median, standard deviation, and multiple percentiles.
2.5. Sentinel-1
Sentinel-1 radar data add monthly structural information, while canopy-height metrics and annual foundation-model embeddings provide complementary classification inputs.
- Sentinel-1 inputs: Sentinel-1 C-band SAR complements Sentinel-2 with information primarily related to forest structure.All available Danish imagery from 2020–2022 is used, including ascending and descending orbits.
- Sentinel-1 processing: Monthly statistics are computed for VV, VH, and VH–VV difference and integrated into the shared data cube.The statistics include mean, median, standard deviation, and the 25th and 75th percentiles.
- Canopy height: Canopy height is estimated by subtracting the digital terrain model from the digital surface model.Both elevation components have 0.4 m spatial resolution and characterize vegetation height above ground.
- Canopy height: Canopy-height metrics summarize each 10 m grid cell using median, maximum, standard deviation, and upper percentiles.These metrics are added as inputs to the classification models.
- Input representations: The STF representation combines 10-day Sentinel-2 series and seasonal summaries with monthly Sentinel-1 statistics, whereas FM embeddings are annual and canopy-height metrics are static.The study evaluates annual AlphaEarth and TESSERA embeddings from 2020, 2021, and 2022 without fine-tuning.
3. Methodology
The study formulates dominant tree species mapping as pixel-based multi-class classification using NFI labels and multiple EO representations, while separating pure and mixed plots for training and evaluation. Models are compared under consistent conditions, predictions are aggregated to plot level, and the selected model is used for national mapping.
- Input representations: STF inputs concatenate multi-year Sentinel-2 spectral features, seasonal statistics, Sentinel-1 statistics, and canopy-height information into a 3,387-feature vector.The manually engineered representation uses observations from 2020, 2021, and 2022.
- Input representations: AlphaEarth and TESSERA embeddings from 2020–2022 are concatenated across years without temporal aggregation or fine-tuning.Both embedding representations are evaluated as alternatives to manually engineered features.
- Data and sample extraction: NFI plots provide dominant species labels, and four central 10 m × 10 m pixels represent each plot for feature extraction.The extracted inputs include spectral-temporal features, foundation-model embeddings, and canopy-height metrics.
- Data and sample extraction: Pure and mixed plots are processed separately because mixed plots create greater uncertainty when plot labels are transferred to individual pixels.Classifiers are trained only on pure plots to reduce label noise, while mixed plots are retained for evaluation.
- Modeling and validation: Plot-level evaluation uses majority voting: a plot is correct when at least two of its four pixel predictions match the reference label.This aggregation aligns model evaluation with the spatial scale of the NFI labels.
- Modeling and validation: RF, XGBoost, and MLP classifiers are trained for each input representation and evaluated with macro-averaged and class-level accuracy metrics.Hyperparameters are tuned independently while experimental conditions and input sets remain consistent across classifiers.
4.1. Spectral-Temporal Features vs Foundation Model Embeddings
STF-based models provide the strongest overall classification, while TESSERA embeddings remain competitive for pure stands and show similar class-confusion structure. Mixed stands are substantially harder for every representation, with a larger performance gap between TESSERA and STF.
- Pure and mixed test plots: STF MLP achieves the best classification performance across all metrics when sufficient labeled data are available.TESSERA MLP is 1.1 percentage points below the best-performing model in F1 score.
- Pure and mixed test plots: TESSERA MLP trails the best model by 1.1 percentage points in F1 score, whereas AlphaEarth MLP trails by 10.5pp.MLP consistently outperforms RF and XGBoost across input representations.
- Pure and mixed test plots: TESSERA MLP achieves the highest F1 scores for underrepresented OBL and maple classes, suggesting pre-training may partially offset limited class-specific training data.STF MLP leads for most species, with STF XGBoost performing better for pine and spruce.
- Pure and mixed test plots: Mixed-plot performance is substantially lower across all models, with STF MLP showing a 19pp overall F1 drop relative to pure plots.The reduction is consistent across species and classifiers.
- Pure and mixed test plots: The TESSERA–STF MLP F1 gap widens to 5pp on mixed plots, indicating weaker performance for TESSERA on compositionally heterogeneous stands.The comparison concerns the finer spectral and structural distinctions required for mixed-stand classification.
- Confusion patterns: STF MLP and TESSERA MLP exhibit broadly similar inter-class confusion patterns across pure and mixed forests.This similarity supports the interpretation that TESSERA embeddings encode discriminatory information aligned with handcrafted features.
4.2. Single year vs Multi-year input
Using three years of observations yields the highest classification performance for both STF MLP and TESSERA MLP, while single-year results vary more for STF than for TESSERA. The findings also indicate greater year-selection robustness for foundation-model embeddings.
- Temporal coverage: Three-year inputs achieve the highest F1 scores for both STF MLP and TESSERA MLP on pure and mixed test plots.The comparison includes single-year, consecutive two-year, and full three-year configurations.
- Temporal coverage: STF MLP pure-plot F1 varies by 7.6pp between its highest and lowest single-year configurations.The 2020 input performs best and the 2022 input worst among single-year STF configurations.
- Temporal coverage: TESSERA MLP pure-plot single-year F1 ranges from 0.771 to 0.797, showing less variation across years than STF.Adding a second year gives modest improvements, while the full three-year configuration improves further.
- Temporal coverage: The results indicate that acquisition year matters, while foundation-model embeddings are more robust to year selection than STF.This conclusion follows the observed variation across individual-year configurations.
4.3. Ablation Study
Ablation results show that the complete STF feature set performs best for both pure and mixed plots, with Sentinel-1 providing the largest individual contribution. Canopy height and spectral indices add smaller but consistent improvements.
- Component contributions: Removing canopy height or spectral indices causes smaller but consistent performance reductions for both forest categories.Their contributions are weaker individually than Sentinel-1 but remain complementary to the other inputs.
- Component contributions: Removing all three supplementary components reduces F1 by 7.2pp for pure plots and 3.5pp for mixed plots.This comparison retains only Sentinel-2 spectral bands.
4.4. Forest Maps
The study produces Denmark’s first 10 m national tree-species map and evaluates its spatial patterns, area estimates, and classification accuracy against reference information and national statistics.
- The final STF MLP was applied across Denmark to produce the first national-scale tree-species map at 10 m resolution.
- The map broadly reproduces forest patterns and captures finer within-stand variation than stand-level polygon references, especially near boundaries and in heterogeneous areas.
- Pine and OBL mapped areas agree nationally within ±3%, while beech, spruce, and fir show the largest positive discrepancies.
- Nordjylland shows the closest regional correspondence, whereas Syddanmark has the largest discrepancies from reference statistics.
- 79.9% (±1.3%) overall accuracy was estimated using area-adjusted validation combining pure and mixed test plots.
- Area-adjusted estimates bring most species within a ±10% agreement range by accounting for classification errors, although low-PA classes such as maple and larch remain problematic.
5. Discussion
The discussion compares engineered spectral-temporal features with foundation-model embeddings, examining performance, data requirements, operational trade-offs, and limitations in heterogeneous forests.
- Classification performance: Manually engineered spectral-temporal representations remain particularly effective for fine-grained species discrimination under heterogeneous forest conditions.
- Classification performance: Lower accuracy for minority species groups is associated with scarce labeled samples, especially for maple, larch, and other broadleaves.
- Classification performance: TESSERA’s performance gap relative to STF is modest, while both representations share broadly similar species-confusion patterns.
- Operational trade-offs: STF construction requires a resource-intensive, domain-expertise-dependent Sentinel-1 and Sentinel-2 processing workflow, whereas FM embeddings are analysis-ready.
- Interpretability: STF ablations offer more transparent input attribution than FM embeddings, whose dimensions entangle information from multiple modalities.
- Training-data requirements: Below approximately 25% of training data, TESSERA MLP consistently achieves higher macro F1 scores than STF MLP, with the advantage increasing as data decrease.
- Temporal inputs: Three-year inputs produce the highest F1 scores for both STF and TESSERA on pure and mixed plots.
- Temporal inputs: For pure plots, TESSERA scores 0.793 versus STF’s 0.743 with 2022 inputs, and 0.809 versus 0.787 with 2021–2022 inputs.
6. Conclusion
The study finds that dense spectral-temporal Sentinel representations maximize national tree species classification accuracy, while TESSERA embeddings offer competitive performance and greater robustness with limited training data. The best-performing model produced Denmark’s first open-access 10 m national tree species map, validated at 79.9% overall accuracy.
- The STF MLP achieved the highest overall performance, whereas TESSERA embeddings remained competitive, particularly for pure forest stands.TESSERA required substantially less preprocessing and no task-specific feature engineering.
- TESSERA embeddings outperformed STF models when fewer than 25% of available reference data were used for training.This indicates greater robustness under limited training data conditions.
- Multi-year observations improved classification performance over single-year inputs, and ablations highlighted contributions from Sentinel-1 SAR, canopy height, and spectral indices.The results support combining complementary information sources for tree species discrimination.
- Separate evaluations showed that mixed forest stands were more difficult to classify than pure stands.The finding underscores the relevance of forest compositional complexity in operational evaluations.
- 79.9% overall accuracy was achieved by the open-access 10 m national tree species map produced with the best-performing STF MLP.The map is presented as Denmark’s first open-access national tree species map at this resolution.
- The findings position dense spectral-temporal Sentinel representations as effective for maximizing accuracy and TESSERA embeddings as an alternative under limited training data.The study notes that foundation models may reduce reliance on task-specific feature-engineering workflows as they mature.
CRediT authorship contribution statement
The contribution statement assigns authorship roles across writing, conceptualization, methodology, software, validation, analysis, investigation, visualization, resources, and supervision. The supplied passages also state that the Denmark dominant-tree-species map is available.
- Alkiviadis Koukos is credited with original-draft writing and broad methodological, software, validation, analysis, investigation, data-curation, and visualization contributions.
- Spyros Kondylatos is credited with review and editing plus conceptualization, methodology, software, validation, analysis, investigation, and visualization.
- Thomas Nord-Larsen contributes through review and editing, resources, validation, and data curation, while Lotte Nyborg contributes validation, resources, funding acquisition, and project administration.
- Christian Tøttrup is credited with review and editing, conceptualization, methodology, validation, analysis, investigation, visualization, and supervision.
- The map of Denmark’s dominant tree species is available through the stated product resource.