Source-linked AI summary
A Methodology to Derive Global Maps of Leaf Traits Using Remote Sensing and Climate Data
Alvaro Moreno-Martinez, Gustau Camps-Valls, Jens Kattge, Nathaniel Robinson, Markus Reichstein, Peter van Bodegom, Koen Kramer, J. Hans C. Cornelissen, Peter Reich, Michael Bahn, Ulo Niinemets, Josep Peñuelas, Joseph Craine, Bruno E. L. Cerabolini, Vanessa Minden, Daniel C. Laughlin, Lawren Sack, Brady Allred, Christopher Baraloto, Chaeho Byun, Nadejda A. Soudzilovskaia, Steven W. Running
TL;DR
Global plant-trait mapping requires methods that address sparse, irregular, and individually sampled observations that do not directly represent coarser spatial scales. The paper develops a modular chain combining gap-filled trait databases, community-level aggregation, remote sensing, and climate data to produce global 500 m maps of five leaf traits. The resulting maps can support Earth-system and biodiversity modelling, while remaining constrained by species-level sampling bias and the distinction between in-situ measurements and pixel-representative estimates.
Problem
Global trait databases provide extensive observations, but measurements are sparse and irregular, collected at individual-plant scale, and not necessarily representative of coarser spatial variability.
Method
The paper combines random-forest gap filling, functional-type abundance estimates, community weighted means, remote sensing, and climate data to spatialize in-situ leaf traits.
Results
The processing chain produces global 500 m maps of SLA, LDMC, leaf nitrogen and phosphorus concentrations, and nitrogen-to-phosphorus ratios.
Takeaways & Limitations
The maps could replace static PFT maps and support modelling of photosynthetic capacity, fluorescence, vegetation nutrient responses, and soil fertility inference.
Takeaways & Limitations
The approach does not definitively resolve species-level sampling bias in global trait databases, and mapped pixel estimates need not coincide with unweighted in-situ measurements.
Abstract
from arXiv · showhide
This paper introduces a modular processing chain to derive global high-resolution maps of leaf traits. In particular, we present global maps at 500 m resolution of specific leaf area, leaf dry matter content, leaf nitrogen and phosphorus content per dry mass, and leaf nitrogen/phosphorus ratio. The processing chain exploits machine learning techniques along with optical remote sensing data (MODIS/Landsat) and climate data for gap filling and up-scaling of in-situ measured leaf traits. The chain first uses random forests regression with surrogates to fill gaps in the database ($> 45 \% $ of missing entries) and maximize the global representativeness of the trait dataset. Along with the estimated global maps of leaf traits, we provide associated uncertainty estimates derived from the regression models. The process chain is modular, and can easily accommodate new traits, data streams (traits databases and remote sensing data), and methods. The machine learning techniques applied allow attribution of information gain to data input and thus provide the opportunity to understand trait-environment relationships at the plant and ecosystem scales.
1. Introduction
The paper addresses the mismatch between local plant-trait observations and the need for spatially continuous trait information. It combines trait databases, remote sensing, climate data, and machine learning to produce global 500 m maps of five leaf traits.
- Plant traits influence individual establishment, fitness, and survival, while environmental and biogeochemical processes shape plant communities.
- Trait databases support modelling continuous spatial trait variability, but their observations are collected at individual-plant scale and may not represent coarser-scale variability.
- Biogeographical and remote-sensing approaches provide two main strategies for extrapolating local plant-trait measurements across space.
- The study presents and validates a combined remote-sensing and biogeographic approach for spatializing key leaf traits globally.
- The processing chain integrates plant-trait databases, satellite data, climate data, and machine learning to produce global maps at 500 m resolution.
- It maps specific leaf area, leaf dry matter content, leaf nitrogen and phosphorus concentrations, and leaf nitrogen-to-phosphorus ratio.
- Random forests with surrogates first fill gaps in TRY, after which functional-type abundances and community weighted means support global trait-map generation.
2. Materials and methods
The methodology combines heterogeneous trait, remote-sensing, land-cover, and climate data through sequential gap filling, community-level aggregation, and spatial modelling. It addresses sparse and irregular trait observations by estimating pixel-level functional composition and community weighted means before mapping traits globally.
- The workflow uses TRY trait data and multiple remote-sensing and climate products to spatialize in-situ trait estimates.
- TRY contains extensive but heterogeneous measurements, with limited georeferencing, incomplete species coverage, nonstandardized protocols, and substantial within-canopy and across-site variability.
- Gap filling: SLA has around 90,000 measurements from more than 7,000 species, whereas LNPR has 21,000 measurements from approximately 2,000 species, leaving abundant gaps.
- Gap filling: An ensemble of boosted random-forest models with surrogates uses trait relationships and correlations to fill missing database entries.
- Community weighted means: Community weighted means combine trait observations with relative abundances of dominant plant functional types within 500 m MODIS pixels.
- Community weighted means: A 30 m Landsat-based land-cover map estimates functional-type abundances inside each MODIS pixel, using MODIS land cover as reference data for classifier training and validation.
- Community weighted means: Trait observations are filtered by pixel composition and proximity before pixel-specific functional-type trait estimates are averaged.
- Community weighted means: A 100 km maximum distance was selected as a stable heuristic threshold for deriving the community weighted mean estimates.
3. Results
The processing chain was evaluated through gap filling, PFT abundance mapping, and trait prediction. Results showed high gap-filling and classification performance, generally accurate trait predictions, and interpretable spatial patterns with some residual biases.
- 3.1. Gap filling of the TRY database: High gap-filling accuracy was obtained for all five traits, with climate and plant growth form among the five most relevant explanatory variables.The method also incorporated trait-trait covariances and additional explanatory variables beyond taxonomic information and transformed-trait distributions.
- 3.2. Abundance of PFTs at a MODIS pixel level: 96% overall accuracy and a 0.85 Cohen’s Kappa coefficient were achieved for the out-of-sample PFT classification.Shrubs and grasses were the most confounded PFTs, although their corresponding accuracy remained high.
- 3.2. Abundance of PFTs at a MODIS pixel level: The RF-generated classification reproduced the original MODIS data while enabling PFT proportions within 500 m pixels to be estimated from finer-resolution Landsat imagery.The comparison used a selected heterogeneous mountain region in the northern Iberian peninsula.
- 3.3. Precision of the trait prediction models: Trait prediction showed good correlations, virtually no bias, and close agreement between best-fit and one-to-one lines across the evaluated traits.The models were assessed with an independent 20% out-of-sample test set after parameter selection on the remaining data.
- 3.3. Precision of the trait prediction models: Residual ranges were about 10% on average, but some PFTs showed asymmetric residual distributions that may introduce potentially significant biases.Error patterns varied by trait and vegetation type, including higher shrubland errors for LNC and LPC and larger LDMC uncertainties in evergreen needle-leaf forests and grasslands.
- 3.4. Global trait maps visualization: The global maps showed trait patterns consistent with ecological expectations, including high LNPR in equatorial areas and low-to-medium LNPR in temperate and dry regions.SLA values also varied among vegetation types, with the lowest values in boreal needle-leaf evergreen trees and high values in grasslands and savannas.
4. Discussion
The maps generally reproduce expected trait means, ranges, and PFT variability, but comparisons are limited by sparse, spatially biased observations and scale differences between leaf measurements and pixel estimates.
- Independent validation is constrained because few comparable global trait maps are available.
- Trait variability within PFTs reflects intrinsic variation and discrepancies between categorical land-cover classes and continuous vegetation patterns.
- The calculated trait maps generally match expected means and ranges from lookup tables and leaf-level PFT measurements.
- Leaf-level measurements and pixel-level community-weighted estimates can differ because they represent different spatial scales.
- Northern Hemisphere observations dominate the database, with gaps across boreal regions, the tropics, Africa, South America, and Asia.
- Trait distributions need not coincide because database entries are not abundance-weighted, whereas maps represent pixels.
- Differences between latitudinal measurements and maps can identify undersampled areas and help target future in-situ sampling.
5. Summary and conclusions
The paper presents a processing chain for producing 500 m global maps of several leaf traits from in-situ measurements, remote sensing, and climatological information. The chain addresses sparse, irregular trait observations while remaining modular for future data and methodological updates.
- 500 m global maps represent specific leaf area, leaf dry matter content, leaf nitrogen and phosphorus concentrations, and nitrogen-phosphorus ratios.
- The processing chain is modular, allowing maps, remote-sensing and climatic inputs, and processing steps to be improved or updated as new data become available.
- Sparse and irregular global trait observations require spatializing local measurements from leaves to canopy-level community estimates.
- Remote sensing supplies continuous spatial and temporal coverage for scaling leaf measurements to landscapes and regional levels.
- The maps could replace static plant functional type maps and support models of photosynthetic capacity and fluorescence.
Appendix A. Climate data
The study uses WorldClim bioclimatic data because climatological information in the TRY database lacks a consistent structure. These data support both database gap filling and spatialization of pixel-representative trait estimates.
- WorldClim interpolated bioclimatic variables for current conditions provide the climate data used in the study.
- Climate data contribute to two steps: filling gaps in the trait database and spatializing pixel-representative trait estimates.
Appendix B. Description and comparison of machine learning regression methods
This appendix describes the theory behind the machine-learning methods and compares them numerically across precision, fit, bias, and robustness to training-data quantity.
- The appendix summarizes the theory underlying the machine-learning methods used in the study.
- The methods are compared numerically in terms of precision, fit, and bias.
- The comparison also evaluates robustness to the number of training data.
Appendix B.1. Machine learning methods
The appendix describes regularized linear regression, random forests, neural networks, and kernel methods used for plant-trait regression. These methods differ in model structure, flexibility, and computational requirements.
- Regularized linear regression models plant traits as weighted sums of input features while penalizing weight energy to impose smoothness.The method estimates feature weights by least-squares minimization with an added regularization term.
- Random forests predict by averaging many decision-tree outputs built from different feature subsets.This ensemble design helps reduce overfitting and supports heterogeneous, incomplete, large-scale inputs.
- Extreme Learning Machine fixes the network structure and randomly selects input-to-hidden weights, optimizing only hidden-to-output weights by least squares.This substantially reduces the computational cost of neural-network training.
- Kernel methods relate input radiance vectors to plant traits through explicit kernel-based predictive models.The paper considers Kernel Ridge Regression and Gaussian Process Regression; Gaussian processes provide predictive means and variances.
- The kernel prediction combines weighted similarities between a new spectrum and training spectra with a regression bias.The kernel function is parameterized by hyperparameters learned through marginal likelihood maximization.
Appendix B.2. Accuracy and robustness of all considered regression models
All regression methods were trained after feature standardization, with separate data used for hyperparameter selection and out-of-sample testing. Robustness was assessed by varying the amount of cross-validation data.
- 80% of the data formed the cross-validation set for hyperparameter selection, while 20% served as an independent out-of-sample test set.Models were evaluated under the standard split and additional training configurations.
- Cross-validation data proportions were varied from 10% to 80% to assess model robustness.
- All input features were standardized before model training.
Appendix B.2.1. Model evaluation: measuring precision and bias
Model evaluation uses mean error, root-mean-square error, and Pearson’s correlation coefficient to assess bias, precision, and goodness-of-fit. Random forests and kernel methods provide the strongest overall results, with method-specific advantages across traits.
- Model evaluation: Mean error measures estimation bias, root-mean-square error measures precision, and Pearson’s correlation coefficient measures goodness-of-fit.
- Numerical comparison: Random forests and kernel methods are the most precise and least biased methods across the evaluated leaf traits.They outperform regularized linear regression and extreme learning machine in all cases.
- Numerical comparison: Random forests perform best for SLA, LPC, and LNPR, whereas kernel machines perform best for LNC and LDMC.Numerical differences in R and RMSE between random forests and kernel machines are not significant.
- Numerical comparison: Table B.1 reports cross-validation results for all methods, evaluation scores, and leaf traits.The table highlights the best results in bold.
Appendix B.2.3. Models’ robustness to number of training samples
When training data are reduced, random forests and kernel methods retain higher precision than regularized linear regression and extreme learning machine. Random forests show consistent performance across traits and are selected as the default option.
- Models’ robustness to number of training samples: Random forests and kernel machines maintain higher precision than regularized linear regression and extreme learning machine across reduced training-data rates.This pattern appears in both Pearson’s R and RMSE for all evaluated traits.
- Models’ robustness to number of training samples: Random forests show high consistency across plant traits and precision measures under reduced training-data conditions.
- Models’ robustness to number of training samples: The authors select random forests as the preferred default option because of their consistency across traits and precision measures.
- Models’ robustness to number of training samples: Figure B.1 reports test-set results for all methods and leaf traits as a function of the training-data rate.
Appendix C. Sensitivity analysis in the gap filling of the TRY database
The gap-filling sensitivity analysis evaluates predictor importance in surrogate-split random forests and finds taxonomy, especially species and genus, crucial for predicting all traits.
- Predictor importance is calculated by summing split-related changes in MSE across each tree and dividing by the number of branch nodes.
- Because the forests use surrogate splits, importance includes all splits at each branch node, including surrogate splits.
- Table C.1 ranks the five most relevant variables for gap filling, including temperature and precipitation-related bioclimatic variables.Bio1–Bio11 represent temperature variables, while Bio12–Bio19 represent precipitation variables.
- Taxonomy is crucial for gap filling across all traits, with species names and genus as the most influential predictors.The findings confirm the effectiveness of taxonomic hierarchy information for predicting trait values.
Appendix D. Sensitivity analysis in the trait prediction models
The trait prediction sensitivity analysis assesses climatic and remote-sensing predictor importance, while also summarizing trait measurements and their variation across plant functional types.
- The mapping models rank the seven most important climatic and remote-sensing predictors for globally mapped canopy traits.Predictor importance is based on split-related MSE changes normalized by the number of branch nodes, including surrogate splits.
- Table D.1 distinguishes temperature and precipitation variables from MODIS band medians and vegetation-index summary statistics.The listed vegetation-index statistics include maximum, minimum, annual sum, and standard deviation.
- Trait measurements are summarized by mean and standard deviation for each plant functional type using TRY database information to associate species with PFTs.
- Table E.1 reports in-situ leaf-level trait means and standard deviations for each considered plant functional type.