Source-linked AI summary
Maize Yield and Nitrate Loss Prediction with Machine Learning Algorithms
Mohsen Shahhosseini, Rafael A. Martinez-Feria, Guiping Hu, Sotirios V. Archontoulis
TL;DR
Simulation models support scenario planning but require substantial expertise and data and can be slow to run. This study evaluated five ML meta-models trained on more than three million APSIM scenarios, finding that they predicted maize yield reasonably well but not N loss at planting time.
Problem
Simulation crop modeling requires substantial expertise and data, while scenario analysis can be impractical because of long runtimes and storage constraints.
Method
The study compared five ML meta-models using more than three million simulated genotype, environment, and management scenarios with pre-season soil and weather inputs.
Results
Random forests and XGBoost were the strongest meta-models overall, while ML predicted yield reasonably but not N loss; yield error decreased 10%–40% with larger training datasets, whereas N-loss error showed no consistent pattern.
Takeaways & Limitations
ML meta-models can inform pre-season maize-yield prediction and future decision-support tool development, but N-loss prediction requires a different approach and likely more in-season information.
Takeaways & Limitations
Annual N loss cannot be reliably predicted from information available up to planting time because its prediction error was about four times higher than yield error.
Abstract
from arXiv · showhide
Pre-season prediction of crop production outcomes such as grain yields and N losses can provide insights to stakeholders when making decisions. Simulation models can assist in scenario planning, but their use is limited because of data requirements and long run times. Thus, there is a need for more computationally expedient approaches to scale up predictions. We evaluated the potential of five machine learning (ML) algorithms as meta-models for a cropping systems simulator (APSIM) to inform future decision-support tool development. We asked: 1) How well do ML meta-models predict maize yield and N losses using pre-season information? 2) How many data are needed to train ML algorithms to achieve acceptable predictions?; 3) Which input data variables are most important for accurate prediction?; and 4) Do ensembles of ML meta-models improve prediction? The simulated dataset included more than 3 million genotype, environment and management scenarios. Random forests most accurately predicted maize yield and N loss at planting time, with a RRMSE of 14% and 55%, respectively. ML meta-models reasonably reproduced simulated maize yields but not N loss. They also differed in their sensitivities to the size of the training dataset. Across all ML models, yield prediction error decreased by 10-40% as the training dataset increased from 0.5 to 1.8 million data points, whereas N loss prediction error showed no consistent pattern. ML models also differed in their sensitivities to input variables. Averaged across all ML models, weather conditions, soil properties, management information and initial conditions were roughly equally important when predicting yields. Modest prediction improvements resulted from ML ensembles. These results can help accelerate progress in coupling simulation models and ML toward developing dynamic decision support tools for pre-season management.
1. Introduction
Machine-learning meta-models are proposed to provide faster, more flexible pre-season predictions than detailed crop simulations, addressing simulation requirements and runtime constraints. The study evaluates their performance, training-data needs, input-variable importance, and ensemble benefits for maize yield and N-loss outcomes.
- Motivation: Simulation models require substantial expertise and data, while scenario analysis is constrained by long runtimes and storage needs.Simulations must also be rerun when new information becomes available or when extrapolating beyond originally simulated conditions.
- Motivation: Meta-models learn relationships from computationally expensive simulations and can provide faster execution, reduced storage needs, and greater flexibility across spatial and temporal scales.The paper investigates ML algorithms as meta-models for crop-production decision-support systems.
- Study aim: The study targets pre-season predictions of maize yield and N loss so farmers can access production and environmental-quality information before planting.The proposed framework aims to support robust, fast, and dynamic forecasting systems when information is most needed.
- Study objectives: The study evaluates four ML dimensions: algorithm performance, training-data requirements, input-data importance, and whether ensembles outperform individual algorithms.These objectives cover prediction accuracy, data needs, variable ranking, and ensemble performance.
2. Materials and methods
The study used APSIM simulations across diverse Midwest sites and factorial management, soil, weather, and initial-condition scenarios to create a large dataset for testing machine-learning meta-models. Models were trained and evaluated with pre-season inputs using time-wise validation and multiple performance and sensitivity analyses.
- APSIM was calibrated with experimental maize yield and drainage N-loss data from seven US Midwest locations before generating simulation scenarios.
- More than 3 million scenarios combined crop, soil, nitrogen-management, weather-year, and initial-condition factors from 1983–2016.Each simulation represented a full factorial combination, with annual runs reset on 20 October.
- Only pre-season weather from approximately mid-October to April was used, excluding growing-season weather unavailable at planting time.Weather features included temperature and precipitation summarized across five fallow-period intervals.
- Four machine-learning meta-models—Ridge, LASSO, random forests, and XGBoost—were evaluated alongside multiple linear regression.
- Performance was assessed with RMSE, RRMSE, R2, and mean bias error, while permutation importance and partial dependence examined input sensitivity and marginal effects.
- The dataset was split into 2.7 million training observations from 1983–2012 and 0.4 million hold-out test observations from 2013–2016.Time-wise 5-fold look-forward cross-validation tuned hyperparameters while preserving temporal structure.
3. Results
Tree-based meta-models generally performed best, with yield predicted more accurately than N loss. Results also showed model-specific responses to training-data size and input variables, while ensembles provided additional improvements.
- XGBoost and random forests outperformed the other meta-models across locations for maize yield and N-loss prediction.
- 13.9% RRMSE was achieved for random-forest yield prediction, compared with 54.5% RRMSE for N loss.Random forests had R2 values of 0.44 for yield and 0.78 for N loss.
- Weather was the most important average input for yield prediction, followed by management and soil properties, although model sensitivities differed substantially.Random forests and XGBoost were especially sensitive to temperature and rainfall features.
- Later planting dates were associated with lower predicted maize yields, while higher initial soil nitrate produced higher predicted nitrate loss.These relationships were observed in partial dependence plots for random forests and XGBoost.
- Weather explained 43% of variance in random-forest N-loss predictions, while initial conditions explained 38%.Other models also generally ranked weather and initial conditions among the most important inputs.
- Weighted ensembles improved maize-yield and nitrate-loss prediction over the best individual meta-model, with the optimal ensemble reaching 12.3% yield RRMSE.
- Yield prediction error decreased as training data increased to approximately 1.6 million observations, after which additional data produced little benefit.XGBoost was most sensitive and random forests least sensitive to training-dataset size.
4. Discussion
The ML meta-models reproduced maize yield more reliably than nitrate loss, while performance varied with training-data size, input variables, locations, and ensemble design. These results inform data collection and model selection for faster crop decision-support systems.
- N loss prediction showed no consistent relationship with training-data size: regression models benefited, whereas random forests and XGBoost were negatively affected.
- Weather was the most important averaged input feature, while model sensitivities differed across weather, management, soil, and initial-condition variables.
- 13%–14% RRMSE for end-of-season yield was comparable to the simulation model’s field-data fit using only information available through planting.
- N loss RRMSE was about four times higher than yield error, indicating that annual N loss cannot be reliably predicted from planting-time information alone.
- Selecting only the most influential features did not improve prediction, with random-forest N loss RRMSE reaching 85.4% using ten features and 95.5% using five.
- Optimized ensembles improved yield prediction over the best single model, reducing RRMSE from 13.4% to 12.3%.
ORCID iDs
The paper lists ORCID identifiers for its authors.
- The authors’ ORCID identifiers are provided for Mohsen Shahhosseini, Rafael A. Martinez-Feria, Guiping Hu, and Sotirios V. Archontoulis.
Data availability statement
The supplied passages consist primarily of bibliographic references and a data availability statement.
- The references cover machine learning, crop modeling, nitrogen loss, sensitivity analysis, ensembles, and agricultural decision support.
- The article states that data supporting the findings are included within the article.