Source-linked AI summary

Profit based evaluation of machine learning for nitrogen recommendations in winter wheat

Xulong Wang, Po Yang

arXiv:2608.27205v1cs.LG

TL;DR

Winter-wheat nitrogen recommendations are made before prices and weather are known, while standard advice ignores prices and ML is usually judged by prediction accuracy. The paper evaluates recommendations by lost profit on 892 measured response curves across price ratios. ML does not replace standard advice profitably at normal prices, but post hoc correction and a damped hybrid improve advice, transfer across sites, and price emission cuts.

  • Problem

    Standard UK nitrogen advice does not respond to price changes, while prediction accuracy does not measure whether ML recommendations are profitable.

  • Method

    The paper scores advice by lost profit on 892 measured yield-response curves from two UK experiments while sweeping the nitrogen-to-grain price ratio.

  • Results

    A post hoc correction yields a 24% improvement and transfers to a second site with a 43% improvement, whereas model and feature changes yield none.

  • Takeaways & Limitations

    ML pays as a profit-scored correction to standard advice, using a price sweep and damped hybrid, rather than as its replacement.

  • Takeaways & Limitations

    Training data come from Broadbalk alone; Woburn is used only to test transfer, and the planned 53-year mean tests fail.

Abstract

from arXiv · show

Nitrogen rates for winter wheat are set before the season, under unknown prices and weather. The standard UK advice does not respond to prices, yet recent price swings moved the most profitable rate by tens of kilograms per hectare. Machine learning is often proposed as the fix. However, it is usually judged on prediction accuracy, and accurate prediction does not by itself make the recommended rate more profitable. Our insight is to score nitrogen advice directly by the profit it forgoes on measured yield response curves. We build a test bench on 892 such curves from two long running UK experiments, and sweep the nitrogen to grain price ratio to cover all price scenarios. On this bench, machine learning fails as a predictor. No model recovers the best rate within farm tolerance, and the benchmark noise shows none can. At normal prices, every model also loses to the standard advice on profit. The gain sits elsewhere. A simple correction step applied after the model cuts profit losses by a quarter, while better models and extra features give no gain. The same frozen correction cuts losses by 43% at the second site without any retraining. A hybrid of standard advice plus a damped correction removes bias and trims rare large losses. The same price sweep also prices emission cuts, at a cost comparable to current carbon prices. Machine learning therefore pays as a profit scored correction to standard advice, not as its replacement.

1 Introduction

Winter-wheat nitrogen advice must account for changing prices, but standard guidance and common machine-learning evaluations do not directly measure profit. This paper evaluates recommendations on measured response curves and finds that machine learning is most useful as a correction to standard advice, not as its replacement.

  • Motivation: UK nitrogen advice is committed before weather and prices are known, while RB209 does not adjust for prices.Ammonium nitrate rose from £234 to £841 per tonne between January 2020 and January 2023, and RB209 required an emergency revision for crisis prices.
  • Motivation: Yield-prediction accuracy can diverge from profitable advice because the best nitrogen rate depends on the response curve’s slope.A prediction with 1.50 t ha−1 error can recommend the best rate, whereas a prediction with 0.24 t ha−1 error can recommend a rate 50 kg too high and lose 0.11 t ha−1 of profit.
  • Approach: The proposed framework scores nitrogen recommendations by forgone profit on 892 measured yield-response curves across a sweep of nitrogen-to-grain price ratios.The curves come from two long-running field experiments, allowing recommendations to be evaluated in the decision-relevant unit under varied price scenarios.
  • Findings: Rate error is ill posed because the benchmark contains about 23 kg N ha−1 of noise, while six machine-learning families lose to standard advice on profit at normal prices.The evaluation therefore distinguishes prediction or rate accuracy from the economic consequence of following a recommendation.
  • Findings: A post hoc correction provides the entire 24% improvement in profit, and a frozen correction transfers to a second site with a 43% improvement.Changing model families or adding features provides no gain; a hybrid with standard advice and a damped correction removes bias and trims rare large losses.
  • Implications: The price-ratio sweep also converts nitrogen optimisation into emission control, with worked abatement costs of £32 to 72 per tonne of CO2e.This connects profit-scored nitrogen advice with pricing emission reductions within the same optimisation framework.

2 Data

The study uses measured nitrogen-response curves from two long-running UK experiments. Broadbalk supplies training data, while Woburn is reserved as an independent transfer test.

  • Broadbalk: Broadbalk contributes 402 quality-screened curves across 53 harvest years from 1968 to 2022.Each curve is fitted through yields at 0, 144, and 288 kg N ha−1.
  • Woburn: Woburn contributes 490 curves across 43 years and is never used for training or tuning.It is a pure test of transfer on lighter soil.
  • Data representation: Nine features accompany each curve, covering weather, previous crop, cultivar era, and field section.Both experiments use the same curve-fitting rules.

3 A profit based test bench

The paper evaluates nitrogen advice by profit loss on held-out measured yield curves across a sweep of price ratios. This exposes benchmark noise that makes a conventional rate-error pass threshold unattainable.

  • Profit in grain terms: Profit loss is the profit at the best rate minus the profit at the advised rate, evaluated on held-out curves across price ratio b.The sweep from b = 3 to 12 brackets observed UK scenarios of 4.1 to 10.7.
  • Why rate error is the wrong score: 23 kg N ha−1 is the benchmark’s median best-rate shift under one-observation refitting, exceeding the 20 kg pass threshold.At b = 5, a 20 kg error costs a median 0.023 t ha−1, while a 30 kg error costs 0.052 t ha−1.
  • Why rate error is the wrong score: Rate error can punish harmless disagreement and hide rare expensive misses, whereas profit loss prices the consequences of both.The paper motivates profit scoring because yield curves are flat near the best rate but can impose larger losses for some misses.

4 Machine learning alone does not clear the bar

Across six model families, machine learning does not outperform RB209 at normal prices when advice is scored by profit loss. The tested gains come from correcting model advice afterward rather than changing models or features.

  • Where the gain comes from: A damped ML correction added to RB209 matches its median loss, removes bias, and trims the worst losses.The comparison uses held-out Broadbalk curves; the hybrid’s small median edge is interpreted as no worse, with gains concentrated in bias and tail behavior.
  • Machine learning alone does not clear the bar: At b ≤6, every ML family has higher median profit loss than RB209; ML falls below RB209 only at b ≥7.The normal-price regime is the one RB209 was built for, while the advantage appears only at crisis prices.
  • Machine learning alone does not clear the bar: 0.070 t ha−1 is ridge’s median loss, versus 0.035 for RB209, while ridge also overshoots the best rate by 25 kg N ha−1.These values are reported for unconstrained ridge advice on held-out curves.

5 The gain lives in the correction step

The tested improvements came from correcting model advice after prediction, not from changing models or features. The correction retained model information while improving the evaluation score.

  • 0.061 t ha−1 was the final correction score, 24% below the TabPFN baseline of 0.080 t ha−1.Model and feature changes gave no reliable gain; the planned mean tests were nonsignificant (Wilcoxon p = 0.70).
  • The only gain that transferred was produced by the post hoc correction group, while model and feature changes produced none.The correction group left the model unchanged and adjusted its advice using damping, shifting, and related tools.
  • Damping all the way towards group averages scored 15% worse than the plain model, so the useful correction preserves field-to-field differences.The tuned correction moves advice only part way towards the group average; the model also supplies information used by the uncertainty cap.
  • The correction does not make the model redundant; the model remains necessary to capture curve slopes and field-to-field differences.The paper’s interpretation is that better use of model output, rather than further modelling effort, generated the improvement.

6 The correction transfers to a new site

A correction trained on Broadbalk transferred unchanged to Woburn, where it reduced losses and failed under scrambled field-group assignments. This supports field-group information as the mechanism rather than generic averaging.

  • 43%: the frozen correction reduced median loss on 490 Woburn curves without retraining, from 0.074 to 0.042 t ha−1 (Wilcoxon p = 0.034).Woburn differs in soil, rotations, and experimental design; the pipeline was applied unchanged.
  • Scrambling field-group labels returned loss to the uncorrected level of 0.074, while all other pipeline components stayed identical.The control mismatched curves and group labels, so advice was pulled towards averages from the wrong groups.
  • The transfer control indicates that the correction uses real agronomic information carried by field groups, not a generic averaging effect.The paper describes this independent transfer as its strongest single piece of evidence.
  • The correction targets yield-curve slope, whereas better prediction can mainly improve yield-curve height without changing the recommended rate.Adding a constant to every yield prediction leaves advice unchanged, explaining why prediction accuracy and advice quality can separate.
  • A price-ratio increase from b = 5 to b = 8 corresponds to a £66 per tonne carbon price and saves about 180 kg CO2e per hectare.The conversion uses a fertiliser-emissions factor of 9.10 kg CO2e per kg N.

7 Profit, prices, and emissions on one dial

The economic case for the hybrid is concentrated in avoiding rare large losses and responding to price-driven nitrogen shifts, with the same price dial also valuing emission reductions. Median gains remain small because profit is flat near the best rate.

  • RB209’s worst 10% of losses reach 1.7 to 3.5 t ha−1, worth £275 to 970 per hectare, while the hybrid reduces large errors from 42 to 37.The hybrid’s median remains level with RB209, so its advantage is reduced tail risk rather than a median-profit premium.
  • At 2023 prices (b ≈10.7), the best rate is about 38 kg N ha−1 below the b = 5 rate, saving about £82 per hectare in fertiliser spend.Profit barely changes because the yield-response curve is flat near the optimum; fixed advice captures none of this shift.
  • Each kilogram of fertiliser N carries about 9.10 kg CO2e, allowing carbon prices to be represented as increases in the effective nitrogen price.At grain £200 per tonne, £66 per tonne CO2e moves the price ratio from b = 5 to b = 8.
  • £32 per tonne of CO2e is the cost of shifting advice from b = 5 to b = 8, which saves about 246 kg CO2e per hectare at a £7.9 per hectare yield cost.A deeper shift to b = 12 costs £72 per tonne of CO2e saved; both estimates are at or below recent UK carbon price levels.
  • The price-response opportunity is worth on the order of £100M in a high-price year across about 1.6 million hectares of UK winter wheat.This estimate concerns the price-response component alone.

8 Limitations •

The study’s limitations concern transfer scope, statistical support for the hybrid’s median edge, rerun variability, default emission factors, and incomplete farm-cost accounting.

  • Training uses data from one site, Broadbalk; Woburn evaluates transfer rather than providing a second training site.
  • The hybrid’s median edge over RB209 is untested statistically, so its case rests on bias reduction and tail-risk control.
  • The correction gain fails the planned mean tests at 53 years; the Woburn transfer carries the statistical weight with p = 0.034.
  • One ledger entry reversed direction between runs and six of thirteen moved, although the final ranking and main conclusion held in every rerun.
  • Emission factors are defaults rather than field measurements, limiting the directness of the emission estimates.
  • Grain terms omit application costs, protein premiums, and storage, so pound conversions are rescalings rather than farm budgets.

9 Conclusion

The paper concludes that machine learning pays for nitrogen advice only as a profit-scored, price-responsive correction to standard advice, not as its replacement.

  • Machine learning is a conditional decision aid when scored in profit on measured curves, applied as a damped correction to standard advice, and evaluated across price ratios.
  • Under these conditions, bias falls from 6 to 1 kg N ha−1, rare large losses shrink, and the correction transfers unchanged with 43% improvement at a new site.
  • The same price dial values emission cuts at £32 to 72 per tonne.

Reproducibility

The evaluation used locked truth tables and held-out years, with materials planned for release.

  • The evaluation used locked truth tables and held out years throughout.
  • Curve tables, ledgers, per-curve losses, and scripts will be released.

Data availability

The study’s two field datasets are open-access records from the electronic Rothamsted Archive.

  • Both field datasets are open access under a CC BY 4.0 licence.
  • The datasets are Broadbalk wheat yields from 1968 to 2022 and Woburn Ley arable wheat yields from 1976 to 2018.
  • Both datasets are published by the electronic Rothamsted Archive at Rothamsted Research, Harpenden, UK.
Loading 2608.27205v1…