Source-linked AI summary
WeedMap: A large-scale semantic weed mapping framework using aerial multispectral imaging and deep neural network for precision farming
Inkyu Sa, Marija Popovic, Raghav Khanna, Zetao Chen, Philipp Lottes, Frank Liebisch, Juan Nieto, Cyrill Stachniss, Achim Walter, Roland Siegwart
TL;DR
Large UAV vegetation maps must preserve fine plant details despite DNN memory and resolution constraints. The paper uses aligned, radiometrically calibrated multispectral orthomosaic tiles processed with a sliding window, achieving higher AUC with nine channels than an RGB SegNet baseline and releasing an expert-labeled dataset.
Problem
Large-scale crop/weed mapping is limited by DNN memory, downsampling-related resolution loss, and multispectral alignment challenges.
Method
The framework processes channel-aligned, radiometrically calibrated multispectral orthomosaic tiles sequentially with a sliding window, matching tile size to the DNN input.
Results
The nine-channel model achieved AUC [bg = 0.839, crop = 0.863, weed = 0.782], outperforming the RGB SegNet baseline [0.607, 0.681, 0.576].
Takeaways & Limitations
The approach generates complete field maps and provides an expert-labeled multispectral sugar beet/weed dataset for further research.
Takeaways & Limitations
Reliable spatiotemporal modeling remains challenging because plant appearance changes across maturity stages and weed species are difficult to representatively capture.
Abstract
from arXiv · showhide
We present a novel weed segmentation and mapping framework that processes multispectral images obtained from an unmanned aerial vehicle (UAV) using a deep neural network (DNN). Most studies on crop/weed semantic segmentation only consider single images for processing and classification. Images taken by UAVs often cover only a few hundred square meters with either color only or color and near-infrared (NIR) channels. Computing a single large and accurate vegetation map (e.g., crop/weed) using a DNN is non-trivial due to difficulties arising from: (1) limited ground sample distances (GSDs) in high-altitude datasets, (2) sacrificed resolution resulting from downsampling high-fidelity images, and (3) multispectral image alignment. To address these issues, we adopt a stand sliding window approach that operates on only small portions of multispectral orthomosaic maps (tiles), which are channel-wise aligned and calibrated radiometrically across the entire map. We define the tile size to be the same as that of the DNN input to avoid resolution loss. Compared to our baseline model (i.e., SegNet with 3 channel RGB inputs) yielding an area under the curve (AUC) of [background=0.607, crop=0.681, weed=0.576], our proposed model with 9 input channels achieves [0.839, 0.863, 0.782]. Additionally, we provide an extensive analysis of 20 trained models, both qualitatively and quantitatively, in order to evaluate the effects of varying input channels and tunable network hyperparameters. Furthermore, we release a large sugar beet/weed aerial dataset with expertly guided annotations for further research in the fields of remote sensing, precision agriculture, and agricultural robotics.
1. Introduction
UAV multispectral orthomosaics support large-area vegetation mapping while preserving plant-level detail, but their size creates resolution and memory challenges for DNN processing. The paper addresses this with aligned, calibrated maps, sliding-window tiles, and a complete weed-mapping system and dataset.
- UAV multispectral imaging provides high-resolution remote-sensing data for vegetation monitoring and potentially supports site-specific weed management.Early weed detection can support herbicide savings, reduced environmental impact, and increased crop yield.
- Large-area applications require maps covering hectares while preserving fine plant-distribution details for subsequent weed-management actions.
- Orthomosaic maps provide georeferenced metric-scale representation, aligned multispectral channels, and globally calibrated reflectance maps.
- GPU memory limits make huge orthomosaics difficult to process without resolution loss, motivating a sliding-window technique that processes small map portions.
- The paper presents a weed-mapping system for orthomosaics covering more than 16,500 m2 and releases expertly guided sugar beet/weed aerial datasets.
2. Related Work
UAV remote sensing and DNN-based segmentation have advanced crop/weed mapping, but complex agro-ecosystems and limited single-image processing remain important challenges. Prior work spans handcrafted classifiers, CNN pipelines, multispectral inputs, and the proposed multi-channel approach for more complete weed maps.
- UAV Remote Sensing: UAV remote sensing supports high-resolution vegetation mapping and precision-farming applications.Research in this area covers plant detection, classification, and dense semantic segmentation using DNNs.
- Traditional Machine Learning: 75–87% overall classification accuracy was achieved by a multispectral crop, weed, and soil patch classifier using pixel intensities and crop-row geometry.The system categorized image patches into distinct crop, weed, and soil classes.
- Traditional Machine Learning: 95%+ classification accuracy was reported for invasive grasses and vegetation using RGB imagery, a decision tree, and handcrafted features.Other work reported 96% overall accuracy for crop-versus-weed detection and up to 86% for crop-versus-multiple-weed-species segmentation.
- Motivation for DNNs: Local variation in environments, soil, and crop or weed species makes agro-ecosystems difficult to characterize with handcrafted features and conventional machine learning.These systems are challenged by the multivariate, complex, and unpredictable nature of agricultural ecosystems.
- DNN-Based Segmentation: Dense semantic segmentation assigns human-interpretable labels to every image pixel, with CNNs forming the dominant approach.Earlier CNN segmentation methods commonly used region proposals followed by a sub-network that inferred labels for each proposal.
- Multispectral CNNs: Most prior systems processed only one RGB image at a time, whereas the proposed approach handles multi-channel inputs to produce more complete weed maps.Related CNN systems also used RGB and NIR imagery with sliding-window, pixel-wise classification pipelines.
3. Methodologies
The methodology combines multispectral UAV data collection, calibrated orthomosaic generation, and resolution-preserving tiled DNN inputs for large-scale crop/weed mapping.
- 3.1. Data Collection Procedures: Eight multispectral orthomosaic maps were collected across sugar beet fields using UAV campaigns and separately operated aerial platforms.The datasets cover fields in Eschikon and Rheinbach, with flight paths recorded at similar altitudes and times of day.
- 3.1. Data Collection Procedures: The dataset covers 1.6554 ha and contains 1.76 billion pixels across 10,196 images, with 1.39 billion training and 367 million testing pixels.RedEdge-M data provide 12 channels and Sequoia data provide eight channels, including RGB, CIR, NDVI, and sensor bands.
- 3.2. Dataset Processing: The DNN receives tiles matching the input resolution, avoiding down-sizing that can discard visual information needed to distinguish small crops and weeds.The maps achieve approximately 1 cm GSD; crops occupy about 15–20 pixels and weeds about 5–10 pixels at high zoom.
- 3.1. Data Collection Procedures: Orthomosaic processing composes RGB, CIR, and NDVI channels from aligned multispectral imagery before DNN input construction.Each channel is treated as an image, allowing the composed and individual channels to be processed independently by subsequent convolution layers.
- 3.4. Dense Semantic Segmentation: The tiling framework crops fixed-size images from aligned orthomosaic maps and feeds 12 composited tile channels into the DNN.This stand sliding-window approach addresses the memory limitations of processing complete high-resolution orthomosaic maps directly.
- 3.3. Orthomosaic Reflectance Maps: Radiometric calibration converts raw multispectral pixels into band reflectance using calibration-panel reflectance, radiance, exposure, gain, black level, and vignette corrections.The resulting reflectance maps are globally calibrated for consistent illumination and vignette compensation across input images.
4. Experimental Results
The experiments evaluate crop/weed segmentation across datasets while varying input channels and network hyperparameters.
- The experiments investigate classifier performance under different input-channel configurations and network hyperparameters.
4.1. Experimental Setup
The study uses annotated multispectral orthomosaic datasets from two cameras, with separate training and testing splits because their bands are not matched.
- Eight multispectral orthomosaic maps are paired with manually annotated labels for three classes: bg, crop, and weed.
- RedEdge-M uses datasets [000, 001, 002, 004] for training and 003 for testing, while Sequoia uses [006, 007] for training and 005 for testing.
- The camera datasets are kept separate because their multispectral bands differ in center wavelength, bandwidth, and sensor sensitivity.
- Training uses learning rate = 0.001, max. iterations = 40,000, momentum = 9.9, weight decay = 0.0005, and gamma = 1.0.Inputs are horizontally mirrored for two-fold data augmentation.
4.2. Performance Evaluation Metric
Performance is evaluated with AUC from precision-recall curves, using pixel-wise class probabilities and threshold-based confusion counts.
- AUC of a precision-recall curve is used for performance evaluation.
- AUC avoids selecting one optimal threshold by evaluating precision and recall across thresholds.
- The network outputs 480 × 360 × 3 pixel-wise class probabilities for the three defined classes.
- True-positive, true-negative, false-positive, and false-negative counts are obtained after thresholding class probabilities.
- Alternative segmentation metrics rely on specific thresholds or maximum-probability labels when comparing predictions with ground truth.
4.3. Results Summary
Twenty models are compared across two cameras, input-channel choices, batch sizes, and class balancing, revealing strong effects from multispectral information and dataset conditions.
- Results Summary: Twenty models vary in input channels, batch size, class balancing, and class-specific AUC across the RedEdge-M and Sequoia datasets.
- Results Summary: The RedEdge-M evaluation identifies RGB-only SegNet as the baseline and marks the best model separately from a one-NDVI-input model.
- Results Summary: Model 5 performs best with nine input channels, while Model 1, despite using all available input data, slightly underperforms it by less than 2%.
- Results Summary: Larger batch size yields better results, but GPU memory limits the maximum practical batch size to five in most training procedures.
- Results Summary: NDVI contributes substantially to vegetation classification, and Model 12 substantially outperforms Model 4 when the former uses only NDVI and the latter excludes it.
- Results Summary: Performance curves include sharp points when precision and recall remain unchanged across thresholds; the same perfcurve rule is applied for fair comparison.
- Results Summary: Sequoia results confirm that NDVI plays a significant role in crop/weed detection, while smaller crop and weed instances produce a 10% worse weed-detection performance.
- Results Summary: Class balancing has dataset-dependent effects: Model 14 significantly outperforms Model 15 without balancing, unlike the RedEdge-M comparison between Models 1 and 3.
4.4. Qualitative Results
Qualitative evaluation shows accurate crop-row predictions but weaker weed classification, especially at fine resolution and across unseen fields. RedEdge-M and Sequoia exhibit similar trends.
- The best models achieved AUC values of [0.839, 0.863, 0.782] for RedEdge-M and [0.951, 0.957, 0.621] for Sequoia.Model 5 was used for RedEdge-M and Model 16 for Sequoia.
- RedEdge-M Analysis: Crop classification performed reasonably, with crop rows clearly visible and magnified views showing visually accurate predictions.
- RedEdge-M Analysis: Weed classification produced more false positives and false negatives than crop classification.
- RedEdge-M Analysis: Wide views estimated weed distributions and densities consistently, with high precision but low recall in weed-dense field regions.
- RedEdge-M Analysis: Weed performance was limited by small weed footprints, class imbalance, and testing fields unseen during training, suggesting possible overfitting or insufficient weed variation.
- Sequoia Analysis: Sequoia showed similar good-crop and relatively poor-weed predictions, with smaller plant footprints, growth-stage variation, and fewer weed pixels than crop pixels.
5. Discussion on Challenges and Limitations
The discussion identifies challenges in generalizing weed mapping across growth stages and fields, while noting practical runtime and downstream agricultural applications. Data augmentation may help, but can be counter-productive when classes look similar.
- Reliable spatiotemporal models spanning plant maturity stages and multiple farm fields remain challenging because plant appearance changes during growth.
- Early-season crop–weed similarity and weeds’ diverse species, shapes, sizes, and appearances make representative high-quality spatiotemporal datasets difficult to obtain.
- More aggressive augmentation, including scaling and random rotations, could improve classification but may be counter-productive when target classes are visually similar.
- Forward inference takes about 200 ms per input image on a NVIDIA Titan X, while total map generation depends on the number of orthomosaic tiles.RedEdge-M testing took 18.8 s and Sequoia testing took 42 s; post-processing time was omitted.
- Generated weed maps can support prescription maps transferred to fertilizer or herbicide boom sprayers.The stated application is intended to minimize chemical usage and labor cost while maintaining agricultural productivity.
6. Conclusions
The paper presents a complete multispectral, DNN-based pipeline for large-scale semantic weed mapping at approximately 1 cm GSD. It reports stronger nine-channel performance than RGB and releases an expert-labeled dataset, while weed segmentation remains limited.
- The pipeline tiles 16,550 m2 multispectral orthomosaics into DNN-sized inputs, preserving approximately 1 cm GSD for sequential sliding-window crop/weed classification.
- The approach generates complete field maps that can be exploited for site-specific weed management.
- Nine-channel input achieved AUC [bg = 0.839, crop = 0.863, weed = 0.782], versus [0.607, 0.681, 0.576] for the RGB SegNet baseline.The analysis varied input channels and network hyperparameters; NDVI significantly helped distinguish crops and weeds.
- The released datasets contain high-resolution multispectral sugar beet/weed imagery with expert labeling for supervised research.
- Weed segmentation remains limited because weed instances are small and vary naturally in shape, size, and appearance.
Abbreviations
The manuscript defines abbreviations used for sensors, imaging concepts, neural networks, positioning systems, and agricultural applications.
- UAV means Unmanned Aerial Vehicle, while DNN means Deep Neural Network and CNN means Convolutional Neural Network.
- NDVI means Normalized Difference Vegetation Index, NIR means Near-Infrared, and GSD means Ground Sample Distance.
- SSWM means Site-Specific Weed Management, and DSM means Digital Surface Model.