Source-linked AI summary
A Dataset-Centric Benchmark of Deep Learning Methods for Grape Leaf Disease Classification and Detection
Petar Canoski, Vlatko Spasev, Ivica Dimitrovski, Ivan Kitanovski, Petre Lameski
TL;DR
Grape leaf disease recognition lacks evaluation evidence that adequately reflects heterogeneous vineyard conditions and dataset differences. This paper builds a dataset-centric benchmark across classification and detection settings, finding near-saturated performance on several controlled datasets but sharp cross-dataset degradation, especially for detection. The results support treating provenance, field realism, and annotation compatibility as central to interpreting model performance.
Problem
Many studies rely on controlled datasets and image-level classification, leaving limited evidence about recognition under heterogeneous field conditions and across differing annotation tasks.
Method
The study analyzes dataset properties and evaluates representative models in image-level classification, region-level classification, and object detection, including cross-dataset experiments.
Results
The benchmark finds near-saturated within-dataset classification on several controlled or derivative datasets, while cross-dataset transfer degrades sharply and detection varies substantially with dataset and annotation definition.
Takeaways & Limitations
Shared disease labels and high within-dataset performance do not ensure transferable recognition tasks, so benchmark interpretation requires attention to provenance, visual domain, and annotation protocols.
Takeaways & Limitations
Region-level accuracy reflects discrimination among manually localized crops rather than complete-image disease recognition or automatic disease localization.
Abstract
from arXiv · showhide
Grape leaf disease recognition is important for precision agriculture, enabling early diagnosis, timely intervention, and improved vineyard management. Although deep learning has achieved strong results, many studies rely on few datasets, often acquired under controlled conditions, and may not reflect real vineyard challenges such as complex backgrounds, variable illumination, occlusion, leaf pose, disease severity, and device differences. This paper presents a dataset-centric benchmark of deep learning methods for grape leaf disease classification and detection. We analyze publicly available datasets in terms of disease categories, annotation types, acquisition conditions, image characteristics, class distributions, provenance, and task suitability. Representative models are evaluated in three settings: image-level classification, region-level classification, and object detection. Classification is assessed using accuracy, while detection is evaluated using mAP@50 and mAP@50:95. Cross-dataset experiments further examine transfer between datasets with compatible disease categories but different visual and annotation characteristics. Results show near-saturated classification performance on several controlled or derivative datasets, greater difficulty on heterogeneous datasets, and substantial variation in detection performance across annotation settings. Cross-dataset performance drops sharply, especially for object detection, indicating that shared disease labels do not necessarily define equivalent recognition tasks. The benchmark emphasizes dataset provenance, realistic field evaluation, annotation compatibility, and external validation for reliable vineyard disease recognition.
1. Introduction
Grapevine disease recognition matters for precision viticulture, but controlled datasets and image-level evaluation may not represent complex field conditions. This paper addresses the gap with a dataset-centric benchmark spanning complementary tasks, datasets, and annotation settings.
- Motivation: Manual vineyard diagnosis is time-consuming, subjective, labor-intensive, and difficult to apply consistently across large areas.Reliability can also vary with observer experience, disease severity, symptom similarity, illumination, occlusion, and spatially variable infection.
- Dataset challenges: Controlled datasets with simple backgrounds and stable illumination do not fully represent field images containing complex backgrounds, overlapping leaves, variable lighting, shadows, and motion blur.
- Task and annotation suitability: Object detection better suits images with multiple leaves or localized symptoms, but detection datasets require clear target definitions and comprehensive annotations that are costly to produce.Public detection datasets are less common and often differ in annotation schemes, disease categories, acquisition conditions, and target definitions.
- Benchmark design: It evaluates representative models for image-level classification, region-level classification, and object detection using accuracy, mAP@50, and mAP@50:95 under consistent task-specific protocols.Cross-dataset experiments examine transfer after harmonizing compatible disease categories.
- Benchmark design: The benchmark analyzes publicly available datasets by disease taxonomy, acquisition conditions, image characteristics, annotation granularity, class distribution, provenance, and task suitability.
- Research objective: The study focuses on how dataset properties affect evaluation outcomes, cross-dataset robustness, and the interpretation of recognition performance rather than simply ranking architectures.It jointly considers acquisition conditions, provenance, class balance, label compatibility, annotation granularity and completeness, and domain shift.
2. Related Work
Related work progresses from handcrafted, segmentation-dependent pipelines to CNN and transfer-learning methods, which often achieve high within-dataset classification performance. However, inconsistent task terminology, dataset bias, provenance overlap, and heterogeneous annotation protocols limit comparability and cross-dataset generalization, motivating a dataset-centric benchmark.
- Task definitions and detection: Grape disease terminology is inconsistent, so the benchmark distinguishes image-level classification, region-level classification, and object detection by their outputs and localization requirements.Image-level classification assigns one label to a complete image, region-level classification evaluates a manually localized crop, and object detection predicts locations and labels jointly.
- Classical image processing and machine learning: Classical grape disease systems combined segmentation, handcrafted color or texture features, and machine-learning classifiers in multistage pipelines.Representative approaches used color segmentation, diseased-region extraction, clustering, feature selection, and classifiers such as SVM, AdaBoost, and random forest.
- Classical image processing and machine learning: Classical methods demonstrated automated recognition but depended strongly on acquisition conditions, preprocessing quality, segmentation accuracy, and manually selected features.These dependencies motivated the transition toward CNN-based and transfer-learning approaches.
- CNN and transfer learning methods: CNN and transfer-learning studies generally report high classification performance on grape disease datasets, especially the four-class PlantVillage grape setting.The common classes are healthy, black rot, Esca or black measles, and leaf blight; some work extends the setting with phylloxera.
- CNN and transfer learning methods: Most deep-learning evaluations remain within a single dataset, limiting conclusions about generalization to different acquisition conditions, external datasets, or field images.High accuracy on controlled or semi-controlled datasets can therefore coexist with limited evidence of real-world robustness.
- Dataset bias and benchmark motivation: Dataset provenance can bias reported generalization when repackaged, augmented, renamed, or derivative datasets share exact or near-duplicate images.Overlap between training and evaluation data can produce information leakage and overly optimistic estimates.
- Task definitions and detection: Spatial annotations support different tasks because selected regions may enable region-level classification, whereas comprehensive instance coverage is needed for quantitative object-detection evaluation.Annotation semantics and completeness matter in addition to whether annotations are boxes or masks.
- Dataset bias and benchmark motivation: The benchmark addresses a gap by jointly studying image-level classification, manually localized-region classification, complete-image detection, dataset overlap, and cross-dataset transfer.It treats datasets as experimental factors affecting task formulation, comparability, and generalization.
3. Public Datasets for Grape Leaf Disease Recognition
The benchmark treats publicly available datasets as experimental factors, distinguishing their taxonomies, acquisition conditions, annotation units, class distributions, and suitability for image classification, region classification, or detection. These datasets vary substantially in visual composition, scale, balance, provenance, and annotation semantics, requiring dataset-specific interpretation and compatible cross-dataset comparisons.
- The benchmark analyzes datasets by taxonomy, annotation granularity and completeness, acquisition conditions, preprocessing, class balance, provenance, and task suitability.
- Image-level classification uses complete images, region-level classification uses crops from manually localized annotations, and object detection predicts categories and locations in complete images.The presence of boxes or masks alone does not determine the appropriate task.
- Classification datasets differ substantially in acquisition conditions, image dimensions, preprocessing, label taxonomy, dataset scale, and predefined data splits.
- GVLiD and GL-Portugal represent natural vineyard conditions most directly, whereas NGLD is field-derived but standardized through resizing and background removal.NGLD contains 2,726 images, while GVLiD contains 3,477 and GL-Portugal 5,267 images.
- Most classification datasets use a common four-class taxonomy, while NGLD, GLHCD, and GL-Portugal add partially overlapping disease and mite categories.Alternative taxonomies broaden disease coverage but limit direct comparison to compatible class subsets.
- The HERMOS region dataset contains 5,060 healthy, 1,882 dead-arm, 1,392 downy-mildew, and 5,518 powdery-mildew regions, so region counts are class-imbalanced and differ from source-image counts.
- Detection datasets differ in annotation scale, class overlap, target semantics, and class balance, spanning local symptoms, disease-related leaves, and complete-leaf or confounder categories.Mildew symptom and LDD share downy mildew and powdery mildew, whereas FD-Confounders uses a distinct taxonomy.
4. Benchmark Methodology
The benchmark evaluates grape leaf disease recognition through image-level classification, region-level classification, and object detection under distinct input and annotation settings. These tasks are reported separately because they address different practical objectives and spatial information.
- Evaluation settings: The benchmark evaluates image-level classification, region-level classification, and object detection as three complementary recognition settings.The settings use complete images, manually localized regions, or complete images with predicted spatial targets, respectively.
- Image-level classification: Image-level classification assigns each complete image to a disease or healthy category from a dataset-specific class set.Possible categories include healthy leaves and diseases such as black rot, Esca, leaf blight, downy mildew, and powdery mildew.
- Region-level classification: Region-level classification extracts annotated regions from source images and predicts their region labels after localization has been provided.The benchmark partitions data at the original-image level before crop extraction to prevent information leakage.
- Object detection: Object detection predicts bounding boxes, class labels, and confidence scores for diagnostically relevant leaves, lesions, spots, or symptomatic regions.Only datasets with sufficiently consistent and comprehensive spatial annotations are included, with segmentation polygons converted to enclosing boxes when available.
- Task interpretation: The three settings are interpreted separately because they differ in input unit, annotation type, and evaluation objective.Object detection addresses localization and classification together, whereas region-level classification evaluates discrimination using manually supplied localization.
4.2. Classification Models
The benchmark compares representative lightweight and moderately sized classification architectures across controlled experimental protocols. The model set spans residual, efficient, mobile, modern convolutional, transformer, and hybrid CNN–Transformer families.
- Model selection: The classification benchmark evaluates representative lightweight and moderately sized architectures rather than exhaustively testing all available classifiers.The selected models are drawn from the timm library and provide coverage across several architectural families.
- Architectural coverage: The evaluated families include residual CNNs, efficient CNNs, mobile-oriented CNNs, modern convolutional networks, vision transformers, and hybrid CNN–Transformer models.Table 5 summarizes these models together with approximate parameter counts, GMACs, and intended comparative roles.
- Residual baselines: ResNet-18 and ResNet-50 provide residual baselines for examining how increased residual-network capacity affects classification performance.ResNet-18 is the lower-cost reference, while ResNet-50 is the higher-capacity residual architecture.
- Efficient models: EfficientNet-B0, EfficientNet-B3, and MobileNetV3-Large represent efficiency-oriented alternatives for accuracy and resource-constrained inference.EfficientNet-B0 is compact, EfficientNet-B3 has higher capacity, and MobileNetV3-Large targets mobile and embedded devices.
- Modern and transformer models: ConvNeXt-Tiny, ViT-S/16, DeiT-S/16, Swin-T, and MobileViT-S extend the comparison to modern convolutional, transformer, and hybrid designs.These models examine alternatives involving patch-based self-attention, data-efficient transformer training, hierarchical attention, or combined convolutional and transformer representations.
- Experimental control: All architectures use identical partitions, preprocessing, augmentation, optimization, checkpoint selection, and evaluation metrics within each classification setting and dataset.This protocol controls within-dataset comparisons while allowing dataset conditions and provenance to be examined separately from architecture.
4.3. Qualitative Classification Interpretation
Grad-CAM provides a qualitative view of the spatial evidence contributing to selected classification predictions. The analysis illustrates model behavior without serving as segmentation ground truth or an additional quantitative metric.
- Grad-CAM interpretation: Grad-CAM weights convolutional feature maps using target-class gradients to produce a coarse class-discriminative activation map.The map is overlaid on the input image to show regions contributing most strongly to the selected class score.
- Visualization setup: ResNet-50 is used for the qualitative visualization because its final convolutional stage provides a direct target for Grad-CAM.The visualization is independent of quantitative model ranking and is applied to a correctly classified GL-Portugal test image.
- Scope: The activation map is not treated as symptom segmentation, localization ground truth, or an additional quantitative evaluation metric.Its role is limited to qualitative illustration of model behavior.
4.4. Object Detection Models
The object-detection benchmark compares lightweight YOLO detectors across generations and capacities with a transformer-based RF-DETR alternative. Its design separates architectural evolution, capacity scaling, and prediction-paradigm comparisons.
- Detector set: The detector set comprises YOLOv8n, YOLO11n, YOLO26n, YOLO26s, and RF-DETR-Nano.The models compare detector generation, capacity, and architectural paradigm.
- Architectural differences: The YOLO models retain a backbone–neck–head organization, while their generations differ in feature blocks, attention, box regression, and inference procedures.YOLOv8 uses C2f modules and an anchor-free decoupled head; YOLO11 adds C3k2 blocks and C2PSA attention; YOLO26 changes the detection pipeline and removes Distribution Focal Loss.
- Training design: YOLO26 includes progressive loss weighting and small-target-aware label assignment, but their practical benefit is evaluated empirically rather than assumed.These mechanisms are relevant because several datasets contain numerous small, spatially distributed symptom regions.
- Capacity scaling: YOLO26n and YOLO26s isolate model-capacity effects within the same detector generation.The small variant increases network width and depth to provide greater feature capacity at approximately four times the parameter count and computational complexity.
- Prediction paradigms: RF-DETR-Nano contrasts YOLO-style dense multiscale prediction with transformer-based query-driven set prediction.It uses a DINOv2 vision-transformer backbone, object queries, and a fixed set of predictions.
- Generation comparison: YOLOv8n, YOLO11n, and YOLO26n enable a broadly size-matched comparison across three YOLO generations.This tests whether newer architectural changes improve disease detection without substantially increasing model size.
4.5. Dataset Provenance and Image-Overlap Analysis
The study checks whether datasets share identical visual content before interpreting cross-dataset transfer. Exact decoded-image hashing identifies overlap despite differences in filenames, directory structures, or file encodings.
- Exact decoded-image hashing detects identical visual content across datasets despite different filenames, directory structures, or file encodings.
- Overlap analysis distinguishes independent source-to-target evaluation from transfer involving shared image content.
- Detected overlaps do not change within-dataset partitions but are reported when interpreting cross-dataset classification results.
4.6. Within- and Cross-Dataset Evaluation
The benchmark compares within-dataset performance with transfer across datasets that differ in visual, statistical, or annotation characteristics. It preserves native detection annotations to test both domain and annotation-protocol shifts.
- Within-dataset evaluation uses samples from the same processed dataset and acquisition and annotation protocol.
- Cross-dataset evaluation tests source-trained representations directly on target test data without target-domain fine-tuning or adaptation.
- Image-level transfer uses datasets with compatible canonical disease taxonomies, including healthy leaves, black rot, Esca or black measles, and leaf blight or Isariopsis leaf spot.
- Object-detection transfer between Mildew symptom and LDD uses shared downy-mildew and powdery-mildew categories in both directions.
- Native bounding-box protocols are retained, combining visual-domain shift with annotation-protocol shift rather than treating shared class names as sufficient equivalence.
- Region-level classification remains within-dataset because no second benchmark dataset has sufficiently compatible labels and annotation interpretation.
- Combining within-dataset testing, overlap analysis, and cross-dataset transfer separates single-source performance from robustness to independent images and annotation conventions.
5. Experimental Setup
The experimental setup applies shared protocols within each task while maintaining separate evaluation settings for image-level classification, region-level classification, and object detection. Models use common partitions, preprocessing, and task-specific training procedures.
- The benchmark evaluates image-level classification, region-level classification, and object detection under a common experimental framework.
- All models use the same dataset partitions and task-specific metrics, with ImageNet-pretrained classifiers and shared Ultralytics detector configuration.
- Datasets without predefined classification partitions receive stratified 70%, 15%, and 15% train, validation, and test splits using seed 42.
- HERMOS regions are split at the source-image level so related regions remain in one partition.
- Classification inputs use 224×224 pixels, with random crops and augmentation during training and deterministic center crops for validation and testing.
- Classification models use a two-stage transfer-learning protocol that first trains the head, then jointly fine-tunes the unfrozen network.
- Weighted cross-entropy with label smoothing of 0.1 uses training-partition class weights to reduce class-imbalance influence.
- YOLO detectors train for 100 epochs at 1,024×1,024 resolution, while RF-DETR-Nano uses an effective batch size of 16 with gradient accumulation.
6.1. Image-level Classification Results
Within-dataset classification is near-perfect on several controlled or derivative datasets but substantially harder on heterogeneous field-oriented datasets. Cross-dataset results show that high in-domain accuracy does not guarantee transfer, especially to GVLiD.
- Several controlled or derivative datasets achieve near-perfect or perfect accuracy across most evaluated architectures.
- 1.0000 accuracy is achieved by every evaluated model on the New Plant Diseases grape subset.
- GL-Portugal is an intermediate case: natural vineyard illumination is present, but leaves are fully visible in most images, reducing scene complexity.
- GVLiD reaches a maximum accuracy of 0.8966, while GLHCD reaches 0.8530 with EfficientNet-B0 and EfficientNet-B3.
- GVLiD difficulty is concentrated in Esca/black measles and leaf blight, whereas healthy leaves and black rot are classified correctly in all selected Swin-T test cases.
- GLHCD remains challenging because brown spot and mites are difficult classes amid broader visual heterogeneity and symptom similarity.
- Complete overlap among several controlled datasets explains part of their near-perfect transfer, limiting its interpretation as independent external validation.
- Transfer involving GVLiD is weak, with controlled-to-GVLiD accuracies of 0.1667–0.3218 and reverse accuracies of 0.2731–0.3250.
6.2. Region-Level Classification Results
Region-level classification evaluates disease recognition on manually localized HERMOS crops, isolating classification from region localization. Accuracies are tightly grouped, with transformer models slightly leading while convolutional models remain competitive.
- HERMOS region-level classification supplies manually localized crops and measures recognition among healthy regions, dead arm, downy mildew, and powdery mildew.
- 0.8837–0.9148 accuracy spans the evaluated models, a difference of 3.11 percentage points across the complete model set.
- 0.9148 accuracy is achieved by DeiT-S/16, followed closely by Swin-Tiny at 0.9138 and ViT-S/16 at 0.9054.
- EfficientNetB0 and MobileNetV3-Large remain competitive at 0.9015 and 0.9010 accuracy, while increased model scale does not consistently improve results.
- These results should not be directly compared with complete-image classification because the HERMOS models receive ground-truth localized regions.
6.3. Object Detection Benchmark Results
Object-detection performance varies more across datasets than across models evaluated on the same dataset, reflecting differences in task definition and annotation structure. The benchmark therefore requires dataset-specific interpretation of mAP@50 and mAP@50:95.
- mAP@50 evaluates detections at IoU 0.5, whereas mAP@50:95 averages thresholds from 0.5 to 0.95 and emphasizes precise localization.
- RF-DETR-Nano leads the Mildew symptom dataset with 0.7853 mAP@50 and 0.5209 mAP@50:95.
- RF-DETR-Nano achieves the best LDD results at 0.6114 mAP@50 and 0.4273 mAP@50:95, while YOLO26s reaches 0.5939 and 0.4160.
- FD-Confounders is hardest under mAP@50:95, with YOLO26s achieving 0.5662 mAP@50 and 0.3109 mAP@50:95.
- FD-Confounders combines overlapping leaves with fine-grained symptoms, making precise leaf boundaries harder than approximately correct detections.
- YOLO26s is the strongest YOLO-family model across all three datasets, but the transformer detector’s advantage is dataset-dependent rather than universal.
- 0.5209, 0.4273, and 0.3109 are the best mAP@50:95 values on the Mildew symptom, LDD, and FD-Confounders datasets, respectively.
6.4. Cross-Dataset Object Detection Generalization
Cross-dataset detection tests YOLO26s under zero adaptation between datasets sharing downy- and powdery-mildew labels but differing in annotation and visual properties. Transfer collapses in both directions, showing that shared labels do not ensure equivalent detection tasks.
- YOLO26s is used in both transfer directions with the same architecture and training protocol, isolating dataset and annotation shift.
- The transfer study compares the only two detection datasets with directly harmonizable categories: downy mildew and powdery mildew.
- The datasets use different annotation procedures: direct symptom boxes for Mildew symptom versus enclosing boxes derived from LDD polygon annotations.
- 0.0010 and 0.0003 mAP@50 and mAP@50:95 are obtained by the LDD-trained model on Mildew symptom, while the reverse transfer reaches 0.0016 and 0.0008.
- The near-zero target-domain results reflect severe transfer loss rather than failure to learn the source tasks, which retain measurable within-dataset performance.
- Class composition, framing, backgrounds, symptom severity, illumination, acquisition distance, variety, and contextual leaf tissue differ across the datasets.
- Strong within-dataset mAP alone is not evidence that a detector will generalize independently; compatibility also requires annotation semantics, target scale, box construction, acquisition protocol, and symptom appearance.
- Potential remedies include annotation harmonization, multi-dataset training, domain adaptation, or fine-tuning with labelled target-domain examples.
7. Conclusions
The benchmark shows that dataset characteristics and annotation definitions strongly shape apparent model performance, while cross-dataset transfer remains unreliable despite shared disease labels. Reliable deployment therefore requires realistic data, compatible annotations, provenance analysis, and external validation.
- Image-level classification: Controlled or derivative datasets produced near-perfect or perfect image-level accuracy for most evaluated architectures, leaving little separation among modern models.This near-saturation was observed on PlantVillage, PDR2018, the Grape Leaf Disease Dataset, its augmented version, and the New Plant Diseases grape subset.
- Image-level classification: GVLiD and GLHCD were more difficult because of greater field variability, scene complexity, broader taxonomy, and visual heterogeneity.Provenance analysis also identified exact image overlap among several controlled or derivative datasets, requiring cautious interpretation of apparent transfer.
- Cross-dataset evaluation: Cross-dataset classification transfer to GVLiD caused substantial accuracy degradation, showing that shared class names and strong within-dataset performance do not ensure field-domain generalization.The result highlights the gap between controlled or derivative datasets and an independently acquired field domain.
- Region-level classification: DeiT-S/16 achieved 0.9148 accuracy in HERMOS region-level classification, narrowly exceeding Swin-Tiny at 0.9138.Because annotated regions were extracted before classification, these results measure discrimination after localization and are not directly comparable with complete-image classification or detection.
- Object detection: Detection performance depended more strongly on dataset and annotation definition than detector generation alone, with RF-DETR-Nano and YOLO26s leading different datasets.RF-DETR-Nano reached mAP@50:95 values of 0.5209 on the Mildew symptom dataset and 0.4273 on the LDD leaf-only subset; YOLO26s reached 0.3109 on FD-Confounders.
- Cross-dataset evaluation: Despite shared downy- and powdery-mildew categories, cross-dataset detection nearly collapsed: mAP@50:95 was 0.0003 from LDD to Mildew symptom and 0.0008 in reverse.Source-domain values of 0.2079 and 0.4972 indicate the collapse reflected poor transferability rather than failed source-domain training.
- Limitations: The benchmark is limited by incompatible detection annotation units, restricted taxonomy compatibility, derivative or overlapping classification content, non-exhaustive HERMOS annotations, and incomplete coverage of real-world conditions.It evaluates available public datasets but does not establish clinical or agronomic validity across all cultivars, disease stages, regions, devices, or seasons.
- Future directions: Future benchmarks should use larger, diverse field datasets with expert-verified labels, exhaustive spatial annotations, documented acquisition protocols, standardized terminology, and harmonized annotation semantics.The paper also identifies multi-dataset training, domain adaptation, self-supervised pretraining, and target-domain fine-tuning as directions for improving comparison and transfer.