Source-linked AI summary

Precision in Rice Variety Classification using Stacking-Based Ensemble Learning

Md. Masudul Islam, Galib Muhammad Shahriar Himel, Md. Golam Moazzam, Mohammad Shorif Uddin

arXiv:2609.10524v1cs.CV

TL;DR

Rice variety diversity creates a need for accurate visual classification, while existing research provides limited robust coverage. This study combines a 20-variety dataset with a stacking-based ensemble and smartphone deployment, achieving 100% accuracy while noting image-quality, overfitting, and uncontrolled-environment limitations.

  • Problem

    Existing rice-variety research falls short of robust and efficient classification based on external characteristics such as color, size, and texture.

  • Method

    The study curates 20 rice varieties, extracts deep features with tuned transfer-learning models, combines classifiers using stacking, and deploys the model in an Android application.

  • Results

    100% accuracy was achieved on the proposed rice-variety classification task.

  • Takeaways & Limitations

    The integrated model and smartphone application support automated rice-variety identification for users capturing grain images.

  • Takeaways & Limitations

    The approach remains vulnerable to overfitting, dependence on high-quality images, similar varieties, and uncontrolled environments.

Abstract

from arXiv · show

Rice, a staple food for a significant portion of the global population, exhibits remarkable diversity in its varieties, presenting substantial challenges for accurate identification by consumers, traders, and farmers. This complexity often facilitates fraudulent practices, such as the unauthorized mixing of rice types, which undermines quality and trust in the supply chain. Despite its critical importance, existing research falls short of providing robust and efficient methods for precise rice variety classification based on external characteristics like color, size, and texture. To address this gap, our study introduces a comprehensive rice variety identification framework designed to enhance transparency and quality assurance. We developed a stacked ensemble model tailored for rice variety classification and curated a comprehensive dataset comprising 20 rice varieties, each distinguished by unique visual attributes. The proposed approach achieved an unprecedented classification accuracy of 100%. Furthermore, we integrated our model into a mobile application, enabling even novice users to effortlessly identify rice varieties using grain images from a smartphone camera. These findings underscore the transformative potential of advanced machine learning techniques in mitigating fraudulent practices and ensuring stringent rice quality control. Our work holds significant implications for agricultural stakeholders, paving the way for automated crop identification systems and advancing precision agriculture practices.

1. Introduction

Rice variety diversity makes accurate visual identification difficult, while existing methods leave gaps in robust classification. The study addresses these gaps with a 20-variety dataset, deep-feature extraction, stacking, and smartphone deployment.

  • The study targets automated rice variety identification amid extensive variety diversity and the importance of rice production and consumption.Rice includes about 40,000 varieties, while Asia accounts for 90% of global cultivation.
  • The authors developed a structured dataset covering 20 distinct rice varieties and extracted deep features using hyperparameter-optimized models.
  • A stacking-based ensemble combines the best-performing baseline models to improve rice classification accuracy.
  • The model was integrated into a smartphone application that identifies rice varieties from grain images.

2. Literature Review

Prior rice-classification studies use varied datasets, features, and learning methods but often lack balanced coverage of sample size, variety diversity, and grain differences. This study responds with a larger, more diverse dataset and advanced deep-learning methods.

  • Prior studies applied neural networks, image processing, hyperspectral imaging, and conventional classifiers across rice datasets with varied sizes and variety counts.Reported examples include nearly 90% precision, 93.34% accuracy, 98.57% accuracy, and 99.85% accuracy in different studies.
  • Many reviewed studies had small sample sizes, limited variety coverage, or insufficient variation in morphology, color, and texture.
  • The study contributes a comprehensive dataset spanning 20 rice varieties with diverse sizes, shapes, and color-texture variations.
  • The proposed work addresses underused deep-learning methodologies through an extensive dataset and advanced deep-learning techniques for rice variety classification.

3. System Architecture

The proposed system lets users capture a rice-grain image with a smartphone and receive a predicted variety through an Android application. The back-end model processes the image and displays the result.

  • Users capture a rice-grain image with at least 5× smartphone-camera zoom, after which the Android app processes it automatically.
  • The installed model predicts the rice variety name and displays the result on the application interface.
  • Figure 1 presents the architecture of the proposed smartphone-based rice variety identification expert system.

4. Methodology

The methodology combines a curated 20-variety rice dataset, tuned transfer-learning models, and stacking-based ensemble classification. It also incorporates regularization, cross-validation, multiclass evaluation metrics, and external validation to assess model reliability.

  • Datasets: Data augmentation expanded the dataset from approximately 4,500 images to 27,000 images while preserving key grain characteristics through flips and rotations.The dataset was split into training and testing sets at an 80:20 ratio.
  • Datasets: The methodology acknowledges that controlled images do not fully represent mixed-grain samples or environmental variability, motivating future data expansion.The stated future direction includes non-controlled environments and additional globally sourced varieties.
  • Proposed System Architecture: The framework combines deep-feature extraction from tuned transfer-learning models with stacking-based ensemble learning using selected high-performing classifiers.Fourteen meta-learners integrate the extracted feature sets, with XGBoost selected as the final classifier.
  • Evaluation Metrics: The evaluation uses precision, recall, F1-score, accuracy, multiclass AUC-ROC, and cross-validation to assess classification performance and overfitting.The cross-validation procedure repeatedly trains on m−1 subsets and validates on the remaining subset.

5. Result and Discussion

The experiments evaluate baseline models, stacked ensembles, cross-validation robustness, external validation, and mobile deployment. The stacking approach achieved 100% accuracy across reported top-model combinations, while external validation remained nearly perfect for five rice varieties.

  • Baseline Evaluation: More than 95% accuracy was used to select baseline models, with the first three exceeding 99% and DenseNet201 showing the lowest validation loss.The selected baselines were evaluated using accuracy curves, ROC curves, confusion matrices, and classification reports.
  • Ensemble Results: 100% accuracy was achieved by the XGBoost meta-classifier across all top-model combinations.The top baseline models were EfficientNetV2L, VGG16, MobileNetV2, and DenseNet201.
  • Robustness Evaluation: Five-fold cross-validation averaged accuracy, precision, recall, and F1-score across five validation iterations to assess robustness.Each fold served once as validation data while the other four folds were used for training.
  • External Validation: External validation produced 100% accuracy for Arborio, Basmati, Ipsala, and Karacadag, and 99.71% for Jasmine.The external dataset contained 700 samples per class across five rice varieties; Jasmine had two misclassifications.
  • Model Deployment: The model was optimized for deployment through TensorFlow Lite conversion, quantization, and cloud-based inference access.The smartphone application introduces challenges from variable lighting, camera quality, and user image-capture errors.
  • Limitations: The study reports 100% accuracy on its self-curated dataset but identifies overfitting, image-quality dependence, and visually similar varieties as limitations.These limitations are especially relevant in uncontrolled environments.

6. Comparative Result Analysis

The comparative analysis situates the proposed ensemble model within prior rice-variety classification research. It emphasizes differences in dataset scale, variety coverage, and methodological performance across studies.

  • Comparative Analysis: The study positions its dataset as addressing imbalances between the number of varieties, sample size, and morphological and color-textural diversity.The comparison notes that some prior studies used many varieties with small datasets, while others used large datasets with limited variety diversity.
  • Comparative Analysis: The comparison table reports the proposed method alongside results from other rice-variety classification studies.The supplied table passage identifies the comparison scope but does not provide individual cell values.

7. Conclusions and Future Directions

The study combines tuned transfer learning, deep-feature extraction, and stacking ensembles with an Android expert system for rice variety classification, achieving 100% accuracy. Future work targets broader real-world robustness through more diverse data, augmentation, and lightweight models.

  • 100% accuracy was achieved by combining tuned transfer learning, deep-feature extraction, and stacking-based ensemble classification.The approach uses rice grain images and integrates the resulting model into an Android-based expert system.
  • Larger, geographically diverse datasets with uncontrolled-environment images and additional varieties are needed to better reflect real-world conditions.The stated target conditions include mixed-grain samples, harvesting variation, and environmental factors.
  • Synthetic augmentation that simulates noise and occlusions is proposed to improve model adaptability.

Funding

The authors report no organizational support for the submitted work.

  • No organizational funding or support was received for the submitted work.
Loading 2609.10524v1…