Source-linked AI summary

AT-ViT: Area-Targeted Multi-View Vision Transformer with Cross-Attention and Multi-Scale Patching for Plant Trait Recognition in Herbarium Images

Amani Sedrat, Takieddine Chehhat, Youcef Sklab, Hanane Ariouat, Abderrazak Sebaa, Eric Chenin, Jean-Daniel Zucker, Edi Profiti

arXiv:2608.21067v1cs.CVcs.AI

TL;DR

Herbarium trait classifiers can rely on background artifacts instead of plant morphology, limiting reliable recognition. AT-ViT combines raw and segmented views with cross-attention and segmentation-guided patch weighting, improving accuracy, plant-region attention alignment, and robustness to background noise.

  • Problem

    Herbarium classifiers may attend to background regions rather than biologically relevant plant structures, motivating plant-focused recognition that preserves raw morphological information.

  • Method

    AT-ViT jointly processes raw and segmented herbarium views with multi-scale cross-attention fusion and segmentation-guided patch weighting.

  • Results

    AT-ViT consistently improves fine-grained trait accuracy, plant-region attention alignment, and robustness to synthetic background noise.

  • Takeaways & Limitations

    AT-ViT provides a more plant-focused and interpretable framework for reliable image-based plant phenotyping.

  • Takeaways & Limitations

    Segmentation masks may contain residual background pixels or preprocessing artifacts that models can exploit, limiting segmented-input reliability.

Abstract

from arXiv · show

Automated plant traits recognition from herbarium images is essential for plant sciences, yet remains challenging because background elements (e.g., textual labels, mounting artifacts, and color charts) can introduce shortcut learning, leading models to rely on spurious non-plant cues rather than plant morphology. This bias degrades both generalization and interpretability. In this paper, we introduce AT-ViT, a dual-branch Vision Transformer that jointly encodes raw herbarium scans and their segmented-derived counterparts via a multi-scale, multi-view cross-attention fusion scheme. AT-ViT further incorporates a mask-guided patch weighting mechanism that amplifies plant-relevant regions and attenuates background-driven features. By learning from the original scans while being guided by segmentation masks through the mask-guided patch reweighting mechanism, the model is encouraged to focus on plant organs and learn plant-centric representations more effectively. Across multiple trait classification tasks (e.g., leaf base shape, thorns), AT-ViT delivers consistent accuracy gains, improves attention localization on plant regions, and exhibits increased robustness under synthetic background perturbations. Specifically, AT-ViT substantially improves spatial attention grounding, boosting plant-region alignment (Avg IoU_p: +15.66 to +18.03 pp) while reducing background overlap (Avg IoU_b: -27.92 to -31.02 pp) relative to CrossViT, and remains markedly more robust to background perturbations, outperforming ResNet101 by up to +32.32 accuracy points and CrossViT by up to +5.07 points under background-noise conditions.

1. Introduction

Herbarium images enable large-scale plant science but challenge automated trait recognition because models may rely on background artifacts instead of localized plant morphology. AT-ViT addresses this by combining raw and segmented views with segmentation-guided patch weighting and cross-attention to improve trait classification, attention alignment, interpretability, and robustness.

  • Motivation: Digitized herbarium specimens support large-scale research in taxonomy, trait evolution, biodiversity monitoring, and phenological analysis.Platforms such as GBIF and ReColNat provide millions of images spanning centuries, continents, and taxa.
  • Motivation: Fine-grained traits such as leaf base shape and small thorns require precise localization of plant organs, yet residual background biases can persist after segmentation.Plant material typically occupies only 20% of the image area, and models may encode residual cues, dataset priors, or preprocessing artifacts.
  • Motivation: Background artifacts can induce shortcut learning, causing models to attend to non-plant cues rather than biologically relevant structures.Labels, mounting materials, scale bars, and color charts are identified as sources of attention bias.
  • Proposed approach: AT-ViT uses a multi-view, multi-scale dual-branch Vision Transformer to preserve raw-scan morphology while incorporating cleaner structural guidance from segmentation.A cross-attention fusion mechanism combines fine-grained texture and appearance information from raw scans with segmentation-derived guidance.
  • Contributions: AT-ViT improves fine-grained trait accuracy, plant-region attention alignment, and robustness to synthetic background noise.The framework is presented as a scalable and interpretable approach that encourages plant-centric representations through segmentation-informed patch weighting.

2. Related Work

Related work progresses from CNN-based herbarium classification and organ-level analysis to transformers, segmentation, and multi-branch models. These studies motivate segmentation-guided transformers that retain raw-scan detail while emphasizing plant regions for fine-grained trait recognition.

  • CNN-based methods: CNNs established herbarium-image classification and later supported organ-level localization, detection, and segmentation across varied sheet layouts and mounting conditions.Foundational studies classified plant taxa and leaf traits, while subsequent pipelines addressed leaves and stems.
  • Vision Transformers: ViTs use self-attention to capture long-range interactions, improving generalization and class balance relative to CNN baselines and supporting fine-grained leaf classification with limited training data.Their performance suggests capacity for capturing subtle morphological cues.
  • Segmentation: Segmentation pipelines evolved from coarse thresholding and color-space filtering to deep architectures such as U-Net and DeepLabv3+ for isolating plant regions.U-Net++ was used for leaf segmentation and counting.
  • Multi-branch architectures: Multi-modal and multi-branch architectures combine complementary representations, including CNN features with segmented-region point clouds and dual-branch transformer image streams.SIM-Net used 2D CNN features and 3D point clouds, while ROI-ViT integrated an image stream.
  • Research motivation: Collectively, prior studies motivate segmentation-guided transformers that preserve raw-scan detail while explicitly biasing models toward plant regions for fine-grained trait recognition.The present work integrates segmentation-informed spatial priors into a multi-view transformer framework.

3. Approach

AT-ViT is a dual-branch, multi-scale transformer that jointly processes raw and segmented herbarium images, using segmentation-derived patch weighting to emphasize plant-relevant regions. Bidirectional cross-attention then aligns fine-grained raw-image textures with segmented morphological structures for binary trait classification.

  • Dual-branch multi-view architecture: AT-ViT jointly processes raw herbarium images and segmented counterparts through synchronized small and large transformer branches.The architecture is redesigned from single-input CrossViT to support multi-view processing.
  • Dual-branch multi-view architecture: The branches use multi-scale patching, with 400 384-dimensional tokens for raw 240 × 240 inputs and 196 768-dimensional tokens for segmented 224 × 224 inputs.Different spatial resolutions and embedding dimensions target localized detail and global structure.
  • Mask-guided patch weighting: Patch saliency is computed from each patch’s non-black plant-pixel ratio and transformed with a sigmoid to distinguish densely and sparsely covered regions.The ratio is defined over [0,1], and the scaling (6 · ratio − 3) recenters the sigmoid for mid-density discrimination.
  • Mask-guided patch weighting: Weights are applied once after linear patch embedding and before the first encoder block, amplifying morphology-rich patches while suppressing background and irrelevant regions.The [CLS] token remains unweighted, while the modulation propagates through subsequent encoder blocks.
  • Cross-attention fusion: Bidirectional cross-attention uses one branch’s [CLS] token to query the other branch’s patch tokens, and concatenates the resulting [CLS] tokens into a 1152-dimensional binary-classification representation.This aligns raw-image textures with segmented morphological structures.

4. Experiments

Experiments use ReColNat herbarium scans to evaluate three binary morphological-trait classification tasks spanning distinctive and fine-grained visual features. Performance is measured by classification accuracy and plant/background attention IoU, with preprocessing and training settings tailored to AT-ViT’s two branches.

  • Dataset and Tasks: The ReColNat dataset contains high-resolution MNHN herbarium scans annotated for three binary traits: thorns, acute leaf base, and acuminate leaf tip.The data are designed to isolate plant material from labels, mounting paper, and color scales.
  • Dataset and Tasks: The selected traits span visual complexity from distinctive thorns to fine-grained discrimination of acute leaf bases and acuminate leaf tips.Trait selection reflects data availability rather than a specific ecological or taxonomic motivation.
  • Implementation: Training used AdamW with weight decay 0.1, learning rate 0.00005, 70 epochs, batch size 16, standard augmentation, and a scheduler halving the rate every 5 epochs.Augmentations included random horizontal flips, affine transformations, and color jittering.
  • Implementation: Raw inputs were resized to 240×240 and segmented masks to 224×224, matching the resolution requirements of AT-ViT’s respective branches.The experiments were performed on a high-performance node with 2×NVIDIA H100 GPUs, 192 CPU cores, and 768 GB of RAM.
  • Evaluation: Evaluation reports classification accuracy plus IoUp for plant-mask overlap and IoUb for background-mask overlap.Binary attention masks were formed by thresholding attention maps at 0.5, with pixels ≥ 0.5 assigned to 1.
  • Evaluation: The mask-guided weighting mechanism is expected to increase IoUp and decrease IoUb by amplifying plant-dense patches and down-weighting background patches before attention redistribution.The IoU metrics range from 0 to 1 and quantify spatial attention alignment.

5. Results and Analysis

AT-ViT consistently improves plant-focused spatial alignment over ResNet101 and CrossViT, while classification gains remain modest. Visualizations, embedding structure, and synthetic-noise tests further indicate more localized, discriminative, and background-robust representations.

  • Quantitative comparison: Over 30 percentage points: AT-ViT increases IoUp versus ResNet101 across all three morphological traits.AT-ViT achieves markedly superior spatial alignment across the evaluated traits while maintaining moderate background attention.
  • Attention localization: AT-ViT concentrates activations on plant structures, unlike baseline attention that is diffuse or extends into labels, scale bars, and other background areas.The patch-weighting strategy suppresses background artifacts and focuses feature activations on plant-relevant regions.
  • Representation and robustness: AT-ViT produces enhanced clustering and separation of thorns-trait embeddings and consistently outperforms both baselines when noise predominantly affects background regions.The robustness result indicates that patch weighting suppresses irrelevant background information and improves resilience to visual clutter.
  • Attention localization: AT-ViT’s attention concentrates near the image center, matching the predominantly central distribution of plant pixels shown by aggregated segmentation masks.ResNet101 emphasizes a small central area, whereas CrossViT tends toward peripheral and background regions.

6. Conclusion

AT-ViT integrates raw and segmented herbarium images through multi-scale, multi-view encoding, cross-attention fusion, and segmentation-guided patch weighting. The study reports improved classification, biologically meaningful attention alignment, and robustness to background noise, while future work targets spatial alignment constraints and broader evaluation.

  • Conclusion: AT-ViT integrates raw and segmented herbarium images through multi-scale, multi-view encoding, cross-attention fusion, and segmentation-guided patch weighting.The architecture is described as a dual-branch Vision Transformer.
  • Conclusion: AT-ViT consistently improves classification accuracy, aligns attention more closely with biologically meaningful regions, and increases robustness to background noise and visual clutter.These outcomes are reported as experimental results in the conclusion.
  • Future Work: Future work will incorporate spatial alignment constraints such as Intersection-over-Union (IoU) into training and evaluate AT-ViT on larger, more taxonomically diverse datasets spanning more morphological traits.The planned objective is to promote closer correspondence between attention maps and true plant regions.

Funding information

The research was funded by the Agence Nationale de la Recherche through the e-Col+ Project.

  • The Agence Nationale de la Recherche funded the e-Col+ Project under Grant/Award Number ANR-21-ESRE-0053.
Loading 2608.21067v1…