Source-linked AI summary
VID-AD: A Dataset for Image-Level Logical Anomaly Detection under Vision-Induced Distraction
Hiroto Nakata, Yawen Zou, Shunsuke Sakai, Shun Maeda, Chunzhi Gu, Yijin Wei, Shangce Gao, Chao Zhang
TL;DR
Logical anomaly detection is difficult to evaluate robustly because visual distractions can obscure rule-level violations while existing benchmarks rarely control for them. VID-AD introduces controlled capture variations and a text-based contrastive framework, which consistently outperforms vision-based baselines across capture conditions.
Problem
Existing benchmarks rarely test logical anomaly detection under controlled background, illumination, and blur variations while preserving the underlying logical state.
Method
VID-AD evaluates one-class logical anomaly detection across five unchanged-rule capture conditions, while its framework learns from normal-image text descriptions and constrained contradictory rewrites.
Results
The proposed method consistently outperforms existing vision-based baselines across all evaluated capture conditions.
Takeaways & Limitations
VID-AD provides a controlled benchmark for assessing whether logical anomaly detectors remain robust when visual appearance changes but scenario rules do not.
Takeaways & Limitations
The benchmark assumes models learn logical constraints using only normal samples, without real anomalous training images.
Abstract
from arXiv · showhide
Logical anomaly detection in industrial inspection remains challenging due to variations in visual appearance (e.g., background clutter, illumination shift, and blur), which often distract vision-centric detectors from identifying rule-level violations. However, existing benchmarks rarely provide controlled settings where logical states are fixed while such nuisance factors vary. To address this gap, we introduce VID-AD, a dataset for logical anomaly detection under vision-induced distraction. It comprises 10 manufacturing scenarios and five capture conditions, totaling 50 one-class tasks and 10,395 images. Each scenario is defined by two logical constraints selected from quantity, length, type, placement, and relation, with anomalies including both single-constraint and combined violations. We further propose a language-based anomaly detection framework that relies solely on text descriptions generated from normal images. Using contrastive learning with positive texts and contradiction-based negative texts synthesized from these descriptions, our method learns embeddings that capture logical attributes rather than low-level features. Extensive experiments demonstrate consistent improvements over baselines across the evaluated settings. The dataset is available at: https://github.com/nkthiroto/VID-AD.
1. Introduction
The introduction identifies a gap in logical anomaly detection: patch-centric visual methods are distracted by low-level environmental variation, while existing benchmarks rarely test controlled capture conditions. It presents VID-AD and a text-based contrastive framework designed to learn logical consistency from normal-image descriptions without anomalous training images.
- Motivation: Logical anomalies arise from global constraint violations such as changes in quantity, length, type, placement, or relation rather than obvious structural defects.
- Motivation: Patch-centric visual detectors struggle with global consistency and produce spurious responses when background, illumination, or blur distracts their local representations.EfficientAD is described as producing strong anomaly responses on normal samples while failing to clearly localize the actual logical anomaly.
- Dataset contribution: VID-AD introduces controlled vision-induced distractions for evaluating logical anomaly detection robustness across varying capture conditions.The benchmark addresses the limited coverage of logical violations under environmental variation in existing datasets.
- Method contribution: The proposed text-based framework learns logical consistency solely from language representations of normal images, bypassing real anomalous training images.It uses constrained, replacement-only rewrites that enforce attribute-level contradictions while preserving the original text structure.
- Method contribution: Contrastive learning with positive and semantically perturbed negative descriptions suppresses irrelevant visual cues while preserving global structural semantics.
- Results: The method consistently outperforms existing vision-based baselines across all capture conditions and achieves state-of-the-art performance on VID-AD.
2. Related Work
Existing industrial anomaly-detection benchmarks and methods primarily target structural defects through local visual evidence. VID-AD addresses the scarcity of evaluations combining logical constraint violations with controlled visual distraction using a text-based, logic-focused framework.
- Datasets: MVTec AD, BTAD, and VisA cover diverse structural defects across industrial categories, while COCO-AD adapts general-purpose data for visual anomaly detection.These datasets have advanced visual anomaly detection but remain centered on structural defect types such as scratches, dents, contamination, and surface damage.
- Datasets: Recent benchmarks broaden evaluation through multi-view capture, diverse viewpoints, wider industrial scenarios, and more challenging imaging conditions.Examples include PAD, Real-IAD, MANTA, CableInspect-AD, MVTec AD 2, and AutVI.
- Datasets: MVTec LOCO AD introduced systematic logical anomalies, but lacks controlled low-level visual distractors while logical states remain unchanged.This leaves joint evaluation of logical constraint violations and vision-induced distraction scarce.
- Methods: Existing methods use patch-level deviation, reconstruction discrepancy, distillation, component consistency, or few-shot matching, excelling mainly on structurally explicit local irregularities.Their reliance on local visual patterns can miss logical violations requiring global consistency reasoning across objects or regions.
- Methods: The proposed framework converts images into logic-focused descriptions and models semantic consistency in language embedding space to target rule violations beyond visual cues.By avoiding visual feature embeddings, the approach is designed to remain robust to appearance-induced distractions.
3. VID-AD Dataset
VID-AD is a controlled one-class benchmark for logical anomaly detection that fixes scenario rules while varying capture conditions, enabling evaluation under vision-induced distraction. It covers ten manufacturing scenarios, five conditions, and 50 tasks with structured constraint violations and normal-only training.
- Dataset design: VID-AD contains 10 manufacturing scenarios and five capture conditions, yielding 50 independent one-class tasks trained exclusively on normal samples.The benchmark varies environmental conditions while preserving each scenario’s underlying logical rules.
- Benchmark scope: VID-AD focuses exclusively on logical anomalies and evaluates consistency across five capture conditions rather than conflating logical violations with structural defects.This design distinguishes it from benchmarks that primarily target structural defects or mix structural and logical anomalies.
- Logical anomalies: Each scenario pairs two constraints from Quantity, Length, Type, Placement, and Relation, while anomalies violate the first, second, or both constraints.Normal samples satisfy both paired constraints simultaneously, enabling separate analysis of single-aspect and combined violations.
- Capture conditions: The five capture conditions are White BG, Cable BG, Mesh BG, Low-light CD, and Blurry CD, representing background, illumination, and lens-related variations.These variations are designed to distract vision-centric detectors without changing the logical state.
- Dataset protocol: 10,395 images comprise 2,500 training samples and 7,895 testing samples, with each task using normal-only training and mixed normal–anomalous testing.A typical task includes 50 training normal images, 50 testing normal images, and approximately 110 testing anomalous images.
4. Proposed Method
The proposed method addresses vision-induced distraction in one-class logical anomaly detection by converting images into logic-focused text descriptions. A frozen Vision-Language Model guided by scenario-specific prompts enables detection based on semantic descriptions rather than raw pixels.
- Motivation: The method targets decoupling logical states from significant appearance variations in a one-class setting trained only on normal samples.VID-AD requires models to learn logical constraints using normal samples while handling vision-induced distractions.
- Framework: The vision-to-text framework performs detection using semantic descriptions instead of raw images to prioritize logical consistency over irrelevant pixel-level fluctuations.This design is intended to reduce the influence of appearance variation during detection.
- Framework: A frozen Vision-Language Model converts each image into a logic-focused text description guided by scenario-specific prompts.The conversion process is illustrated in Fig. 3.
S S S
The proposed unsupervised detector converts images into logic-focused textual descriptions, learns logical consistency from positive and synthesized contradictory texts, and scores test samples against normal training embeddings.
- Text-based detection framework: A frozen VLM generates one scenario-specific, logic-focused text per image, while BERT is fine-tuned using dropout-based positives and text-only contradictory negatives.The contrastive objective pulls consistent descriptions together and pushes contradiction-based descriptions apart.
- Logic-focused description generation: Prompts target object type, color, count, region, relative length, and spatial relations to suppress irrelevant background and lighting variation.The same prompt and structured output format are used during training and testing, supporting stable embeddings in tasks with roughly 50 images.
- Negative text synthesis: For each normal description, text-only attribute substitution produces a fluent negative description that introduces logical inconsistencies without anomalous images.The synthesized negatives preserve formatting constraints while representing contradictory logical states.
- Contrastive representation learning: BERT mean-pools final-layer token states and independently learns a compact embedding region for logically consistent descriptions in each one-class task.Dropout creates two stochastic views of the same positive text, making the encoder less sensitive to minor feature variation while retaining contradiction sensitivity.
- Similarity-based anomaly scoring: During inference, the detector compares each test-text embedding with stored normal embeddings using a k-nearest-neighbor distance score, with k = 5 in all experiments.Higher normality scores indicate proximity to the training distribution, whereas lower scores indicate logical deviation.
5. Experiments
Experiments on VID-AD compare the proposed language-based method with established one-class visual anomaly detectors and evaluate robustness across capture conditions, scenarios, and VLM choices. The method achieves the strongest and most stable condition-wise performance, while robustness depends on scenario rule expressibility and description quality.
- The evaluation compares the method with PaDiM, PatchCore, AnoGAN, VAE, EfficientAD, and CSAD across feature-space, reconstruction, distillation, and component-consistency paradigms.
- Condition-wise robustness: 0.132 to 0.207 AUROC: the proposed method outperforms CSAD across all five capture conditions and achieves the best AUROC under each condition.Its cross-condition standard deviation is 0.013, competitive with UniVAD while retaining higher absolute performance.
- Condition-wise robustness: 0.129 AUROC: EfficientAD shows a large best-to-worst condition gap, while other vision-centric models fluctuate by 0.070 to 0.101 and UniVAD has a 0.033 gap at lower accuracy.These gaps indicate sensitivity to low-level visual variations.
- Scenario-wise sensitivity: 0.006 and 0.009: the method has low scenario-wise sensitivity for Fruits and Ropes, whereas Sticks reaches 0.096 because descriptions can omit rule-critical relations such as relative length.Fruits descriptions remain consistent across conditions, while Sticks descriptions sometimes fail to preserve rule-relevant relations.
- VLM comparison: 0.831: Qwen2 7B achieves the best overall mean AUROC and smallest cross-condition variance, exceeding Llama3.2 11B at 0.822±0.036 and LLaVA 13B at 0.701±0.034.Qwen2 7B remains strong under both background variations and degraded capture conditions, whereas Llama3.2 11B degrades under Low-light CD and Blurry CD.
6. Discussion
The discussion links robustness to representation invariance: language embeddings emphasizing logical attributes reduce sensitivity to background changes and capture degradations. It also frames linguistic expressibility and VLM choice as factors in language-based detection, evaluated across scenarios and capture conditions.
- Robustness Under Vision-Induced Distraction: Language-embedding detection emphasizes logical attributes instead of raw visual features, reducing sensitivity to background changes and capture degradations.The discussion explicitly connects robustness under vision-induced distraction with representation invariance.
- Qualitative Linguistic Expressibility: Qualitative VLM-generated positive texts compare a stable Fruits scenario with a challenging Sticks scenario under vision-induced distraction.The examples are selected using scenario-wise mean AUROC and condition-wise variability.
- VLM Choice: The VLM ablation generates one positive description per image and one corresponding negative per positive example, reporting mean AUROC across 10 scenarios and capture conditions.The table compares Qwen2 7B, Llama3.2 11B, and LLaVA 13B, with mean and standard deviation across capture conditions.
- Sensitivity Across Capture Conditions: Scenario-wise sensitivity is visualized across five capture conditions using AUROC dots, min–max vertical bars, and condition-mean diamond markers.The figure summarizes per-scenario variation in the proposed method on VID-AD.
7. Conclusion
The paper addresses logical anomaly detection under fixed logical states and varying low-level visual conditions by introducing VID-AD and a language-based detection framework. VID-AD provides a one-class benchmark, while the framework learns from text descriptions rather than visual features.
- Contributions: VID-AD is a one-class benchmark with 50 tasks spanning 10 manufacturing scenarios and five capture conditions.The benchmark targets logical anomaly detection when low-level visual variations occur while the underlying logical state remains fixed.
- Contributions: The proposed detection framework learns from text descriptions rather than visual features.Its training uses contrastive learning between normal descriptions and semantically related text, as stated in the supplied passage.
- Motivation: The work targets the challenge of detecting logical anomalies under low-level visual variations with fixed underlying logical states.This motivates both the benchmark and the language-based framework.
CRediT authorship contribution statement
The authors’ contributions span conceptualization, methodology, validation, investigation, writing, resources, project administration, funding acquisition, and supervision.
- Hiroto Nakata led conceptualization, methodology, validation, investigation, and original-draft writing.
- Chao Zhang contributed methodology, investigation, resources, review and editing, project administration, funding acquisition, and supervision.
- Yawen Zou, Shunsuke Sakai, Shun Maeda, Chunzhi Gu, Yijin Wei, and Shangce Gao contributed writing review and editing.
Declaration of competing interest
The authors declare no competing interests relevant to the article and no financial or nonfinancial affiliations or involvement with organizations connected to its subject matter or materials.
- Declaration of competing interest: The authors declare no relevant competing interests.They certify having no financial or nonfinancial affiliations or involvement with organizations or entities connected to the manuscript’s subject matter or materials.