Source-linked AI summary

LogiScope-VQA: Benchmarking Vision-Language Models for Logistics Hazard Identification in Industrial Scenarios

Hanjing Zhou, Mingze Yin, Ying Lian, Jun Ma, Chang-Yu Hsieh, Yanbing Zhou

arXiv:2609.09790v1cs.CVcs.AIcs.CL

TL;DR

Existing LMM benchmarks provide limited evidence about safety-oriented performance in real-world industrial logistics, where data scarcity and diverse operational requirements remain important gaps. The paper introduces LogiScope-VQA, a real-world benchmark combining multimodal logistics data, structured VQA tasks, expert annotation, and risk-bias analysis. Results show that current models remain substantially below logistics experts, with fine-grained perception and conservative risk bias emerging as major challenges.

  • Problem

    Limited real-world industrial data and incomplete multi-task evaluation leave LMM safety-oriented perception and reasoning in logistics scenarios insufficiently assessed.

  • Method

    The paper curates LogiScope-VQA from industrial surveillance data and expert annotations, covering perception, warehouse knowledge, risk reasoning, and risk-bias evaluation.

  • Results

    Current LMMs show a pronounced gap relative to human experts, while fine-grained industrial perception is a shared bottleneck and conservative risk bias is prevalent.

  • Takeaways & Limitations

    LogiScope-VQA provides a comprehensive testbed for evaluating real-world logistics deployment readiness and identifying visual, reasoning, and bias-related weaknesses.

  • Takeaways & Limitations

    Current models particularly struggle with low-resolution, wide-angle, heavily occluded, and densely cluttered logistics surveillance footage, while most over-report risks.

Abstract

from arXiv · show

Large Multimodal Models (LMMs) large-scale deployment in industrial warehouse settings specifically necessitates that models exhibit human-expert-level hazard-oriented perception, understanding, and reasoning capabilities. However, the scarcity of real industrial data, tightly coupled to commercial terms, significantly hampers further advancement. To bridge this gap, we curate LogiScope-VQA to investigate the practical applicability of mainstream LMMs in real-world logistics operations. LogiScope-VQA comprises 2,476 images and 2,918 videos primarily sourced from real-world logistics parks, along with 10,274 VQAs meticulously curated and validated by human annotators. Grounded in 18 core objects and 20 risk types, we devise 39 subtasks aligned with three principal themes: industrial element perception, warehouse knowledge understanding, and potential risk reasoning. Furthermore, we incorporate dynamic thinking-budget configurations and dual-dimensional risk bias analyses to elucidate the properties of LMMs. Extensive experiments unveil that even powerful proprietary models, including GPT-5.5, Gemini-3.1-Pro, and Claude-Opus-4.7, exhibit a significant gap relative to human performance. The unique challenge of jointly integrating perception, understanding, and reasoning for hazard identification poses substantial headroom for further improvement on LogiScope-VQA. We additionally reveal the pervasive security bias issue that impedes LLMs' practical deployment in real-world settings. The industrial dataset is publicly available under the CC BY-NC-SA 4.0 license.

1. Introduction

LogiScope-VQA addresses the limited evaluation of LMMs in real-world industrial logistics, where safety-oriented requirements and dense spatial targets create underexplored challenges. It introduces a benchmark spanning perception, warehouse knowledge, and risk reasoning to assess practical deployment value.

  • Industrial logistics evaluation remains limited because real-world data collection is constrained by commercial cost, privacy, safety, and liability concerns.
  • The benchmark evaluates industrial element perception, warehouse knowledge understanding, and potential risk reasoning through a progressive curriculum.
  • Real-world logistics safety assessment must cover diverse dimensions, including object perception, spatial reasoning, path planning, and action recognition.
  • LogiScope-VQA contains 2,476 images, 2,918 videos, 1,0274 VQAs, 20 risk factors, and 18 logistics objects.
  • A holistic evaluation of 20 mainstream LMMs finds substantial performance disparity relative to human experts despite parity between leading open-source and proprietary models.
  • The benchmark is intended to bridge simulation-validated models and practical industrial application while exposing visual-perception challenges, risk-reasoning characteristics, and response biases.

2. LogiScope-VQA

LogiScope-VQA is a real-world industrial benchmark designed to test whether foundation LMMs can perceive logistics objects, apply domain knowledge, integrate multimodal information, and identify safety risks. Its curation combines large-scale surveillance data, structured question generation, and multi-round expert validation.

  • 2.1. Overview of Logistics Benchmark: The benchmark evaluates foundation LMM perception and reasoning for logistics safety in real-world industrial scenarios.
  • 2.1. Overview of Logistics Benchmark: Its visual corpus begins with 3.5 million Cainiao surveillance clips and is distilled into 5,394 high-quality visual samples after cleaning, filtering, and augmentation.
  • 2.1. Overview of Logistics Benchmark: LogiScope-VQA covers single-choice, multiple-choice, and open-ended VQA problems grounded in 18 core objects and 20 risk types.
  • 2.1. Overview of Logistics Benchmark: The benchmark challenges models to perceive dense objects, operationalize logistics knowledge, integrate multimodal information, and anticipate safety risks through sophisticated thinking.
  • 2.1. Overview of Logistics Benchmark: Unlike simulated or narrowly focused prior benchmarks, LogiScope-VQA primarily uses real-world warehouse data and provides comprehensive multidimensional task coverage.
  • 2.3. Dataset Curation Pipeline: Its hierarchical curation pipeline aggregates, cleans, edits, annotates, generates, refines, and cross-validates VQA pairs for coverage and annotation quality.

3. Experiment

Experiments evaluate diverse proprietary and open-source LMMs on LogiScope-VQA, including human reference participants, thinking-mode configurations, time budgets, and safety-risk bias analyses. Results show persistent fine-grained perception gaps, task-dependent thinking benefits, saturation with longer budgets, and systematic safety-risk biases.

  • Evaluation Setups: The benchmark evaluates six proprietary models, multiple open-source model families, chance-level baselines, and novice and expert human references.The evaluation uses a unified prompting and scoring setup, with micro-averaged accuracy and LLM-as-a-Judge processing for open-ended answers.
  • Main Performance Analysis: Top open-source performance essentially ties the strongest proprietary systems, while leading models remain far below the logistics expert.Top models surpass the untrained novice human but remain substantially behind the experienced logistics specialist.
  • Main Performance Analysis: Fine-grained industrial perception is a shared bottleneck, with top models clustering in a narrow low-performing band on single- and cross-instance perception tasks.Wide-angle, low-resolution, cluttered surveillance imagery is identified as a major challenge, contrasting with stronger performance on knowledge-intensive tasks.
  • In-depth Thinking Analysis: Thinking mode has mixed, task-dependent effects: industrial element perception responds least, warehouse knowledge gains modestly, and potential risk reasoning benefits most.The reasoning-heavy risk task aligns particularly well with explicit chain-of-thought, especially for stronger models such as Qwen3.5-Plus.
  • In-depth Thinking Analysis: Performance on potential risk reasoning saturates across models as thinking time increases, with severe degradation at low budgets and marginal gains beyond the standard budget.The analysis recommends selectively enabling thinking mode and setting the time budget near the saturation point to balance performance and token consumption.
  • Safety Risk Bias Analysis: Safety-risk judgments show systematic conservative bias: most models have negative Risk Bias Scores, while proprietary and open-source systems diverge in bias patterns.Paired-image analysis finds most models exceed 60% risk recall but vary substantially in false-alarm rates; even models with the best trade-off remain below human experts.

4. Challenges and Future Directions

LogiScope-VQA identifies fine-grained perception and safety-risk bias as two major limitations of current models, defining priorities for future research.

  • Current limitations: Fine-grained perception remains a bottleneck, especially in low-resolution, wide-angle, occluded, and densely cluttered surveillance footage.Suggested directions include pre-training for dense, low-resolution targets and exploiting temporal information across video frames.
  • Current limitations: Most models heavily over-report risks, revealing a prevalent safety-risk bias that industrial applications must address alongside accuracy.The paper suggests balanced data curation or bias-aware training objectives to prevent this bias at its source.
  • Future directions: Future work should improve intrinsic visual capacity and reduce risk bias to support more reliable logistics safety assessment.These directions follow the paper’s identified perception bottleneck and fairness concern.

5. Related Work

Industrial logistics safety is an underexplored evaluation setting because real-world data are constrained, while existing benchmarks provide limited realism or task breadth. LogiScope-VQA addresses this gap with expert-annotated real-world data and comprehensive evaluation.

  • Research context: Industrial logistics safety combines object recognition, spatial reasoning, and safety-critical hazard identification in a practical LMM testbed.These requirements distinguish industrial environments from natural-scene evaluation.
  • Research gap: Commercial confidentiality constrains access to industrial data, limiting systematic evaluation in realistic logistics environments.The paper identifies commercial constraints as a persistent obstacle to industrial benchmark development.
  • Prior benchmarks: IndustryEQA contains approximately 1.3 thousand physics-simulated samples, whereas LogiScope-VQA provides 10,274 expert-annotated VQA instances.The comparison emphasizes the difference between simulated prior data and the released benchmark’s expert-annotated scale.
  • Benchmark scope: Existing logistics benchmarks evaluate narrower task slices, while LogiScope-VQA offers a dedicated suite spanning workflows, equipment states, human behaviors, and safety regulations.The paper characterizes LogiScope-VQA as the first dedicated evaluation suite for logistics scenarios.

6. Conclusion

The paper concludes that LogiScope-VQA comprehensively evaluates LMM capabilities for logistics hazard identification and exposes substantial weaknesses in industrial perception and integrated reasoning. It also describes data-protection measures and a non-commercial licensing boundary.

  • Conclusion: LogiScope-VQA evaluates industrial element perception, warehouse knowledge understanding, and potential risk reasoning.These three capabilities form the benchmark’s comprehensive assessment of logistics hazard identification.
  • Conclusion: LMMs fall short of expectations in industrial visual perception, while effective hazard identification depends on coherent sophisticated reasoning.The conclusion presents perception weakness and reasoning requirements as central findings.
  • Authorship and data construction: Approximately 20.5% of the dataset is generated by a large language model, while the conceptual innovation and manuscript were produced independently by human authors.The paper states that LMMs were used only at the final stage for terminology polishing and partly during dataset construction.
  • Responsible AI: Privacy-sensitive content, including faces, warehouse names, and commercial identifiers, was blurred or masked through automated detection and manual verification.The stated process aims to prevent identification of individuals and proprietary entities.

C. Data Usage and Licensing

The supplied material describes LogiScope-VQA’s non-commercial licensing and its hierarchical curation pipeline, which converts real logistics data and structured safety knowledge into audited VQAs. The pipeline emphasizes predefined schemas, balanced coverage, and quality control.

  • Data Usage and Licensing: LogiScope-VQA is released under CC BY-NC-SA 4.0 for academic and non-commercial research only.Commercial use is prohibited, and derivative works must retain the same license.
  • Visual Data Acquisition: The raw corpus contains 3.5 million surveillance clips collected from Cainiao warehouse parks across diverse operational scenarios.The collection spans high racks, loading docks, storage zones, picking shelves, and office areas.
  • Knowledge representation: Safety knowledge is encoded as hierarchical JSON tuples of Object, Attribute, and Value across 18 core warehouse-safety entities.Values are predefined enumerations, such as a person climbing a rack, to constrain the attribute space.
  • Object Attribute Annotation: Annotators localize objects before assigning predefined attributes, with video labeling restricted to the keyframe showing the most salient risk characteristics.Third-party audits and iterative correction cycles support annotation precision.
  • Question Generation: Fifteen question templates transform object attributes into natural-language VQA pairs, followed by rule-driven refinement for disambiguation and fluency.The templates include twelve multiple-choice and three open-ended formats.
  • Dataset Assembly: The dataset is rebalanced across hallucination proportions, core object types, and hazard categories to address long-tail distributions.Ground-truth answers are assigned and low-quality pairs filtered through specialized annotation and auditing.

D.3. Full Task Taxonomy

The section presents the task taxonomy and the prompt-based procedure for generating targeted visual data and evaluating logistics-safety VQA responses. The assistant is instructed to reason from observable evidence without fabricating unsupported details.

  • Task taxonomy: Fig. 9 presents the benchmark’s full task taxonomy.
  • Data generation: Targeted image-editing prompts use multiple placeholders whose allowable values create diverse resulting prompts.The values are comprehensively enumerated in Fig. 11.
  • Evaluation prompting: The logistics safety VQA assistant must answer from observable visual evidence, allowing logical reasoning but forbidding unsupported fabrication.
  • Evaluation prompting: Responses must contain two ordered parts, beginning with a thorough reasoning analysis.

E. Evaluation Details

Evaluation uses prompt templates tailored to question solving and LLM-as-a-Judge assessment. Short-answer correctness depends on containing all essential ground-truth key points.

  • Prompt templates: The left evaluation template handles multiple-choice and open-ended question solving.
  • Prompt templates: The right template supports an LLM-as-a-Judge workflow in which a third-party LMM labels responses “correct” or “incorrect”.
  • Correctness criterion: For short-answer questions, a response is judged correct when it contains all essential ground-truth key points.

E.2. Human Evaluation Protocol

Human performance is measured with two participant groups differing in expertise. Undergraduate volunteers provide generalist evaluation while lacking specialized logistics and safety knowledge.

  • Participant groups: Human evaluation recruits two participant groups with contrasting expertise levels.
  • Generalist evaluators: Ten undergraduate volunteers serve as generalist evaluators with common sense and basic visual-recognition abilities.
  • Generalist evaluators: The students lack specialized logistics-operation and safety-protocol knowledge and answer approximately 1,000 questions each.

E.3. Reliability of LLM-as-Judge Evaluation

The study evaluates LLM-as-a-Judge reliability through repeated judgments and comparison with human experts. On 100 VQA pairs, repeatability exceeds 93% and agreement with experts exceeds 87%.

  • Reliability protocol: Reliability is assessed through intra-judge repeatability and inter-judge agreement with independent human judgments.The assessment uses a randomly sampled subset of 100 open-ended VQA questions.
  • Reliability protocol: Intra-judge agreement compares three evaluations of each question at temperature = 0.7.
  • Reliability protocol: Inter-judge agreement compares LLM-judge decisions with independent human-expert judgments.
  • Results: Intra-judge repeatability exceeds 93%, while agreement with human experts exceeds 87%.The reported results support the judge’s stability and alignment with human judgment.
Loading 2609.09790v1…