Source-linked AI summary
PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
Zixin Zhang, Kanghao Chen, Xingwang Lin, Lutao Jiang, Xu Zheng, Yuanhuiyi Lyu, Litao Guo, Yinchuan Li, Ying-Cong Chen
TL;DR
Although MLLMs are used in embodied AI, the depth of their physical-tool understanding remains largely unquantified. PhysToolBench evaluates that understanding across tool recognition, operation, and creation, and its evaluation finds a substantial gap between MLLMs and human performance.
Problem
The depth of MLLMs’ physical-tool comprehension remains largely unexplored despite their use in embodied planning and VLA models.
Method
PhysToolBench is a VQA benchmark of over 1,000 image-text pairs that assesses physical-tool understanding across recognition, understanding, and creation.
Results
32 MLLMs showed a clear performance ceiling: even the most advanced proprietary models scored no higher than 63%, versus human proficiency over 90%.
Takeaways & Limitations
The evaluation identifies a substantial gap in current MLLMs’ ability to reason about physical tools.
Abstract
from arXiv · showhide
The ability to use, understand, and create tools is a hallmark of human intelligence, enabling sophisticated interaction with the physical world. For any general-purpose intelligent agent to achieve true versatility, it must also master these fundamental skills. While modern Multimodal Large Language Models (MLLMs) leverage their extensive common knowledge for high-level planning in embodied AI and in downstream Vision-Language-Action (VLA) models, the extent of their true understanding of physical tools remains unquantified. To bridge this gap, we present PhysToolBench, the first benchmark dedicated to evaluating the comprehension of physical tools by MLLMs. Our benchmark is structured as a Visual Question Answering (VQA) dataset comprising over 1,000 image-text pairs. It assesses capabilities across three distinct difficulty levels: (1) Tool Recognition: Requiring the recognition of a tool's primary function. (2) Tool Understanding: Testing the ability to grasp the underlying principles of a tool's operation. (3) Tool Creation: Challenging the model to fashion a new tool from surrounding objects when conventional options are unavailable. Our comprehensive evaluation of 32 MLLMs-spanning proprietary, open-source, specialized embodied, and backbones in VLAs-reveals a significant deficiency in tool understanding. Furthermore, we provide an in-depth analysis and propose preliminary solutions. Code and dataset are publicly available.
1 INTRODUCTION
Physical-tool comprehension is important for embodied interaction, yet its depth in MLLMs remains largely unexplored. PhysToolBench evaluates that understanding across recognition, reasoning about tools, and creating solutions from available objects; evaluation reveals a substantial performance gap.
- 1 INTRODUCTION: Physical-tool use is a prerequisite for embodied agents to complete physical tasks successfully and efficiently.The introduction illustrates this with a robot needing a hammer to drive a nail.
- 1 INTRODUCTION: Although MLLMs serve as embodied planners and VLA backbones, the depth of their physical-tool comprehension remains largely unexplored.The paper notes that prior work has demonstrated only preliminary tool understanding.
- 1 INTRODUCTION: PhysToolBench evaluates tool recognition, operational understanding, and the ability to fashion tools from surrounding objects when standard options are unavailable.Its medium tier includes tool selection under constraints, multi-tool tasks, and assessment of physical viability.
- 1 INTRODUCTION: Across 32 MLLMs evaluated on PhysToolBench, even the most advanced proprietary models scored no higher than 63%, compared with human proficiency over 90%, revealing a substantial tool-understanding disparity.The analysis also found small-model failures, long-tail tool-recognition issues, affordance hallucinations, and inadequate visual reasoning.
2 RELATED WORKS
MLLMs are used for embodied planning and action, where tool use matters for complex tasks, but their physical-tool understanding remains underexplored. Existing digital-tool benchmarks and A4Bench leave room for a more task-oriented physical-tool evaluation.
- 2.1 MLLM AND ITS APPLICATION IN EMBODIED AI: MLLMs process visual inputs and are used both as high-level planners and as backbones for VLA models that output robot actions.The paper notes that embodied models can perform basic household tasks, while more complex objectives may exceed manipulators alone.
- 2.2 PHYSICAL TOOL USE IN EMBODIED AI: Early robot tool-use approaches show rudimentary capability, but the depth of physical-tool understanding in their MLLM ‘brains’ remains largely unexplored.Examples include tool crafting, tool use during task planning, and learning tool manipulation from demonstrations.
- 2.3 RELATED BENCHMARKS: Digital-tool benchmarks do not cover physical tools, while A4Bench asks models to identify an imaged tool’s function rather than choose tools for a task.PhysToolBench instead supplies a task requirement and an image of several tools, testing selection based on observation and reasoning.
3 THE PHYSTOOLBENCH
PhysToolBench is a VQA benchmark that asks models to select tools from an image for a described task, with only depicted objects available. Its tiered design probes progressively deeper understanding, and human-reviewed collection combines generated and photographed scenes.
- 3.1 OVERVIEW: PhysToolBench contains over 1,000 task-and-image pairs, with labeled objects and an instruction to choose only among depicted items or answer “None.”The benchmark spans daily life, industrial, outdoor activities, and professional settings.
- 3.2 DESIGN PRINCIPLES: Easy tests primary-function recognition, Medium tests attributes, tool combinations, and availability, and Hard tests repurposing objects to meet task requirements.The tiers are presented as a progression from basic tool-use planning to complex scenarios and a forward-looking AGI challenge.
- 3.3 DATASET COLLECTION PROCESS: Experts designed task-scene pairs, primarily generated images with GPT-4o-image, staged and photographed some complex objects, then labeled and independently reviewed the dataset.Approximately 90% of images were generated and 10% were physically staged and photographed.
4 EXPERIMENTS ON PHYSTOOLBENCH
The experiments evaluate 32 MLLMs across four model classes using a consistent prompt, with chain-of-thought encouraged except for models with built-in thinking modes. Most models trail human accuracy, and the supplied figures compare open-source model size and embodied models with their base models.
- 4.1 BENCHMARK CANDIDATES: 32 MLLMs across four classes were evaluated, with proprietary models tested through APIs and the others deployed locally.A consistent prompt encouraged chain-of-thought reasoning, except for models with built-in thinking modes; five human participants served as a reference.
- 4.2 OVERALL RESULTS: Below 60%: most MLLMs scored under this mark, compared with humans’ at-least-87.85% overall accuracy.The paper reports that proprietary general-purpose models performed best; GLM-4.5V scored 55.14% and outperformed some proprietary models.
- 4.2 OVERALL RESULTS: GLM-4.5V achieved 55.14%, outperforming its open-source peers and some proprietary models, while proprietary general-purpose MLLMs performed best among evaluated models; VLA-framework MLLM backbones exhibited the weakest performance.
- 4.3 FINDINGS ON PHYSTOOLBENCH: Figure 4 reports a significant correlation between open-source model size and performance; Figure 5 compares embodied models with their base models.The supplied figure descriptions identify the comparisons but do not give additional values.
F.1. A foundational ability to understand tools emerges in large models with sufficient scale.
Tool comprehension correlates with model scale, but even top-tier MLLMs retain long-tail recognition and functional-understanding failures. The reported scale threshold is associated with stronger easy-task performance, not immunity to these errors.
- F.1. A foundational ability to understand tools emerges in large models with sufficient scale.: Models above 10B parameters generally achieve 60–70% accuracy on easy-level tasks, while models below 5B generally score below 50% on easy tasks and below 25% overall.The authors preliminarily identify approximately 10 billion parameters as the scale at which foundational tool-use understanding emerges on easy tasks.
- F.2. A long-tail problem persists in tool recognition and understanding, even for the most ad-: Top-tier MLLMs struggle with long-tail items, especially digital products, and confuse visually similar HDMI/DP cables and Type-C/Lightning ports.This weakness appears in open-source models and remains, though marginally improved, in closed-source models.
- F.2. A long-tail problem persists in tool recognition and understanding, even for the most ad-: Most top-tier proprietary models miss the functional requirement to connect a monitor to a Type-C-only laptop using an HDMI cable and adapter.
F.3. Embodied-specific MLLMs show no significant advantage on PhysToolBench.
VLA MLLM backbones scored below 15% overall, and embodied-task fine-tuning did not improve tool understanding over comparable base models. Models also struggled to recognize when presented tools were nonfunctional, a weakness with stated safety implications.
- F.3. Embodied-specific MLLMs show no significant advantage on PhysToolBench.: RoboBrain2’s 32B and 7B variants and Embodied-R1-3B scored marginally below their equivalent Qwen2.5VL backbones.The authors suggest that current robotic datasets may need more high-quality tool-comprehension data.
- F.3. Embodied-specific MLLMs show no significant advantage on PhysToolBench.: In four illustrated cases, none of the MLLMs identified that a damaged or nonfunctional tool was unavailable.The benchmark’s M3 tier tests this with cases where the correct tool is present but unusable; the authors link the failures to shallow, surface-level associations.
- F.3. Embodied-specific MLLMs show no significant advantage on PhysToolBench.: The authors warn that attempting to use a nonfunctional tool can cause mission failure and safety hazards, including fueling a tractor with gasoline or using a damaged syringe.
- F.5. The MLLM backbones in current VLAs are extremely weak.: Contemporary VLA MLLM backbones all scored below 15% overall on PhysToolBench.The authors argue that modest-scale robotic datasets cannot rectify this limitation and call for more capable backbones and larger, more diverse action datasets.
F.6. Reasoning ability is important and useful, but still insufficient.
Reasoning-focused prompting and training improved reported benchmark performance, but models still showed hallucinations and spatial-reasoning errors. A vision-centric method that analyzes detected object crops improved M3 performance, while the conclusion presents PhysToolBench as a tiered measure of tool understanding.
- F.6. Reasoning ability is important and useful, but still insufficient.: Chain-of-Thought-prompted models had significantly higher accuracy, and reasoning-optimized Ovis-2.5-9B scored 48.52%, comparable to a 72B model at 49.51%.GLM-4.5V’s built-in thinking mode also outperformed other open-source models and some proprietary models, according to the authors.
- 4.4 A PRELIMINARY SOLUTION: Despite these gains, models still hallucinate, and none recognized that a suitably sized flathead screwdriver could unscrew the illustrated Phillips screw.The authors call for greater focus on visual-centric reasoning for high-level planning tasks.
- 4.4 A PRELIMINARY SOLUTION: The proposed agent combines global image-and-query analysis, object detection and cropping, detailed crop analysis, and multi-level evidence integration.It uses DINOX object detection to identify objects and crop them by their bounding boxes.
- 4.4 A PRELIMINARY SOLUTION: GPT-4o and GPT-5 gained 10.24% and 18.06% performance, respectively, on M3 with the same backbone MLLM using vision-centric reasoning.M3 is the difficulty level where existing models perform worst.
- 5 CONCLUSION: PhysToolBench comprises 1,000 image-text pairs across three difficulty tiers, and the authors propose it as a standard for measuring embodied agents’ tool capabilities.Their evaluation covered 32 MLLMs and found all tested models significantly below human performance.
A.1 DATASET CONSTUCTION
The authors constructed PhysToolBench through expert brainstorming, supervised image generation and review, and human-in-the-loop annotation. They filtered the process’s initial cases to produce 1,000 high-quality examples.
- Phase 1: Conceptualization.: Five coauthor experts brainstormed task-scene pairs for three weeks, producing an initial collection of 1,500 cases.
- Phase 2: Image Generation.: GPT-4o generated images from scene descriptions, with human experts reviewing quality, realism, and accuracy and refining many images through one to three iterations.For cases where generation repeatedly failed, particularly complex digital products, the team staged and photographed the scenes physically.
- A.1 DATASET CONSTUCTION: Annotators could flag problematic images or tasks, which a separate reviewer group re-examined and, when needed, regenerated.
- A.1 DATASET CONSTUCTION: After three rounds of construction and review, the team retained 1,000 high-quality cases.
A.2 REALISM EVALUATION OF GENERATED IMAGES
Generated-image realism was assessed by GPT-4o and a 10-participant user study, with both rating the images highly. The appendix also provides full image-question-answer materials and GPT-4o predictions for examples.
- A.2 REALISM EVALUATION OF GENERATED IMAGES: GPT-4o rated 100 images 1.92/2 on average, while 10 participants rated the same images 1.78/2, indicating that the images were generally realistic.Both evaluations used a scale from 0 (unrealistic) to 2 (highly realistic).
- B COMPLETE DEMONSTRATION OF IMAGE–QUESTION–ANSWER TRIPLETS: Figures 12–20 provide full examples, including original images and prompts, ground-truth answers, and GPT-4o outputs.