Source-linked AI summary

CreativityBench: Evaluating Agent Creative Reasoning via Affordance-Based Tool Repurposing

Cheng Qian, Hyeonjeong Ha, Jiayu Liu, Jeonghwan Kim, Jiateng Liu, Bingxuan Li, Aditi Tiwari, Dwip Dalal, Zhenhailong Wang, Xiusi Chen, Mahdi Namazifar, Yunzhu Li, Heng Ji

arXiv:2605.02910v2cs.AIcs.CLcs.LG

TL;DR

Creative tool use in LLMs remains underexplored, especially whether models can ground non-obvious solutions in fine-grained physical affordances. CreativityBench evaluates this ability using a structured affordance knowledge base and 14K grounded tasks, finding that models often choose plausible tools but struggle with parts, affordances, and mechanisms.

  • Problem

    Creative intelligence in LLMs remains poorly defined and insufficiently evaluated, particularly for repurposing objects through fine-grained physical affordances.

  • Method

    The paper builds an affordance knowledge base linking entities, parts, attributes, and uses, then evaluates models on grounded creative tool-repurposing tasks.

  • Results

    Over 60% performance drops when models must ground a plausible tool at the correct part or attribute level, while scaling and standard inference strategies yield limited gains.

  • Takeaways & Limitations

    Creative intelligence is not simply an extension of analytical reasoning or action planning but appears to require grounded comparison and flexible affordance recombination.

  • Takeaways & Limitations

    The paper notes that its metrics are not directly comparable, especially when models select different tools and follow different reasoning paths.

Abstract

from arXiv · show

Recent advances in large language models have led to strong performance on reasoning and environment-interaction tasks, yet their ability for creative problem-solving remains underexplored. We study this capability through the lens of creative tool use, where a model repurposes available objects by reasoning about their affordances and attributes rather than relying on canonical usage. As a first step, we introduce CreativityBench, a benchmark for evaluating affordance-based creativity in LLMs. To this end, we build a large-scale affordance knowledge base (KB) with 4K entities and 150K+ affordance annotations, explicitly linking objects, parts, attributes, and actionable uses. Building on this KB, we generate 14K grounded tasks that require identifying non-obvious yet physically plausible solutions under constraints. Evaluations across 10 state-of-the-art LLMs, including closed and open-source models, show that models can often select a plausible object, but fail to identify the correct parts, their affordances, and the underlying physical mechanism needed to solve the task, leading to a significant drop in performance. Furthermore, improvements from model scaling quickly saturate, strong general reasoning does not reliably translate to creative affordance discovery, and common inference-time strategies such as Chain-of-Thought yield limited gains. These results suggest that creative tool use remains a major challenge for current models, and that CreativityBench provides a useful testbed for studying this missing dimension of intelligence, with potential implications for planning and reasoning modules in future agents.

1 Introduction

The paper frames creative tool use as a distinct form of intelligence: repurposing objects through affordance-based, physically grounded reasoning to devise novel solutions under constraints. CreativityBench operationalizes this capability with a large affordance knowledge base and grounded tasks, revealing severe limits in current models’ exact physical grounding and creative affordance discovery.

  • Creative intelligence requires generating novel yet useful solutions under constraints, especially when success depends on inventing paths by repurposing available resources.
  • Creative tool use infers an object’s affordances from physical attributes and repurposes its parts for unconventional actions beyond intended uses.
  • Existing evaluations underexplore creative tool use because they primarily assess plausible actions, embodied planning, multimodal understanding, or interactive exploration rather than grounded affordance reasoning.
  • Across 14K tasks, models often identify plausible tools but fail to ground their use at specific parts or attributes, causing a performance drop of over 60%.The introduction identifies exact physical grounding as a severe bottleneck and distinguishes it from analytical reasoning and creative affordance discovery.
  • The affordance knowledge base contains 4K entities and 150K+ affordance annotations for grounded task sampling, trajectory construction, training, and evaluation.

2 Related Work

Prior work evaluates LLM creativity through psychological tests and studies physical reasoning, interactive decision-making, and structured affordance knowledge. However, existing approaches often lack scalable, physically grounded modeling of fine-grained object parts, attributes, and mechanisms.

  • Creativity Evaluation: LLMs have demonstrated creative capabilities in narrative and poetry generation, tool and system design, real-world problem modeling, and human brainstorming and ideation.Creativity is framed as enabling robust action in novel and unfamiliar environments.
  • Creativity Evaluation: Psychological creativity assessments measure fluency, originality, and flexibility, but are prompt-sensitive, costly, noisy, and imperfect indicators of model creativity.These evaluations nevertheless suggest that modern models can achieve strong creativity scores.
  • Interactive and Physical Reasoning: Interactive benchmarks such as lagerBench emphasize perception, planning, and coordination, but typically lack physically grounded, fine-grained affordances of objects and components.Their scenario-driven or prompt-generated tasks therefore emphasize planning, reasoning, or multimodal understanding rather than underlying mechanisms.
  • Interactive and Physical Reasoning: PIQA, PROST, and NEWTON evaluate physical commonsense, object attributes, simple affordances, and broader physical reasoning in everyday tasks.These benchmarks represent a line of work on whether AI systems can reason about physical attributes and affordances of everyday objects.
  • Structured Affordance Knowledge: SYNTHIA decomposes objects into parts and associated affordances, but primarily models conceptual part-affordance associations without explicitly representing physical attributes determining whether a part can be used.This work highlights the importance of part-level functional decomposition while leaving physical attribute modeling incomplete.

3 Preliminaries: Structured Reasoning Is Not Enough for Creativity

A controlled comparison on 100 MacGyver creative tool-use tasks tests whether structured affordance-level reasoning improves creativity. Structured CoT modestly improves procedural grounding but not creative affordance reinterpretation, pointing to grounded affordance knowledge as the key limitation.

  • Experimental setup: 100 creative tool-use tasks from MacGyver compare direct prompting with structured affordance-level CoT before adding new knowledge or benchmark resources.The structured guideline includes tool inventory listing and part decomposition.
  • Evaluation: Solutions are evaluated on Correctness, Feasibility, Physical Grounding, Constraint Coverage, Tool Usage, and Creativity.The criteria distinguish task achievement, physical executability, mechanics, constraints, tool selection, and non-obvious affordance reinterpretation.
  • Results: 3.44 to 3.52 Feasibility, 3.44 to 3.54 Physical Grounding, and 4.22 to 4.31 Tool Usage are the absolute gains from structured CoT.These are modest improvements on procedural dimensions.
  • Results: 47% vs. 41% Physical Grounding and 38% vs. 28% Tool Usage are CoT’s relative-evaluation win rates, while Creative Reasoning declines.CoT improves grounding and tool-use outcomes but performs worse on creative reasoning.
  • Interpretation: Structured CoT improves procedural grounding without stronger creative affordance reinterpretation, suggesting that models lack grounded affordance knowledge that can be flexibly recombined.MacGyver is not explicitly structured around affordances, and its evaluation often relies on LLM-as…

4 CreativityBench

CreativityBench evaluates creative tool use as affordance-based reasoning: models must infer how object attributes and parts enable non-obvious actions rather than rely on semantic tool labels. It builds a hierarchical affordance knowledge base and reverse-engineers grounded benchmark tasks from known affordances with verifiable reasoning trajectories.

  • Affordance-based creative tool use: Tools are defined by affordances—action possibilities arising from structure, materials, interfaces, constraints, and accessible resources—rather than names or intended categories.Creative tool use requires identifying useful attributes and combining their affordances to achieve a goal.
  • Affordance knowledge base: The knowledge base hierarchically links entities, non-overlapping parts, physical and state attributes, and part-level affordances.Part-level modeling captures local structural features that may support uses distinct from an object’s canonical function.
  • Affordance knowledge base: Attribute combinations create distinct entity instances because differences in physical or state properties can substantially change their affordances.The knowledge base uses predefined attribute fields plus a flexible field for distinctive traits.
  • Affordance knowledge base: 157K annotated affordances span approximately 4K entities and 26K parts, with 288K physical attributes and 125K state attributes.The annotations are automatically scaled using GPT-5.2 and include affordances with varying typicality levels.
  • Benchmark task construction: Benchmark tasks are reverse-engineered from sampled affordances, requiring models to infer the correct entity, part, and affordance from scenario entities and attributes.The four-stage construction process comprises gold affordance sampling, task synthesis, gold verification, and distractor sampling.
  • Benchmark task construction: Gold affordance sampling clusters semantically similar affordances and varies cluster size and typicality to balance common and long-tail cases while controlling difficulty.Complete-linkage clustering over Text-Embedding-3-Large* embeddings yields approximately 3.5K clusters per scenario.

5 Experiments

Experiments evaluate models on 14K grounded tool-use tasks, requiring entity- and part-level selection plus usage specification. Results show a large entity-to-part grounding gap, plausible actions without sufficient physical grounding, and diminishing gains from scaling.

  • Evaluation Setup: All models are evaluated on 14K tasks requiring selection of the relevant entity and part, distractor filtering, and specification of how to use them.The main setting provides each task description, all scene entities, and every part’s attributes.
  • Metrics: Gold Correct Rate requires the correct entity and part, whereas Entity Correct Rate requires only the correct entity.Gold Correct Rate is defined to be no greater than Entity Correct Rate.
  • Exact Grounded Tool Use: 0.5149 entity correctness falls to 0.1910 gold correctness, a relative decrease of over 60%, showing that part-level grounding remains difficult.For some models, the entity-to-gold correctness gap reaches 0.35.
  • Grounded Reasoning Quality: 3.5860 action feasibility exceeds 3.2003 physical grounding, indicating that plausible actions often lack accurate physical-attribute justification.Models frequently rely on semantic alignment between the tool and task while failing to account for all execution conditions.
  • Scaling Effects: Qwen gold correctness rises from 0.1882 at 4B to 0.2483 at 14B, but reaches only 0.2588 at 32B, demonstrating diminishing returns from scaling.Comparable saturation appears for GPT and Gemini models, while scaling alone does not resolve grounded affordance reasoning.

6 Analysis

Model performance is shaped by affordance commonality, distractor structure, and inference mode. Rare affordances, larger candidate sets, interactive evaluation, and higher-temperature sampling often expose failures in physical grounding, exploration, and candidate comparison.

  • Distractors: Performance consistently decreases as distractor entities increase, with the largest average drop occurring from 3 to 9 distractors.The decline from 9 to 12 distractors becomes less steep, suggesting a diminishing effect at larger candidate-set sizes.
  • Distractors: Similar-affordance distractors achieve higher scores than dissimilar distractors and can partially offset the difficulty of larger candidate sets.Fine-grained analysis shows steadily decreasing performance with dissimilar distractors, whereas similar-distractor performance initially falls and then partially recovers.
  • Inference modes: Higher temperature generally degrades smaller-model performance while mildly helping larger models, primarily increasing hallucinated entities and parts rather than useful exploration.Naive repeated sampling therefore does not consistently improve affordance discovery and may produce ungrounded answers.
  • Inference modes: Interactive evaluation substantially reduces performance because models inspect fewer than three entities, commit early, and often fail to inspect the gold entity before answering.Correct entity-and-part predictions usually involve direct target examination, whereas completely incorrect predictions have gold inspection rates below 20% for most models.
  • Inference modes: Structured CoT produces only modest changes, indicating that the main bottleneck is candidate comparison and affordance-level physical grounding rather than reasoning format.Qwen models show gains typically within around 5%, while GPT-family models show slight declines.

7 Discussion

The paper frames creativity as structured, affordance-grounded reasoning rather than unconstrained imagination or hallucination. Its findings motivate physical-textual dual reasoning with foresight, while suggesting that distribution-sharpening training objectives may suppress the diversity needed for creative problem solving.

  • Difference of Creativity and Hallucination: Creativity is defined as grounded in object attributes and induced affordances, distinguishing it from creative writing, open-ended design, and innovative research ideation.The paper characterizes this as structured creativity rather than freer imagination or productive hallucination.
  • Toward Physical-Textual Dual Reasoning and Foresight Governance: Models often produce plausible-sounding but mechanically incorrect solutions even when given explicit parts, attributes, and affordances.This indicates a missing internal process for anticipating physical consequences before acting.
  • Toward Physical-Textual Dual Reasoning and Foresight Governance: Physical-textual dual reasoning would combine textual affordance recombination with physical imagination that predicts how parts, materials, and states evolve under candidate actions.The proposed loop is intended to improve creative discovery and provide foresight before action.
  • Enhancement of Model Creativity: Existing reinforcement-learning methods often emphasize sampling efficiency and produce distribution sharpening, which can improve reliability but suppress structured diversity for creative problem solving.The discussion therefore suggests that improving model creativity may require a training objective different from those optimized by much current reinforcement learning.

8 Conclusion and Future Work · A Significance, Scope, and Clarifications · A.1 Why CreativityBench Matters

CreativityBench evaluates creative tool repurposing through physical affordance-based reasoning, requiring models to connect objects with relevant parts, attributes, and mechanisms. Its findings expose substantial grounding failures despite advances in general reasoning, model scaling, and inference-time strategies, while providing a structured resource for analyzing grounded creativity and guiding future agents.

  • 8 Conclusion and Future Work: CreativityBench evaluates creative tool repurposing through affordance-based reasoning with a structured knowledge base linking entities, parts, attributes, and uses.The benchmark is designed to assess creative solutions beyond canonical object functions.
  • 8 Conclusion and Future Work: Models often fail to ground creative solutions in the correct part, attribute, and physical mechanism, even when they identify a plausible object.The largest performance drop occurs when evaluation moves from object-level plausibility to finer-grained grounding.
  • A.1 Why CreativityBench Matters: The benchmark targets a capability under-measured by evaluations centered on analytical reasoning, tool execution, and long-horizon planning: creative tool use grounded in physical affordances.Its central question is whether a model can identify which object and physical property make a non-obvious solution usable.
  • A.1 Why CreativityBench Matters: CreativityBench requires models to localize the relevant part, connect it to attributes, and recover the affordance mechanism enabling success.This moves evaluation beyond settings that reward selecting a generally relevant object.
  • A.1 Why CreativityBench Matters: Its structured affordance knowledge base enables scalable task construction while preserving an interpretable latent solution path.The same organization supports more informative failure analysis by distinguishing errors involving entities, parts, attributes, and affordances.
  • A.1 Why CreativityBench Matters: The benchmark highlights limits of strong general reasoning when physical grounding and unconventional repurposing are required.For embodied systems, it frames robust real-world problem solving as recognizing latent affordances under constraints rather than merely following canonical object functions.

A.2 Clarifications of Concerns

The clarifications frame CreativityBench as a structured, household-focused evaluation of mechanism-sensitive tool repurposing rather than unconstrained generation or simple retrieval. They also emphasize methodological choices and limitations that shape how single-gold scoring, human results, and cross-model comparisons should be interpreted.

  • Scope and Methodological Role: LLM assistance scales structured data construction, while benchmark tasks remain anchored in an ontology of entities, parts, attributes, affordances, and conditions.The structured representation is intended to reduce arbitrariness in task generation.
  • Why the Benchmark Measures More Than Retrieval: Each task requires mechanism-sensitive inference linking a goal, candidate part, relevant attributes, and use conditions, beyond retrieving familiar object-use associations.The benchmark targets non-obvious but physically grounded repurposing.
  • Single-Gold Structure: Single-gold evaluation is a measurement choice supported by gold verification and filtering of tasks when an alternative solution is judged preferable.This supports a controlled evaluation set without claiming that real-world creativity has only one valid solution.
  • Domain Coverage: The focused household domain enables tighter control over entity distributions, affordance granularity, and distractor sampling than a heterogeneous open-domain benchmark.The authors present this restricted scope as a current strength rather than a complete solution to creativity evaluation.

B Preliminary Experiments Details … C.3 Core Prompt Details

The paper specifies structured prompts for creative tool-use reasoning and evaluation, compares direct and affordance-CoT prompting across text-only and vision-plus-text settings, and constructs its affordance knowledge base through staged, schema-constrained annotation.

  • B.1 Core Prompt Details: The direct creative tool-use prompt requires feasible solutions using only listed tools, respecting physical constraints, enabling needed tool invention, and providing executable, minimal steps.If complete solution is impossible, the model must return the best partial plan and explain the limitation.
  • B.1 Core Prompt Details: The affordance-CoT prompt requires stating the goal, inventorying tools, identifying relevant parts and physical properties, deriving affordances, planning with part references, and validating constraints.Its output schema includes task goals, constraints, tool inventories, reasoning plans, solution steps, used tools, and constraint handling.
  • B.1 Core Prompt Details: The evaluation prompts score candidate solutions on correctness, feasibility, physical grounding, constraint coverage, tool usage, and creative reasoning using task-specific 1–5 rubrics.Correctness concerns achieving the goal, feasibility concerns physical execution under constraints, and physical grounding concerns object properties, geometry, and mechanics.
  • B.2 Textual v.s. Visual Grounding in Creative Tool Use: The study uses a 2×2 ablation with GPT-4.1-mini crossing Text-only versus Vision+Text inputs with Direct versus Affordance-CoT prompting.Text-only descriptions list goals, tools, and constraints, whereas vision+text inputs provide objects and visible constraints through images while text states goals and non-visible constraints.
  • C.1 Overview of Annotation Stages: The annotation pipeline transforms environment scenarios into object-part datasets through entity grounding, structural decomposition, physical and state characterization, affordance annotation, and entity assembly.Each stage consumes the preceding stage’s output so downstream annotations remain grounded in earlier structure and attributes.
  • C.1 Overview of Annotation Stages: The pipeline begins with eight indoor scenarios and generates specific objects, bounded partonomies, multiple physical variants, and multiple state configurations.The scenarios are Kitchen, Living Room, Bedroom, Bathroom, Garage, Home Office, Dining Room, and Garden.
  • C.2 Core Hyperparameters Details: Core hyperparameter tables define data-generation settings, combinatorial and annotation controls for coverage and tractability, and inference-time settings for scalable, fault-tolerant processing.These controls govern the generation and annotation pipeline described in the surrounding sections.
  • C.3 Core Prompt Details: Core prompts constrain entity decomposition to bounded, non-overlapping parts with explicit functional boundaries and require physical-attribute annotation conditioned on part context and connections.The decomposition prompt allows up to 8 parts and includes hidden components or components with useful affordances; the physical prompt provides all parts, connected parts, and connection descriptions.

D Task Creation Pipeline Details … D.3 Core Prompt Details

The task creation pipeline constructs grounded, creativity-demanding benchmark tasks by selecting a dominant gold affordance, controlling distractors, and packaging judge-checkable solutions. Its prompts enforce realistic first-person narratives, conceal the intended mechanism, and produce structured comparisons and solutions grounded in annotated attributes.

  • D Task Creation Pipeline Details: Each task enforces a preferred entity-part-action gold affordance, controlled semantic distractors, and a grounded first-person narrative with judge-checkable constraints.These requirements define rigor, ambiguity control, and verifiability throughout task generation.
  • D.1 Overview of Sampling Stages: The pipeline clusters affordances by scenario, stratifies gold sampling by cluster size and normal-versus-emergency level, then instantiates answer-hidden first-person task prompts.Semantic lookup spaces preserve traceability to scenarios, entities, parts, and annotation fields.
  • D.1 Overview of Sampling Stages: Gold candidates pass intra-entity dominance checks, while distractors are sampled from near and far semantic pools and filtered into dissimilar, mixed, or similar-but-not-better tiers.Comparison and filtering exclude entities whose parts outperform the gold or create excessive ambiguity.
  • D.1 Overview of Sampling Stages: Final task assembly combines gold annotations, selected entities, judge outputs, scene items, a first-person environment, and a structured four-step solution aligned with recipient, use, environment, and application mechanics.These aligned fields support direct judge-side verification.
  • D.2 Core Hyperparameters Details: Stage-specific prompts use strict JSON schemas to enforce groundedness, comparability, decision consistency, ambiguity reduction, gold concealment, and machine-parsable outputs.The data are generated with GPT-5.2 under human supervision and iterative prompt refinement.
  • D.3 Core Prompt Details: Gold-to-task prompting starts from a concrete recipient and real-world situation, asks what to use or how, and forbids paraphrasing or revealing the gold entity, part, or mechanism.The intended narrative names the recipient, context, problem, and goal while keeping the affordance implicit.
  • D.3 Core Prompt Details: Final packaging creates scene items, an environment description, and exactly four solution keys covering recipient preparation, tool setup, environmental setup, and applying the affordance, with judge-verification notes.All steps must remain grounded in annotated conditions and listed attributes without inventing unsupported actions.

E Experiment Details · F Analysis Details

The experiments use largely deterministic, long-context evaluations with structured outputs, while analysis scores grounded task solutions through objective metrics and rubric-based subjective judgments. The rubric assesses prerequisite coverage, attribute grounding, correctness, and physical feasibility against gold affordances and solutions.

  • E Experiment Details: Most models use temperature 0, while GPT-5-Mini and GPT-5-Nano use default sampling because they lack adjustable temperature.This setting is used for the main-table results.
  • E Experiment Details: All models receive a maximum output length of at least 16K tokens, which is empirically sufficient for generation.The main prompt asks models to identify the best entity part and return exact-match gold_entity and gold_part fields plus detailed how_to_use instructions.
  • E Experiment Details: The task prompt provides environment content, available entities with full descriptions, and other scene items before requesting a creative entity-part solution.The answer format requires case-sensitive exact entity and part names alongside detailed usage instructions.
  • E Experiment Details: The main setting reports two objective metrics and six subjective metrics, uniformly rescaled from 0–2 to 1–5.Subjective metrics are applied only to answers counted as Gold Correct; other failure cases are analyzed later.
  • E Experiment Details: Gemini-3.1-Flash-Lite judges whether predicted how_to_use instructions are feasible against the gold affordance and gold solution.The evaluator is instructed to use strict, evidence-based field-level judgments.
  • E Experiment Details: The rubric separately scores environment, tool-part, and recipient prerequisite coverage, each using 0/1/2 or NA where applicable.Scores distinguish absent, incomplete, and well-covered conditions, with recipient assumptions also marked False when unmet.
  • E Experiment Details: Attributes grounding, prediction correctness, and action feasibility each use a 0–2 scale to assess enabling attributes, alignment with gold solutions, and practical physical workability.Action feasibility is judged independently of the gold solution and penalizes impossible, unsafe, implausible, or incomplete actions.
  • E Experiment Details: Every rubric reason must cite concrete evidence, and unclear or missing evidence receives a stricter lower score.The evaluator must return exactly 12 JSON fields containing reasons and scores for all six rubric dimensions.

F.1 Inference setting’s impact on performance

Evaluation uses a 10% subset for Section 6.3, while full-test validation shows consistent temperature and inference-mode trends, supporting the subset’s representativeness. Interactive mode substantially hurts performance, whereas CoT provides no clear overall benefit.

  • Sampling and evaluation scope: 10% of the dataset, containing about 1.4K examples, is evaluated in Section 6.3 because it supports reliable analysis while limiting substantial experimental costs.Full 14K-example evaluation is especially expensive for interactive mode because it requires multi-turn inference.
  • Sampling and evaluation scope: Full 14K-test-set results show trends consistent with the main text, supporting the representativeness of the sampled subset.Selected models were evaluated across sampling temperatures and inference modes.
  • Temperature sampling: GPT-5.2 and Llama-3-70B tend to improve with temperature sampling, whereas smaller models generally experience performance drops.These temperature effects are observed on the full test set.
  • Inference mode: Interactive mode substantially reduces performance, while CoT mode causes only small fluctuations and offers no clear overall benefit.This pattern appears in the full-set inference-mode experiments.

F.2 Fine-grained analysis on auxiliary metrics

Figures 14–17 provide a fine-grained analysis of auxiliary metrics across four aspects: gold affordance cluster size, gold affordance emergency level, distractor similarity, and distractor number.

  • Fine-grained auxiliary-metric analysis: Figures 14–17 report complete auxiliary-metric results for gold affordance cluster size, gold affordance emergency level, distractor similarity, and distractor number.The analysis covers these four distinct aspects.

F.3 Error analysis details

The error analysis uses deterministic LLM judgments on sampled failures, evaluating predictions both independently and against gold solutions. It decomposes prediction quality into condition coverage, attribute grounding, action feasibility, and comparative reasonableness.

  • Judging protocol: Gemini-3.1-Flash-Lite judges a 10% sample of each model’s failures at temperature zero, following the main-results evaluation criteria.The protocol assesses constraint satisfaction, physical grounding, and action feasibility.
  • Judging protocol: The judge first evaluates each prediction independently, then compares it with the gold solution and its supporting rationale to determine whether the gold is preferable.The comparative judgment assesses whether the prediction remains reasonable relative to the gold choice.
  • Fine-grained metrics: Fine-grained judgment scores environment, use, and recipient condition coverage, plus attributes grounding and action feasibility.Condition fields use 0/1/2 or NA, while attributes grounding and action feasibility use 0/1/2 scales.
  • Fine-grained metrics: The rubric requires concrete evidence for every reasoning field and assigns stricter, lower scores when evidence is unclear or missing.Action feasibility distinguishes impossible or unsafe actions from partially workable and operationally correct actions.

F.4 Attribution analysis details

The appendix validates its attribution judge, specifies a taxonomy for failed physical tool-use predictions, and reports fine-grained model-level breakdowns. Across models, physical invalidity remains the dominant failure reason, with consistent category trends.

  • Method and reliability: Attribution analysis samples 10% of cases, uses Gemini-3.1-Flash-Lite at judging temperature 0.0, and repeats categorization with Qwen3-32B and GPT-5-Mini.The judge receives the gold comparison reason and assigns exactly one primary contributing factor plus additional categories when appropriate.
  • Method and reliability: Gemini’s primary-category predictions overlap with Qwen3-32B in 83.3% of cases and with GPT-5-Mini in 90.91% of cases.These agreement rates support using Gemini as the categorization judge.
  • Attribution taxonomy: The taxonomy covers physical invalidity, practical invalidity, risk or requirement mismatch, and comparative inferiority rather than true failure.Fine-grained categories include hallucinated affordances, affordance and performance mismatches, destructive workarounds, impractical processes, safety risks, constraint violations, and preference-sensitive inferiority.
Loading 2605.02910v2…