Source-linked AI summary

PlanViz: Evaluating Planning-Oriented Image Generation and Editing for Computer-Use Tasks

Junxian Li, Kai Liu, Leyang Chen, Weida Wang, Zhixin Wang, Jiaqi Xu, Fan Li, Renjing Pei, Linghe Kong, Yulun Zhang

arXiv:2602.06663v2cs.CV

TL;DR

UMMs are powerful at natural-image generation, but their ability to plan and manipulate functional computer-use visuals remains underexplored. PlanViz addresses this gap with a human-curated benchmark spanning three planning tasks and PlanScore for correctness, visual quality, and efficiency. Experiments reveal inconsistent performance, with particularly strong limitations in planning-intensive editing and differences between proprietary and open-source models.

  • Problem

    UMMs lack established evaluation evidence for generating and editing functional computer-use visuals that require spatial reasoning, structural precision, and semantic consistency.

  • Method

    PlanViz is a human-annotated benchmark covering route planning, workflow diagramming, and web&UI displaying, with task-adaptive PlanScore evaluation.

  • Results

    Performance is inconsistent across models and tasks, with image editing remaining especially difficult and proprietary models generally outperforming open-source counterparts.

  • Takeaways & Limitations

    PlanViz exposes limitations in satisfying complex planning constraints beyond surface-level image quality and motivates research bridging understanding and generation for computer-use scenarios.

Abstract

from arXiv · show

Unified multimodal models (UMMs) have shown impressive capabilities in generating natural images and supporting multimodal reasoning. However, their potential in supporting computer-use planning tasks, which are closely related to our lives, remain underexplored. Image generation and editing in computer-use tasks require capabilities like spatial reasoning and procedural understanding, and it is still unknown whether UMMs have these capabilities to finish these tasks or not. Therefore, we propose PlanViz, a new benchmark designed to evaluate image generation and editing for computer-use tasks. To achieve the goal of our evaluation, we focus on sub-tasks which frequently involve in daily life and require planning. Specifically, three representative sub-tasks are designed: route planning, work diagramming, and web&UI displaying. We address challenges in data quality ensuring by curating human-annotated questions and reference images, and a quality control process. For detailed and exact evaluation, a task-adaptive score, PlanScore, is proposed. The score helps understanding the correctness, visual quality and efficiency of generated images. Through experiments, we highlight key limitations and opportunities for future research on this topic.

1 Introduction

PlanViz targets an underexplored capability of unified multimodal models: planning-oriented image generation and editing for functional computer-use visuals. It introduces a human-curated benchmark and PlanScore, then finds substantial performance inconsistency, especially for editing and open-ended tasks.

  • Motivation: PlanViz evaluates UMMs on computer-use visuals, where functional content demands more structural precision and semantic consistency than typical image generation.These visuals include graphical user interfaces, structured workflows, and slide layouts used in professional and personal life.
  • Benchmark: The benchmark covers route planning, workflow diagramming, and web&UI displaying, with open-ended and closed-ended question categories.The tasks are selected because they are common in daily life and include planning steps.
  • Contributions: PlanViz combines manually collected and annotated data, quality control, and PlanScore for task-adaptive evaluation of correctness, visual quality, and efficiency.The benchmark is designed to provide a detailed and exact view of image generation and editing capabilities in computer-use tasks.
  • Findings: Experiments find inconsistent performance across nearly all models and tasks, with GPT-Image-1 performing best overall while still struggling with editing.The authors relate this pattern to mostly reactive generation behavior facing explicit planning requirements.
  • Findings: Models generally perform better on closed-ended than open-ended questions, indicating that detailed, fine-grained guidance is important for these tasks.The result is reported as a finding from the benchmark experiments rather than as a general conclusion about all image-generation settings.

2 Related Work

Related benchmarks mainly evaluate natural-image generation, while newer benchmarks extend toward logically meaningful or domain-knowledge-intensive images; PlanViz focuses on both generation and editing for everyday computer-use tasks.

  • Existing benchmarks: Previous UMM benchmarks typically contain natural-image generation, with some recent benchmarks adding logically meaningful images or images requiring domain knowledge.The related-work discussion positions these settings as distinct from everyday computer-use visuals.

3 PlanViz

PlanViz defines planning around translating multiple goals and constraints into functional computer-use visuals, constructs a human-checked benchmark, and evaluates outputs with task-adaptive MLLM judging.

  • Benchmarking Objective: Planning in PlanViz means translating multiple complex goals and constraints into a functional visual outcome, rather than classical state-action search.The benchmark targets route planning, workflow diagramming, and web&UI displaying.
  • Benchmarking Objective: The benchmark uses real-world computer-use scenarios and includes route planning, workflow diagramming, and web&UI displaying as representative planning sub-tasks.Its data distribution and question topics are organized around these three task categories.
  • Data Construction: Nearly 500 collected images were filtered to 60 high-quality images suitable for each editing sub-task, with independent scenarios and recorded web&UI state changes.Collection rules require clear roads, visible layouts, recognizable text, and real-world maps or GUI screens.
  • Data Construction: Blind quality checking gave the annotations a total score of 0.96 for question difficulty and reference-image or keypoint correctness.Independent annotators assessed whether each question was appropriately difficult and whether its references and keypoints were correct.
  • Score Judgement Pipeline: PlanScore measures correctness, visual quality, and efficiency using task-adaptive MLLM-as-judge evaluation with detailed rating criteria and reference images.Correctness uses annotated key points, while visual quality and efficiency are normalized from 0–5 ratings to the [0,1] scale.
  • Data Construction: Data construction proceeds through high-quality collection and cleaning, human annotation, quality checking, and prompt-style transformation.The stages are ordered from collection through annotation and checking to diverse surface-level prompt wording.
  • Evaluation Reliability: Human and cross-model studies support the reliability of the proposed judgment method for reflecting PlanViz performance.The studies compare model judgments with judgments from 10 human annotators and GPT-5 on sampled outputs.

4 Experiment

Experiments evaluate recent UMMs and image-generation/editing models on PlanViz, revealing substantial inconsistencies across tasks, metrics, and model types. Editing, route planning, and precise constraint satisfaction remain especially difficult.

  • Evaluation setup: PlanViz evaluates recent open-source and closed-source UMMs alongside prior image-generation and editing models using correctness, visual quality, efficiency, and average scores.Correctness is treated as the most important score for assessing planning and useful image generation.
  • Main results: 0.61-0.67 is the reported range for SOTA image-editing performance, while proprietary generation models such as Seedream 4.5 reach 0.88 and open-source models typically remain below 0.55.For Cor specifically, editing SOTA models score 0.3-0.5, whereas proprietary generation models approach 0.8-0.9.
  • Main results: Image editing consistently trails generation, with Qwen-Image reaching Cor values of 0.46, 0.77, and 0.77 in generation but no more than 0.15 in editing.Seedream-4.5’s overall score also declines from 0.88 in generation to 0.67 in editing.
  • Main results: 0.81 is the SOTA Cor for route-planning generation, below approximately 0.90 for workflow design and web&UI displaying; editing remains challenging across all three tasks.The authors relate these differences to reactive generation behavior facing explicit planning requirements.
  • Main results: Thinking mechanisms have task-dependent effects: image-generation Cor improves by up to 0.35 for Bagel and 0.13 for UniPic2-Metaquery, but editing gains are marginal or sometimes negative.Step1X-Edit-v1p2 improves overall average scores by only +0.04, while some thinking-mode configurations worsen performance.
  • Main results: High Vis can coexist with near-zero Cor, while higher Cor may coincide with lower Vis and Ef, producing visually plausible but semantically mismatched or degraded outputs.Janus-4o in route-planning editing exemplifies the trade-off between correctness and other scores.
  • Analysis: Score distributions show GPT-Image-1 concentrated near 0.8-1.0 for route-planning generation, whereas open-source models cluster near 0.0-0.2 and editing distributions remain low.Vis distributions are generally high for generation and shift upward relative to Cor in editing, supporting a semantic-visual decoupling.
  • Analysis: Prompt-style experiments find open-source UMMs more sensitive than GPT-Image-1, with higher standard-deviation-to-mean ratios and variable performance.The study uses 10 questions and 10 language styles per question to examine this influence.

5 Conclusion

PlanViz is introduced as a benchmark for planning-oriented image generation and editing in computer-use tasks. It exposes inconsistent model performance and persistent difficulties with planning-intensive editing and constraint satisfaction beyond surface-level image quality.

  • 5 Conclusion: PlanViz combines diverse computer-use subtasks, human-annotated data, and a task-adaptive evaluation metric for benchmarking planning-oriented image generation and editing.The benchmark targets both generation and editing rather than surface-level image quality alone.
  • 5 Conclusion: Current UMMs show inconsistent performance and especially clear limitations on planning-intensive image editing.The conclusion identifies satisfying task constraints as a major challenge beyond surface-level image quality.
  • 5 Conclusion: PlanViz is intended to support research bridging understanding and generation for real-world computer-use scenarios.This consequence is stated as the benchmark’s intended research role.
Loading 2602.06663v2…