paper-with-me

Papers

PlanViz: Evaluating Planning-Oriented Image Generation and Editing for Computer-Use Tasks

2026-02-06 · Junxian Li, Kai Liu, Leyang Chen, Weida Wang, Zhixin Wang, Jiaqi Xu, Fan Li, Renjing Pei, Linghe Kong, Yulun Zhang arxiv

Unified multimodal models (UMMs) have shown impressive capabilities in generating natural images and supporting multimodal reasoning. However, their potential in supporting computer-use planning tasks, which are closely related to our lives, remain underexplored. Image generation and editing in computer-use tasks require capabilities like spatial reasoning and procedural understanding, and it is still unknown whether UMMs have these capabilities to finish these tasks or not. Therefore, we propose PlanViz, a new benchmark designed to evaluate image generation and editing for computer-use tasks. To achieve the goal of our evaluation, we focus on sub-tasks which frequently involve in daily life and require planning. Specifically, three representative sub-tasks are designed: route planning, work diagramming, and web&UI displaying. We address challenges in data quality ensuring by curating human-annotated questions and reference images, and a quality control process. For detailed and exact evaluation, a task-adaptive score, PlanScore, is proposed. The score helps understanding the correctness, visual quality and efficiency of generated images. Through experiments, we highlight key limitations and opportunities for future research on this topic.

📄 PDF Abstract BibTeX arXiv:2602.06663

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal ReasoningSpatial ReasoningImage Generation

Similar Papers 제목 키워드 기반

Propose, Assess, Search: Harnessing LLMs for Goal-Oriented Planning in Instructional Videos

2024-09-30 · Md Mohaiminul Islam, Tushar Nagarajan, Huiyu Wang, Fu-Jen Chu 외

Goal-oriented planning, or anticipating a series of actions that transition an agent from its current state to a predefined objective, is crucial for developing intelligent assistants aiding users in daily procedural tas…

TopKG: Target-oriented Dialog via Global Planning on Knowledge Graph

2022-10-01 · COLING 2022 10 · Zhitong Yang, Bo wang, Jinfeng Zhou, Yue Tan 외

Target-oriented dialog aims to reach a global target through multi-turn conversation. The key to the task is the global planning towards the target, which flexibly guides the dialog concerning the context. However, exist…

Response Generation

Towards Evaluating Plan Generation Approaches with Instructional Texts

2020-01-13 · Debajyoti Paul Chowdhury, Arghya Biswas, Tomasz Sosnowski, Kristina Yordanova

Recent research in behaviour understanding through language grounding has shown it is possible to automatically generate behaviour models from textual instructions. These models usually have goal-oriented structure and a…

SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene Consistency

2025-10-27 · Quanjian Song, Donghao Zhou, Jingyu Lin, Fei Shen 외 arxiv

Recent text-to-image models have revolutionized image generation, but they still struggle with maintaining concept consistency across generated images. While existing works focus on character consistency, they often over…

Story GenerationImage Generation

Target-constrained Bidirectional Planning for Generation of Target-oriented Proactive Dialogue

2024-03-10 · Jian Wang, Dongding Lin, Wenjie Li

Target-oriented proactive dialogue systems aim to lead conversations from a dialogue context toward a pre-determined target, such as making recommendations on designated items or introducing new specific topics. To this …

Dialogue Generation