paper-with-me

Papers

CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration

2026-07-06 · Hairui Zhu, Yiying Yang, Tengjin Weng, Ziyu Lu, Xiao Yao, Xiaoyang Ye, Lin Ma, Wenhao Jiang hf

Complex image creation and editing often require more than a single generation or editing model. A user request may involve synthesizing images, localizing objects, segmenting regions, editing selected content, compositing intermediate assets, reading text, and enhancing the final result. Such tasks shift multimodal agents from perception-augmented reasoning to manipulation-centered visual creation, where tools must actively transform visual states rather than merely inspect them. However, existing multimodal tool-use agents are mostly optimized for perception, search, or domain-specific editing, and lack large-scale supervision for executable image-creation trajectories. In this paper, we introduce CanvasCraft, a large-scale multimodal tool-use dataset for complex image creation and editing, and CanvasAgent, a tool-augmented multimodal agent that learns to orchestrate heterogeneous visual tools through multi-turn interaction. CanvasCraft contains 140K fully annotated executable trajectories and 10K RL task specifications. CanvasAgent is first trained with SFT to learn executable reasoning-action trajectories, and is then optimized with GRPO using a hybrid reward that combines outcome- and process-level signals. During rollout, CanvasAgent inspects intermediate results, tracks visual assets, and adapts tool decisions to the evolving visual state. Experiments evaluate both final image quality and trajectory behavior, demonstrating the effectiveness of CanvasAgent and the proposed dataset for complex multi-tool image creation workflows.

📄 PDF Abstract BibTeX arXiv:2607.05465

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MeshPad: Interactive Sketch-Conditioned Artist-Designed Mesh Generation and Editing

2025-03-03 · Haoxuan Li, Ziya Erkoc, Lei LI, Daniele Sirigatti 외

We introduce MeshPad, a generative approach that creates 3D meshes from sketch inputs. Building on recent advances in artist-designed triangle mesh generation, our approach addresses the need for interactive mesh creatio…

CMD: Controllable Multiview Diffusion for 3D Editing and Progressive Generation

2025-05-11 · Peng Li, Suizhi Ma, Jialiang Chen, YuAn Liu 외

Recently, 3D generation methods have shown their powerful ability to automate 3D model creation. However, most 3D generation methods only rely on an input image or a text prompt to generate a 3D model, which lacks the co…

3D Generation

ACE: Anti-Editing Concept Erasure in Text-to-Image Models

2025-01-03 · CVPR 2025 1 · ZiHao Wang, Yuxiang Wei, Fan Li, Renjing Pei 외

Recent advance in text-to-image diffusion models have significantly facilitated the generation of high-quality images, but also raising concerns about the illegal creation of harmful content, such as copyrighted images. …

HyperDreamer: Hyper-Realistic 3D Content Generation and Editing from a Single Image

2023-12-07 · Tong Wu, Zhibing Li, Shuai Yang, Pan Zhang 외

3D content creation from a single image is a long-standing yet highly desirable task. Recent advances introduce 2D diffusion priors, yielding reasonable results. However, existing methods are not hyper-realistic enough f…

Semantic Segmentation

LAVE: LLM-Powered Agent Assistance and Language Augmentation for Video Editing

2024-02-15 · Bryan Wang, Yuliang Li, Zhaoyang Lv, Haijun Xia 외

Video creation has become increasingly popular, yet the expertise and effort required for editing often pose barriers to beginners. In this paper, we explore the integration of large language models (LLMs) into the video…

Video Editing