paper-with-me

홈 › Papers

From Dense Prediction to Visual Editing: Structured Supervision for Unified Image and Video Creation

2026-08-13 · Zhefan Rao, Bin Zou, Xuanhua He, Chong Hou Choi, Yanheng Li, Rui Liu, Haoxuan Che, Qifeng Chen arxiv

Unified image and video creation requires a model to follow diverse instructions while preserving identity, geometry, and temporal structure from visual context. However, semantic-only conditioning and creation-only training do not explicitly supervise the local structure needed for precise, temporally consistent editing. We therefore formulate depth and surface-normal prediction as image-form denoising targets, using these dense tasks as structured visual supervision within the same creation interface. Our framework decouples semantic interpretation from spatially aligned visual injection while sharing one multimodal diffusion transformer (MMDiT) backbone across all tasks. Mutual Context Attention (MCA), a paired-video data-construction procedure, and a progressive training curriculum then connect the learned structural cues to temporally localized editing and reference-conditioned creation. A single checkpoint obtains the highest overall score in the reported comparison of unified systems (4.15); adding dense supervision improves OpenVE Overall from 3.98 to 4.06 and Local Add from 3.92 to 4.18. These results support a deliberately bounded conclusion: perception-oriented dense supervision transfers useful structural knowledge to downstream creation, especially editing locality and preservation; we do not claim superiority as a standalone dense predictor.

📄 PDF Abstract BibTeX arXiv:2608.14740

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SWEET: Sparse World Modeling with Image Editing for Embodied Task Execution

2026-05-19 · Yiren Song, Yihan Wang, Xiyao Deng, Zhuoran Yan 외 arxiv

Visual prediction has emerged as a promising paradigm for embodied control, where future observations are generated and then translated into actions. However, dense video generation is computationally expensive and often…

Robot ManipulationVideo GenerationImage Editing

ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

2025-01-09 · Xingyu Fu, Minqian Liu, Zhengyuan Yang, John Corring 외

Structured image understanding, such as interpreting tables and charts, requires strategically refocusing across various structures and texts within an image, forming a reasoning sequence to arrive at the final answer. H…

Visual Question Answering (VQA)Visual Reasoning

STRAP: Structured Object Affordance Segmentation with Point Supervision

2023-04-17 · Leiyao Cui, Xiaoxue Chen, Hao Zhao, Guyue Zhou 외

With significant annotation savings, point supervision has been proven effective for numerous 2D and 3D scene understanding problems. This success is primarily attributed to the structured output space; i.e., samples wit…

ObjectScene Understanding

UnsOcc: 3D Semantic Occupancy Prediction in Unstructured Scene via Rendering Fusion

2026-06-02 · Ye Wu, Ruiqi Song, Baiyong Ding, Nanxin Zeng 외 arxiv

Unstructured scenes present unique challenges for autonomous driving, as irregular obstacles and sparse scene layouts undermine the effectiveness of traditional perception methods such as 3D object detection. 3D semantic…

2D Semantic Segmentation3D Object DetectionAutonomous Driving

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing

2026-07-09 · Feng Wang, Canmiao Fu, Zhipeng Huang, Chen Li 외 arxiv

Recent unified multimodal models show a single architecture can jointly perform vision/language understanding and image generation/editing. However, they repeatedly feed all historical visual and textual inputs into a sh…

Reinforcement LearningImage Generation