paper-with-me

Papers

InstanceV: Instance-Level Video Generation

2025-11-28 · Yuheng Chen, Teng Hu, Jiangning Zhang, Zhucun Xue, Ran Yi, Lizhuang Ma arxiv

Recent advances in text-to-video diffusion models have enabled the generation of high-quality videos conditioned on textual descriptions. However, most existing text-to-video models rely solely on textual conditions, lacking general fine-grained controllability over video generation. To address this challenge, we propose InstanceV, a video generation framework that enables i) instance-level control and ii) global semantic consistency. Specifically, with the aid of proposed Instance-aware Masked Cross-Attention mechanism, InstanceV maximizes the utilization of additional instance-level grounding information to generate correctly attributed instances at designated spatial locations. To improve overall consistency, We introduce the Shared Timestep-Adaptive Prompt Enhancement module, which connects local instances with global semantics in a parameter-efficient manner. Furthermore, we incorporate Spatially-Aware Unconditional Guidance during both training and inference to alleviate the disappearance of small instances. Finally, we propose a new benchmark, named InstanceBench, which combines general video quality metrics with instance-aware metrics for more comprehensive evaluation on instance-level video generation. Extensive experiments demonstrate that InstanceV not only achieves remarkable instance-level controllability in video generation, but also outperforms existing state-of-the-art models in both general quality and instance-aware metrics across qualitative and quantitative evaluations.

📄 PDF Abstract BibTeX arXiv:2511.23146

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption

2024-12-12 · CVPR 2025 1 · Tiehan Fan, Kepan Nan, Rui Xie, Penghao Zhou 외

Text-to-video generation has evolved rapidly in recent years, delivering remarkable results. Training typically relies on video-caption paired data, which plays a crucial role in enhancing generation performance. However…

Text-to-Video GenerationVideo Generation

Improving Generalized Visual Grounding with Instance-aware Joint Learning

2025-09-17 · Ming Dai, Wenxuan Cheng, Jiang-Jiang Liu, Lingfeng Yang 외 arxiv

Generalized visual grounding tasks, including Generalized Referring Expression Comprehension (GREC) and Segmentation (GRES), extend the classical visual grounding paradigm by accommodating multi-target and non-target sce…

Generalized Referring Expression ComprehensionSemantic SegmentationVisual Grounding

SegmentAnyTreeV2: Scaling Transformer-Based Tree Instance Segmentation Across Sensors, Platforms, and Forests

2026-06-06 · Maciej Wielgosz, Stefano Puliti, Rasmus Astrup arxiv

We present SegmentAnyTreeV2, a sensor- and platform-agnostic framework for semantic and instance segmentation of forest point clouds. The model combines a serialization-based Point Transformer v3 backbone with a lightwei…

Domain GeneralizationInstance SegmentationPoint Clouds

IM-Zero: Instance-level Motion Controllable Video Generation in a Zero-shot Manner

2025-01-01 · CVPR 2025 1 · YuYang Huang, Yabo Chen, Li Ding, Xiaopeng Zhang 외

Controllability of video generation has been recently concerned in addition to the quality of generated videos. The main challenge to controllable video generation is to synthesize videos based on user-specified inst…

Motion GenerationText-to-Video GenerationVideo Generation

VRMDiff: Text-Guided Video Referring Matting Generation of Diffusion

2025-03-11 · Lehan Yang, Jincen Song, Tianlong Wang, Daiqing Qi 외

We propose a new task, video referring matting, which obtains the alpha matte of a specified instance by inputting a referring caption. We treat the dense prediction task of matting as video generation, leveraging the te…

Image MattingVideo AlignmentVideo Generation