paper-with-me

홈 › Papers

Temporal-Conditional Referring Video Object Segmentation with Noise-Free Text-to-Video Diffusion Model

2025-08-19 · Ruixin Zhang, Jiaqing Fan, Yifan Liao, Qian Qiao, Fanzhang Li arxiv

Referring Video Object Segmentation (RVOS) aims to segment specific objects in a video according to textual descriptions. We observe that recent RVOS approaches often place excessive emphasis on feature extraction and temporal modeling, while relatively neglecting the design of the segmentation head. In fact, there remains considerable room for improvement in segmentation head design. To address this, we propose a Temporal-Conditional Referring Video Object Segmentation model, which innovatively integrates existing segmentation methods to effectively enhance boundary segmentation capability. Furthermore, our model leverages a text-to-video diffusion model for feature extraction. On top of this, we remove the traditional noise prediction module to avoid the randomness of noise from degrading segmentation accuracy, thereby simplifying the model while improving performance. Finally, to overcome the limited feature extraction capability of the VAE, we design a Temporal Context Mask Refinement (TCMR) module, which significantly improves segmentation quality without introducing complex designs. We evaluate our method on four public RVOS benchmarks, where it consistently achieves state-of-the-art performance.

📄 PDF Abstract BibTeX arXiv:2508.13584

Code (0)

등록된 구현이 없습니다.

Tasks

Referring Video Object Segmentation

Similar Papers 제목 키워드 기반

Temporal Prompting Matters: Rethinking Referring Video Object Segmentation

2025-10-08 · Ci-Siang Lin, Min-Hung Chen, I-Jieh Liu, Chien-Yi Wang 외 arxiv

Referring Video Object Segmentation (RVOS) aims to segment the object referred to by the query sentence in the video. Most existing methods require end-to-end training with dense mask annotations, which could be computat…

Referring Video Object Segmentation

Temporally Consistent Referring Video Object Segmentation with Hybrid Memory

2024-03-28 · Bo Miao, Mohammed Bennamoun, Yongsheng Gao, Mubarak Shah 외

Referring Video Object Segmentation (R-VOS) methods face challenges in maintaining consistent object segmentation due to temporal context variability and the presence of other visually similar objects. We propose an end-…

HTRObjectReferring Expression SegmentationReferring Video Object Segmentation+5

EventRR: Event Referential Reasoning for Referring Video Object Segmentation

2025-08-10 · Huihui Xu, Jiashi Lin, Haoyu Chen, Junjun He 외 arxiv

Referring Video Object Segmentation (RVOS) aims to segment out the object in a video referred by an expression. Current RVOS methods view referring expressions as unstructured sequences, neglecting their crucial semantic…

Referring Video Object SegmentationReferring Expression

Temporal-Enhanced Multimodal Transformer for Referring Multi-Object Tracking and Segmentation

2024-10-17 · Changcheng Xiao, Qiong Cao, Yujie Zhong, Xiang Zhang 외

Referring multi-object tracking (RMOT) is an emerging cross-modal task that aims to locate an arbitrary number of target objects and maintain their identities referred by a language expression in a video. This intricate …

Multi-Object TrackingMulti-Object Tracking and SegmentationObject TrackingReferring Multi-Object Tracking+3

Exploring Pre-trained Text-to-Video Diffusion Models for Referring Video Object Segmentation

2024-03-18 · Zixin Zhu, Xuelu Feng, Dongdong Chen, Junsong Yuan 외

In this paper, we explore the visual representations produced from a pre-trained text-to-video (T2V) diffusion model for video understanding tasks. We hypothesize that the latent representation learned from a pretrained …

Referring Video Object SegmentationSemantic SegmentationVideo Object SegmentationVideo Semantic Segmentation+1