paper-with-me

Papers

SOC: Semantic-Assisted Object Cluster for Referring Video Object Segmentation

2023-05-26 · NeurIPS 2023 11 · Zhuoyan Luo, Yicheng Xiao, Yong liu, Shuyan Li, Yitong Wang, Yansong Tang, Xiu Li, Yujiu Yang

This paper studies referring video object segmentation (RVOS) by boosting video-level visual-linguistic alignment. Recent approaches model the RVOS task as a sequence prediction problem and perform multi-modal interaction as well as segmentation for each frame separately. However, the lack of a global view of video content leads to difficulties in effectively utilizing inter-frame relationships and understanding textual descriptions of object temporal variations. To address this issue, we propose Semantic-assisted Object Cluster (SOC), which aggregates video content and textual guidance for unified temporal modeling and cross-modal alignment. By associating a group of frame-level object embeddings with language tokens, SOC facilitates joint space learning across modalities and time steps. Moreover, we present multi-modal contrastive supervision to help construct well-aligned joint space at the video level. We conduct extensive experiments on popular RVOS benchmarks, and our method outperforms state-of-the-art competitors on all benchmarks by a remarkable margin. Besides, the emphasis on temporal coherence enhances the segmentation stability and adaptability of our method in processing text expressions with temporal variations. Code will be available.

📄 PDF Abstract BibTeX arXiv:2305.17011

Code (1)

RobertLuo1/NeurIPS2023_SOC 공식 구현 pytorch

Tasks

cross-modal alignmentObjectReferring Expression SegmentationReferring Video Object SegmentationSegmentationSemantic SegmentationVideo Object SegmentationVideo Semantic Segmentation

Similar Papers 제목 키워드 기반

EventRR: Event Referential Reasoning for Referring Video Object Segmentation

2025-08-10 · Huihui Xu, Jiashi Lin, Haoyu Chen, Junjun He 외 arxiv

Referring Video Object Segmentation (RVOS) aims to segment out the object in a video referred by an expression. Current RVOS methods view referring expressions as unstructured sequences, neglecting their crucial semantic…

Referring Video Object SegmentationReferring Expression

VideoOrion: Tokenizing Object Dynamics in Videos

2024-11-25 · Yicheng Feng, Yijiang Li, Wanpeng Zhang, Sipeng Zheng 외

We present VideoOrion, a Video Large Language Model (Video-LLM) that explicitly captures the key semantic information in videos--the spatial-temporal dynamics of objects throughout the videos. VideoOrion employs expert v…

Language ModelingLanguage ModellingLarge Language ModelObject+2

The Second Place Solution for The 4th Large-scale Video Object Segmentation Challenge--Track 3: Referring Video Object Segmentation

2022-06-24 · Leilei Cao, Zhuang Li, Bo Yan, Feng Zhang 외

The referring video object segmentation task (RVOS) aims to segment object instances in a given video referred by a language expression in all video frames. Due to the requirement of understanding cross-modal semantics w…

Objectobject-detectionObject DetectionReferring Video Object Segmentation+5

Referring Multi-Object Tracking

2023-03-06 · CVPR 2023 1 · Dongming Wu, Wencheng Han, Tiancai Wang, Xingping Dong 외

Existing referring understanding tasks tend to involve the detection of a single text-referred object. In this paper, we propose a new and general referring understanding task, termed referring multi-object tracking (RMO…

Multi-Object TrackingObjectObject TrackingReferring Multi-Object Tracking

SimToken: A Simple Baseline for Referring Audio-Visual Segmentation

2025-09-22 · Dian Jin, Yanghao Zhou, Jinxing Zhou, Jiaqi Ma 외 arxiv

Referring Audio-Visual Segmentation (Ref-AVS) aims to segment specific objects in videos based on natural language expressions involving audio, vision, and text information. This task poses significant challenges in cros…

Object Localization