paper-with-me

Papers

URVOS: Unified Referring Video Object Segmentation Network with a Large-Scale Benchmark

2020-08-01 · ECCV 2020 8 · Seonguk Seo, Joon-Young Lee, Bohyung Han

We propose a unified referring video object segmentation network (URVOS). URVOS takes a video and a referring expression as inputs, and estimates the {object masks} referred by the given language expression in the whole video frames. Our algorithm addresses the challenging problem by performing language-based object segmentation and mask propagation jointly using a single deep neural network with a proper combination of two attention models. In addition, we construct the first large-scale referring video object segmentation dataset called Refer-Youtube-VOS. We evaluate our model on two benchmark datasets including ours and demonstrate the effectiveness of the proposed approach. The dataset is released at \url{https://github.com/skynbe/Refer-Youtube-VOS}.

📄 PDF Abstract BibTeX

Code (1)

skynbe/Refer-Youtube-VOS 공식 구현 pytorch

Tasks

ObjectOne-shot visual object segmentationReferring ExpressionReferring Expression SegmentationReferring Video Object SegmentationSegmentationSemantic SegmentationVideo Object SegmentationVideo Semantic Segmentation

Similar Papers 제목 키워드 기반

Show Me When and Where: Towards Referring Video Object Segmentation in the Wild

2026-03-15 · Mingqi Gao, Jinyu Yang, Jingnan Luo, Xiantong Zhen 외 arxiv

Referring video object segmentation (RVOS) has recently generated great popularity in computer vision due to its widespread applications. Existing RVOS setting contains elaborately trimmed videos, with text-referred obje…

Referring Video Object Segmentation

The 1st Solution for 4th PVUW MeViS Challenge: Unleashing the Potential of Large Multimodal Models for Referring Video Segmentation

2025-04-07 · Hao Fang, Runmin Cong, Xiankai Lu, Zhiyang Chen 외

Motion expression video segmentation is designed to segment objects in accordance with the input motion expressions. In contrast to the conventional Referring Video Object Segmentation (RVOS), it places emphasis on motio…

Inference OptimizationReferring Video Object SegmentationSegmentationSemantic Segmentation+3

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

2025-01-07 · Haobo Yuan, Xiangtai Li, Tao Zhang, Zilong Huang 외

This work presents Sa2VA, the first unified model for dense grounded understanding of both images and videos. Unlike existing multi-modal large language models, which are often limited to specific modalities and tasks, S…

2kLanguage ModelingLanguage ModellingObject+6

Multimodal Referring Segmentation: A Survey

2025-08-01 · Henghui Ding, Song Tang, Shuting He, Chang Liu 외 arxiv

Multimodal referring segmentation aims to segment target objects in visual scenes, such as images, videos, and 3D scenes, based on referring expressions in text or audio format. This task plays a crucial role in practica…

Referring Expression

1st Place Solution for 5th LSVOS Challenge: Referring Video Object Segmentation

2024-01-01 · Zhuoyan Luo, Yicheng Xiao, Yong liu, Yitong Wang 외

The recent transformer-based models have dominated the Referring Video Object Segmentation (RVOS) task due to the superior performance. Most prior works adopt unified DETR framework to generate segmentation masks in quer…

ObjectReferring Video Object SegmentationSegmentationSemantic Segmentation+2