paper-with-me

홈 › Papers

ReferDINO-Plus: 2nd Solution for 4th PVUW MeViS Challenge at CVPR 2025

2025-03-30 · Tianming Liang, Haichao Jiang, Wei-Shi Zheng, Jian-Fang Hu

Referring Video Object Segmentation (RVOS) aims to segment target objects throughout a video based on a text description. This task has attracted increasing attention in the field of computer vision due to its promising applications in video editing and human-agent interaction. Recently, ReferDINO has demonstrated promising performance in this task by adapting object-level vision-language knowledge from pretrained foundational image models. In this report, we further enhance its capabilities by incorporating the advantages of SAM2 in mask quality and object consistency. In addition, to effectively balance performance between single-object and multi-object scenarios, we introduce a conditional mask fusion strategy that adaptively fuses the masks from ReferDINO and SAM2. Our solution, termed ReferDINO-Plus, achieves 60.43 \(\mathcal{J}\&\mathcal{F}\) on MeViS test set, securing 2nd place in the MeViS PVUW challenge at CVPR 2025. The code is available at: https://github.com/iSEE-Laboratory/ReferDINO-Plus.

📄 PDF Abstract BibTeX arXiv:2503.23509

Code (1)

isee-laboratory/referdino-plus 공식 구현 pytorch

Tasks

ObjectReferring Video Object SegmentationSemantic SegmentationVideo EditingVideo Object SegmentationVideo Semantic Segmentation

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

The 1st Winner for 5th PVUW MeViS-Text Challenge: Strong MLLMs Meet SAM3 for Referring Video Object Segmentation

2026-04-01 · Xusheng He, Canyang Wu, Jinrong Zhang, Weili Guan 외 arxiv

This report presents our winning solution to the 5th PVUW MeViS-Text Challenge. The track studies referring video object segmentation under motion-centric language expressions, where the model must jointly understand app…

Referring Video Object Segmentation

1st Place Solution for MeViS Track in CVPR 2024 PVUW Workshop: Motion Expression guided Video Segmentation

2024-06-11 · Mingqi Gao, Jingnan Luo, Jinyu Yang, Jungong Han 외

Motion Expression guided Video Segmentation (MeViS), as an emerging task, poses many new challenges to the field of referring video object segmentation (RVOS). In this technical report, we investigated and validated the …

Referring Video Object SegmentationSegmentationSemantic SegmentationVideo Object Segmentation+2

The 1st Solution for 4th PVUW MeViS Challenge: Unleashing the Potential of Large Multimodal Models for Referring Video Segmentation

2025-04-07 · Hao Fang, Runmin Cong, Xiankai Lu, Zhiyang Chen 외

Motion expression video segmentation is designed to segment objects in accordance with the input motion expressions. In contrast to the conventional Referring Video Object Segmentation (RVOS), it places emphasis on motio…

Inference OptimizationReferring Video Object SegmentationSegmentationSemantic Segmentation+3

SaSaSaSa2VA: 2nd Place of the 5th PVUW MeViS-Text Track

2026-03-28 · Dengxian Gong, Quanzhu Niu, Shihao Chen, Yuanzheng Wu 외 arxiv

Referring video object segmentation (RVOS) commonly grounds targets in videos based on static textual cues. MeViS benchmark extends this by incorporating motion-centric expressions (referring & reasoning motion expressio…

Referring Video Object Segmentation

Report of the 5th PVUW Challenge: Towards More Diverse Modalities in Pixel-Level Understanding

2026-04-28 · Chang Liu, Henghui Ding, Nikhila Ravi, Yunchao Wei 외 arxiv

This report summarizes the objectives, datasets, and top-performing methodologies of the 2026 Pixel-level Video Understanding in the Wild (PVUW) Challenge, hosted at CVPR 2026, which evaluates state-of-the-art models und…

Object Segmentation