paper-with-me

Papers

Decoupling Static and Hierarchical Motion Perception for Referring Video Segmentation

2024-04-04 · CVPR 2024 1 · Shuting He, Henghui Ding

Referring video segmentation relies on natural language expressions to identify and segment objects, often emphasizing motion clues. Previous works treat a sentence as a whole and directly perform identification at the video-level, mixing up static image-level cues with temporal motion cues. However, image-level features cannot well comprehend motion cues in sentences, and static cues are not crucial for temporal perception. In fact, static cues can sometimes interfere with temporal perception by overshadowing motion cues. In this work, we propose to decouple video-level referring expression understanding into static and motion perception, with a specific emphasis on enhancing temporal comprehension. Firstly, we introduce an expression-decoupling module to make static cues and motion cues perform their distinct role, alleviating the issue of sentence embeddings overlooking motion cues. Secondly, we propose a hierarchical motion perception module to capture temporal information effectively across varying timescales. Furthermore, we employ contrastive learning to distinguish the motions of visually similar objects. These contributions yield state-of-the-art performance across five datasets, including a remarkable $\textbf{9.2%}$ $\mathcal{J\&F}$ improvement on the challenging $\textbf{MeViS}$ dataset. Code is available at https://github.com/heshuting555/DsHmp.

📄 PDF Abstract BibTeX arXiv:2404.03645

Code (1)

heshuting555/dshmp 공식 구현 pytorch

Tasks

Contrastive LearningReferring ExpressionReferring Expression SegmentationReferring Video Object SegmentationSentenceSentence EmbeddingsVideo SegmentationVideo Semantic Segmentation

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

MM-Conv: A Multimodal Dataset and Benchmark for Context-Aware Grounding in 3D Dialogue

2026-05-20 · Anna Deichler, Jim O'Regan, Fethiye Irmak Dogan, Lubos Marcinek 외 arxiv

Grounding language in the physical world requires AI systems to interpret references that emerge dynamically during conversation. While current vision-language models (VLMs) excel at static image tasks, they struggle to …

Visual Localization

Cognitive Disentanglement for Referring Multi-Object Tracking

2025-03-14 · Shaofeng Liang, Runwei Guan, Wangwang Lian, Daizong Liu 외

As a significant application of multi-source information fusion in intelligent transportation perception systems, Referring Multi-Object Tracking (RMOT) involves localizing and tracking specific objects in video sequence…

DisentanglementMulti-Object TrackingObjectObject Tracking+2

CoHD: A Counting-Aware Hierarchical Decoding Framework for Generalized Referring Expression Segmentation

2024-05-24 · Zhuoyan Luo, Yinghao Wu, Tianheng Cheng, Yong liu 외

The newly proposed Generalized Referring Expression Segmentation (GRES) amplifies the formulation of classic RES by involving complex multiple/non-target scenarios. Recent approaches address GRES by directly extending th…

Generalized Referring Expression SegmentationObjectReferring ExpressionReferring Expression Segmentation+1

Vision-Motion-Reference Alignment for Referring Multi-Object Tracking via Multi-Modal Large Language Models

2025-11-21 · Weiyi Lv, Ning Zhang, Hanyang Sun, Haoran Jiang 외 arxiv

Referring Multi-Object Tracking (RMOT) extends conventional multi-object tracking (MOT) by introducing natural language references for multi-modal fusion tracking. RMOT benchmarks only describe the object's appearance, r…

Multi-Object Tracking

SDD-4DGS: Static-Dynamic Aware Decoupling in Gaussian Splatting for 4D Scene Reconstruction

2025-03-12 · Dai Sun, Huhao Guan, Kun Zhang, Xike Xie 외

Dynamic and static components in scenes often exhibit distinct properties, yet most 4D reconstruction methods treat them indiscriminately, leading to suboptimal performance in both cases. This work introduces SDD-4DGS, t…

4D reconstruction