paper-with-me

Papers

Training-Free Spatio-temporal Decoupled Reasoning Video Segmentation with Adaptive Object Memory

2026-03-02 · Zhengtong Zhu, Jiaqing Fan, Zhixuan Liu, Fanzhang Li arxiv

Reasoning Video Object Segmentation (ReasonVOS) is a challenging task that requires stable object segmentation across video sequences using implicit and complex textual inputs. Previous methods fine-tune Multimodal Large Language Models (MLLMs) to produce segmentation outputs, which demand substantial resources. Additionally, some existing methods are coupled in the processing of spatio-temporal information, which affects the temporal stability of the model to some extent. To address these issues, we propose Training-Free \textbf{S}patio-temporal \textbf{D}ecoupled Reasoning Video Segmentation with \textbf{A}daptive Object \textbf{M}emory (SDAM). We aim to design a training-free reasoning video segmentation framework that outperforms existing methods requiring fine-tuning, using only pre-trained models. Meanwhile, we propose an Adaptive Object Memory module that selects and memorizes key objects based on motion cues in different video sequences. Finally, we propose Spatio-temporal Decoupling for stable temporal propagation. In the spatial domain, we achieve precise localization and segmentation of target objects, while in the temporal domain, we leverage key object temporal information to drive stable cross-frame propagation. Our method achieves excellent results on five benchmark datasets, including Ref-YouTubeVOS, Ref-DAVIS17, MeViS, ReasonVOS, and ReVOS.

📄 PDF Abstract BibTeX arXiv:2603.01545

Code (0)

등록된 구현이 없습니다.

Tasks

Video Object SegmentationVideo Segmentation

Similar Papers 제목 키워드 기반

STAC: Selective Spatiotemporal Aggregation and Compression for Video Reasoning Segmentation

2026-07-03 · Syed Ariff Syed Hesham, Yun Liu, Guolei Sun, Jing Yang 외 arxiv

Video reasoning segmentation demands pixel-accurate object tracking across hundreds of frames under complex natural language queries, producing dense spatiotemporal tokens whose quadratic self-attention cost makes long-v…

Natural Language QueriesObject Tracking

Spatial-Temporal-Decoupled Masked Pre-training for Spatiotemporal Forecasting

2023-12-01 · Haotian Gao, Renhe Jiang, Zheng Dong, Jinliang Deng 외

Spatiotemporal forecasting techniques are significant for various domains such as transportation, energy, and weather. Accurate prediction of spatiotemporal series remains challenging due to the complex spatiotemporal he…

Time SeriesTraffic Prediction

NCSTR: Node-Centric Decoupled Spatio-Temporal Reasoning for Video-based Human Pose Estimation

2026-03-20 · Quang Dang Huynh, Xuefei Yin, Andrew Busch, Hugo G. Espinosa 외 arxiv

Video-based human pose estimation remains challenged by motion blur, occlusion, and complex spatiotemporal dynamics. Existing methods often rely on heatmaps or implicit spatio-temporal feature aggregation, which limits j…

Pose Estimation

A Decoupled Spatio-Temporal Framework for Skeleton-based Action Segmentation

2023-12-10 · Yunheng Li, Zhongyu Li, ShangHua Gao, Qilong Wang 외

Effectively modeling discriminative spatio-temporal information is essential for segmenting activities in long action sequences. However, we observe that existing methods are limited in weak spatio-temporal modeling capa…

Action SegmentationSkeleton Based Action Segmentation

DeCo-VAE: Learning Compact Latents for Video Reconstruction via Decoupled Representation

2025-11-18 · Xiangchen Yin, Jiahui Yuan, Zhangchi Hu, Wenzhang Sun 외 arxiv

Existing video Variational Autoencoders (VAEs) generally overlook the similarity between frame contents, leading to redundant latent modeling. In this paper, we propose decoupled VAE (DeCo-VAE) to achieve compact latent …

Video Reconstruction