paper-with-me

홈 › Papers

Unleashing the Temporal-Spatial Reasoning Capacity of GPT for Training-Free Audio and Language Referenced Video Object Segmentation

2024-08-28 · Shaofei Huang, Rui Ling, Hongyu Li, Tianrui Hui, Zongheng Tang, Xiaoming Wei, Jizhong Han, Si Liu

In this paper, we propose an Audio-Language-Referenced SAM 2 (AL-Ref-SAM 2) pipeline to explore the training-free paradigm for audio and language-referenced video object segmentation, namely AVS and RVOS tasks. The intuitive solution leverages GroundingDINO to identify the target object from a single frame and SAM 2 to segment the identified object throughout the video, which is less robust to spatiotemporal variations due to a lack of video context exploration. Thus, in our AL-Ref-SAM 2 pipeline, we propose a novel GPT-assisted Pivot Selection (GPT-PS) module to instruct GPT-4 to perform two-step temporal-spatial reasoning for sequentially selecting pivot frames and pivot boxes, thereby providing SAM 2 with a high-quality initial object prompt. Within GPT-PS, two task-specific Chain-of-Thought prompts are designed to unleash GPT's temporal-spatial reasoning capacity by guiding GPT to make selections based on a comprehensive understanding of video and reference information. Furthermore, we propose a Language-Binded Reference Unification (LBRU) module to convert audio signals into language-formatted references, thereby unifying the formats of AVS and RVOS tasks in the same pipeline. Extensive experiments on both tasks show that our training-free AL-Ref-SAM 2 pipeline achieves performances comparable to or even better than fully-supervised fine-tuning methods. The code is available at: https://github.com/appletea233/AL-Ref-SAM2.

📄 PDF Abstract BibTeX arXiv:2408.15876

Code (1)

appletea233/al-ref-sam2 공식 구현 pytorch

Tasks

ObjectSemantic SegmentationSpatial ReasoningVideo Object SegmentationVideo Semantic Segmentation

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Adam 설명 없음
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Weight Decay 설명 없음
Attention 설명 없음
Position-Wise Feed-Forward Layer 설명 없음

Similar Papers 제목 키워드 기반

Unleashing Temporal Capacity of Spiking Neural Networks through Spatiotemporal Separation

2025-12-05 · Yiting Dong, Zhaofei Yu, Jianhao Ding, Zijie Xu 외 arxiv

Spiking Neural Networks (SNNs) are considered naturally suited for temporal processing, with membrane potential propagation widely regarded as the core temporal modeling mechanism. However, existing research lack analysi…

Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning

2026-03-24 · Jiacheng Hua, Yishu Yin, Yuhang Wu, Tai Wang 외 arxiv

Existing Multimodal Large Language Models (MLLMs) struggle with 3D spatial reasoning, as they fail to construct structured abstractions of the 3D environment depicted in video inputs. To bridge this gap, drawing inspirat…

Question AnsweringSpatial Reasoning

TrackVLA++: Unleashing Reasoning and Memory Capabilities in VLA Models for Embodied Visual Tracking

2025-10-08 · Jiahang Liu, Yunpeng Qi, Jiazhao Zhang, Minghan Li 외 arxiv

Embodied Visual Tracking (EVT) is a fundamental ability that underpins practical applications, such as companion robots, guidance robots and service assistants, where continuously following moving targets is essential. R…

Zero-shot GeneralizationSpatial ReasoningVisual Tracking

Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding

2026-03-19 · Xianjin Wu, Dingkang Liang, Tianrui Feng, Kui Xia 외 arxiv

While Multimodal Large Language Models demonstrate impressive semantic capabilities, they often suffer from spatial blindness, struggling with fine-grained geometric reasoning and physical dynamics. Existing solutions ty…

Scene UnderstandingSpatial ReasoningVideo Generation

Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video Grounding

2025-09-18 · Zaiquan Yang, Yuhao Liu, Gerhard Hancke, Rynson W. H. Lau arxiv

Spatio-temporal video grounding (STVG) aims at localizing the spatio-temporal tube of a video, as specified by the input text query. In this paper, we utilize multimodal large language models (MLLMs) to explore a zero-sh…

Spatio-Temporal Video Grounding