paper-with-me

홈 › Papers

Pinpointing Trigger Moment for Grounded Video QA: Enhancing Spatio-temporal Grounding in Multimodal Large Language Models

2025-11-04 · Jinhwan Seo, Yoonki Cho, Junhyug Noh, Sung-eui Yoon arxiv

In this technical report, we introduce a framework to address Grounded Video Question Answering (GVQA) task for the ICCV 2025 Perception Test Challenge. The GVQA task demands robust multimodal models capable of complex reasoning over video content, grounding the resulting answers visually, and tracking the referenced objects temporally. To achieve this capability, our proposed approach decomposes the GVQA task into a three-stage pipeline: (1) Video Reasoning \& QA, (2) Spatio-temporal Grounding and (3) Tracking. Our key contribution is the introduction of a trigger moment, derived from our proposed CORTEX prompt, which pinpoints the single most visible frame of a target object to serve as a robust anchor for grounding and tracking. To this end, we achieve the HOTA score of 0.4968, which marks a significant improvement over the previous year's winning score of 0.2704 on GVQA task.

📄 PDF Abstract BibTeX arXiv:2511.02182

Code (0)

등록된 구현이 없습니다.

Tasks

Video Question Answering

Similar Papers 제목 키워드 기반

TA-Prompting: Enhancing Video Large Language Models for Dense Video Captioning via Temporal Anchors

2026-01-06 · Wei-Yuan Cheng, Kai-Po Chang, Chi-Pin Huang, Fu-En Yang 외 arxiv

Dense video captioning aims to interpret and describe all temporally localized events throughout an input video. Recent state-of-the-art methods leverage large language models (LLMs) to provide detailed moment descriptio…

Dense Video CaptioningMoment Retrieval

Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning

2024-02-18 · Long Qian, Juncheng Li, Yu Wu, Yaobo Ye 외

Large Language Models (LLMs) demonstrate remarkable proficiency in comprehending and handling text-based tasks. Many efforts are being made to transfer these attributes to video modality, which are termed Video-LLMs. How…

Language ModelingLanguage ModellingLarge Language Model

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection

2025-04-20 · Weijun Zhuang, Qizhang Li, Xin Li, Ming Liu 외

Temporal Action Detection and Moment Retrieval constitute two pivotal tasks in video understanding, focusing on precisely localizing temporal segments corresponding to specific actions or events. Recent advancements intr…

Action DetectionDecoderMoment RetrievalNatural Language Queries+3

D&M: Enriching E-commerce Videos with Sound Effects by Key Moment Detection and SFX Matching

2024-08-23 · Jingyu Liu, Minquan Wang, Ye Ma, Bo wang 외

Videos showcasing specific products are increasingly important for E-commerce. Key moments naturally exist as the first appearance of a specific product, presentation of its distinctive features, the presence of a buying…

Highlight DetectionMoment Retrieval

Background-aware Moment Detection for Video Moment Retrieval

2023-06-05 · Minjoon Jung, Youwon Jang, SeongHo Choi, Joochan Kim 외

Video moment retrieval (VMR) identifies a specific moment in an untrimmed video for a given natural language query. This task is prone to suffer the weak alignment problem innate in video datasets. Due to the ambiguity, …

Moment RetrievalNatural Language Moment RetrievalRetrieval