paper-with-me

Papers

A Survey on Video Temporal Grounding with Multimodal Large Language Model

2025-08-07 · Jianlong Wu, Wei Liu, Ye Liu, Meng Liu, Liqiang Nie, Zhouchen Lin, Chang Wen Chen arxiv

The recent advancement in video temporal grounding (VTG) has significantly enhanced fine-grained video understanding, primarily driven by multimodal large language models (MLLMs). With superior multimodal comprehension and reasoning abilities, VTG approaches based on MLLMs (VTG-MLLMs) are gradually surpassing traditional fine-tuned methods. They not only achieve competitive performance but also excel in generalization across zero-shot, multi-task, and multi-domain settings. Despite extensive surveys on general video-language understanding, comprehensive reviews specifically addressing VTG-MLLMs remain scarce. To fill this gap, this survey systematically examines current research on VTG-MLLMs through a three-dimensional taxonomy: 1) the functional roles of MLLMs, highlighting their architectural significance; 2) training paradigms, analyzing strategies for temporal reasoning and task adaptation; and 3) video feature processing techniques, which determine spatiotemporal representation effectiveness. We further discuss benchmark datasets, evaluation protocols, and summarize empirical findings. Finally, we identify existing limitations and propose promising research directions. For additional resources and details, readers are encouraged to visit our repository at https://github.com/ki-lw/Awesome-MLLMs-for-Video-Temporal-Grounding.

📄 PDF Abstract BibTeX arXiv:2508.10922

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Temporal Sentence Grounding in Videos: A Survey and Future Directions

2022-01-20 · Hao Zhang, Aixin Sun, Wei Jing, Joey Tianyi Zhou

Temporal sentence grounding in videos (TSGV), \aka natural language video localization (NLVL) or video moment retrieval (VMR), aims to retrieve a temporal moment that semantically corresponds to a language query from an …

Moment RetrievalRetrievalSentenceTemporal Sentence Grounding

VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos

2024-11-07 · Shehan Munasinghe, Hanan Gani, Wenqi Zhu, Jiale Cao 외

Fine-grained alignment between videos and text is challenging due to complex spatial and temporal dynamics in videos. Existing video-based Large Multimodal Models (LMMs) handle basic conversations but struggle with preci…

DecoderLanguage ModelingLanguage ModellingLarge Language Model+3

VideoGLaMM : A Large Multimodal Model for Pixel-Level Visual Grounding in Videos

2025-01-01 · CVPR 2025 1 · Shehan Munasinghe, Hanan Gani, Wenqi Zhu, Jiale Cao 외

Fine-grained alignment between videos and text is challenging due to complex spatial and temporal dynamics in videos. Existing video-based Large Multimodal Models (LMMs) handle basic conversations but struggle with p…

Large Language ModelVideo SegmentationVideo Semantic SegmentationVisual Grounding

Video-LMM Post-Training: A Deep Dive into Video Reasoning with Large Multimodal Models

2025-10-06 · Yolo Y. Tang, Jing Bi, Pinxin Liu, Zhenyu Pan 외 arxiv

Video understanding represents the most challenging frontier in computer vision, requiring models to reason about complex spatiotemporal relationships, long-term dependencies, and multimodal evidence. The recent emergenc…

Reinforcement Learning

Efficient Temporal Extrapolation of Multimodal Large Language Models with Temporal Grounding Bridge

2024-02-25 · Yuxuan Wang, Yueqian Wang, Pengfei Wu, Jianxin Liang 외

Despite progress in multimodal large language models (MLLMs), the challenge of interpreting long-form videos in response to linguistic queries persists, largely due to the inefficiency in temporal grounding and limited p…

Computational EfficiencyLanguage ModellingOptical Flow EstimationQuestion Answering+1