paper-with-me

홈 › Papers

Harnessing Object Grounding for Time-Sensitive Video Understanding

2025-09-08 · Tz-Ying Wu, Sharath Nittur Sridhar, Subarna Tripathi arxiv

We propose to improve the time-sensitive video understanding (TSV) capability of video large language models (Video-LLMs) with grounded objects (GO). We hypothesize that TSV tasks can benefit from GO within frames, which is supported by our preliminary experiments on LITA, a state-of-the-art Video-LLM for reasoning temporal localization. While augmenting prompts with textual descriptions of these object annotations improves the performance of LITA, it also introduces extra token length and susceptibility to the noise in object-level information. To address this, we propose GO-Tokenizer, a lightweight add-on module for Video-LLMs leveraging off-the-shelf object detectors to encode compact object information on the fly. Experimental results demonstrate that pretraining with GO-Tokenizer outperforms the vanilla Video-LLM and its counterpart, utilizing textual descriptions of objects in the prompt. The gain generalizes across different models, datasets, and video understanding tasks, such as reasoning temporal localization and dense captioning.

📄 PDF Abstract BibTeX arXiv:2509.06335

Code (0)

등록된 구현이 없습니다.

Tasks

Dense Captioning

Similar Papers 제목 키워드 기반

ArrowGEV: Grounding Events in Video via Learning the Arrow of Time

2026-01-10 · Fangxu Yu, Ziyao Lu, Liqiang Niu, Fandong Meng 외 arxiv

Grounding events in videos serves as a fundamental capability in video analysis. While Vision Language Models (VLMs) are increasingly employed for this task, existing approaches predominantly train models to associate ev…

Reinforcement Learning

ActPrompt: In-Domain Feature Adaptation via Action Cues for Video Temporal Grounding

2024-08-13 · Yubin Wang, Xinyang Jiang, De Cheng, Dongsheng Li 외

Video temporal grounding is an emerging topic aiming to identify specific clips within videos. In addition to pre-trained video models, contemporary methods utilize pre-trained vision-language models (VLM) to capture det…

Prompt Learning

VideoExpert: Augmented LLM for Temporal-Sensitive Video Understanding

2025-04-10 · Henghao Zhao, Ge-Peng Ji, Rui Yan, Huan Xiong 외

The core challenge in video understanding lies in perceiving dynamic content changes over time. However, multimodal large language models struggle with temporal-sensitive video tasks, which requires generating timestamps…

Instruction FollowingVideo Understanding

Learning Sample Importance for Cross-Scenario Video Temporal Grounding

2022-01-08 · Peijun Bao, Yadong Mu

The task of temporal grounding aims to locate video moment in an untrimmed video, with a given sentence query. This paper for the first time investigates some superficial biases that are specific to the temporal groundin…

Sentence

Spatio-Temporal Graph for Video Captioning with Knowledge Distillation

2020-03-31 · CVPR 2020 6 · Boxiao Pan, Haoye Cai, De-An Huang, Kuan-Hui Lee 외

Video captioning is a challenging task that requires a deep understanding of visual scenes. State-of-the-art methods generate captions using either scene-level or object-level information but without explicitly modeling …

Knowledge DistillationObjectVideo CaptioningVisual Grounding