paper-with-me

홈 › Papers

On the Consistency of Video Large Language Models in Temporal Comprehension

2024-11-20 · CVPR 2025 1 · Minjoon Jung, Junbin Xiao, Byoung-Tak Zhang, Angela Yao

Video large language models (Video-LLMs) can temporally ground language queries and retrieve video moments. Yet, such temporal comprehension capabilities are neither well-studied nor understood. So we conduct a study on prediction consistency -- a key indicator for robustness and trustworthiness of temporal grounding. After the model identifies an initial moment within the video content, we apply a series of probes to check if the model's responses align with this initial grounding as an indicator of reliable comprehension. Our results reveal that current Video-LLMs are sensitive to variations in video contents, language queries, and task settings, unveiling severe deficiencies in maintaining consistency. We further explore common prompting and instruction-tuning methods as potential solutions, but find that their improvements are often unstable. To that end, we propose event temporal verification tuning that explicitly accounts for consistency, and demonstrate significant improvements for both grounding and consistency. Our data and code will be available at https://github.com/minjoong507/Consistency-of-Video-LLM.

📄 PDF Abstract BibTeX arXiv:2411.12951

Code (1)

minjoong507/consistency-of-video-llm 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation

2025-12-28 · Kai Liu, Jungang Li, Yuchong Sun, Shengqiong Wu 외 arxiv

This paper presents JavisGPT, the first unified multimodal large language model (MLLM) for joint audio-video (JAV) comprehension and generation. JavisGPT has a concise encoder-LLM-decoder architecture, which has a SyncFu…

Omni-RGPT: Unifying Image and Video Region-level Understanding via Token Marks

2025-01-14 · CVPR 2025 1 · Miran Heo, Min-Hung Chen, De-An Huang, Sifei Liu 외

We present Omni-RGPT, a multimodal large language model designed to facilitate region-level comprehension for both images and videos. To achieve consistent region representation across spatio-temporal dimensions, we intr…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+3

Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning

2024-02-18 · Long Qian, Juncheng Li, Yu Wu, Yaobo Ye 외

Large Language Models (LLMs) demonstrate remarkable proficiency in comprehending and handling text-based tasks. Many efforts are being made to transfer these attributes to video modality, which are termed Video-LLMs. How…

Language ModelingLanguage ModellingLarge Language Model

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding

2025-10-30 · Minjoon Jung, Junbin Xiao, Junghyun Kim, Byoung-Tak Zhang 외 arxiv

Do Video-LLMs have consistent temporal understanding when videos capture the same event from different viewpoints? To study this question, we introduce EgoExo-Con(sistency), a benchmark of synchronized egocentric and exo…

Reinforcement Learning

Investigating and Enhancing the Robustness of Large Multimodal Models Against Temporal Inconsistency

2025-05-20 · Jiafeng Liang, Shixin Jiang, Xuan Dong, Ning Wang 외

Large Multimodal Models (LMMs) have recently demonstrated impressive performance on general video comprehension benchmarks. Nevertheless, for broader applications, the robustness of their temporal analysis capability nee…