paper-with-me

홈 › Papers

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models

2025-07-22 · Tz-Ying Wu, Tahani Trigui, Sharath Nittur Sridhar, Anand Bodas, Subarna Tripathi arxiv

In this paper, we introduce VideoNarrator, a novel training-free pipeline designed to generate dense video captions that offer a structured snapshot of video content. These captions offer detailed narrations with precise timestamps, capturing the nuances present in each segment of the video. Despite advancements in multimodal large language models (MLLMs) for video comprehension, these models often struggle with temporally aligned narrations and tend to hallucinate, particularly in unfamiliar scenarios. VideoNarrator addresses these challenges by leveraging a flexible pipeline where off-the-shelf MLLMs and visual-language models (VLMs) can function as caption generators, context providers, or caption verifiers. Our experimental results demonstrate that the synergistic interaction of these components significantly enhances the quality and accuracy of video narrations, effectively reducing hallucinations and improving temporal alignment. This structured approach not only enhances video understanding but also facilitates downstream tasks such as video summarization and video question answering, and can be potentially extended for advertising and marketing applications.

📄 PDF Abstract BibTeX arXiv:2507.17050

Code (0)

등록된 구현이 없습니다.

Tasks

Video Question AnsweringVideo Summarization

Similar Papers 제목 키워드 기반

FlowNar: Scalable Streaming Narration for Long-Form Videos

2026-05-30 · Zeyun Zhong, Manuel Martin, Chengzhi Wu, David Schneider 외 arxiv

Recent Large Multimodal Models (LMMs), primarily designed for offline settings, are ill-suited for the dynamic requirements of streaming video. While recent online adaptations improve real-time processing, they still fac…

TeaserGen: Generating Teasers for Long Documentaries

2024-10-08 · Weihan Xu, Paul Pu Liang, Haven Kim, Julian McAuley 외

Teasers are an effective tool for promoting content in entertainment, commercial and educational fields. However, creating an effective teaser for long videos is challenging for it requires long-range multimodal modeling…

Language ModellingLarge Language Model

Narration Generation for Cartoon Videos

2021-01-17 · Nikos Papasarantopoulos, Shay B. Cohen

Research on text generation from multimodal inputs has largely focused on static images, and less on video data. In this paper, we propose a new task, narration generation, that is complementing videos with narration tex…

Text Generation

EgoAVU: Egocentric Audio-Visual Understanding

2026-02-05 · Ashish Seth, Xinhao Mei, Changsheng Zhao, Varun Nagaraja 외 arxiv

Understanding egocentric videos plays a vital role for embodied intelligence. Recent multi-modal large language models (MLLMs) can accept both visual and audio inputs. However, due to the challenge of obtaining text labe…

DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation

2026-04-29 · Mingji Ge, Qirui Chen, Zeqian Li, Weidi Xie arxiv

Long-term video understanding requires interpreting complex temporal events and reasoning over procedural activities. While instructional video corpora, like HowTo100M, offer rich resources for model training, they prese…

Zero-shot GeneralizationDense Video CaptioningCross-Modal Retrieval