paper-with-me

홈 › Papers

Content Based Video Narration of Gameplay with Vision Language Models

2026-08-14 · Mathew Varghese arxiv

Live game commentary is scarce: it exists for professional esports broadcasts and almost nowhere else. We present a content-based video narration system that produces spoken, esports-style commentary for arbitrary gameplay recordings using a general-purpose vision-language model (VLM) and a text-to-speech back end, with no game-specific instrumentation, no engine telemetry, and no task-specific training. Three mechanisms carry the system. Temporal mosaic packing arranges nine uniformly sampled frames into a single 3x3 image, letting an image-native VLM reason about motion while consuming one image payload per segment instead of nine. Context-conditioned prompting replays the K most recent narrations as assistant-role history, suppressing the repetition that dominates per-segment captioning of static scenes. Duration-conditioned generation and elastic alignment constrain narration length in the prompt, then time-scale or symmetrically pad the synthesized audio so each utterance fills its segment slot exactly, giving frame-accurate muxing without a forced aligner. The implementation supports either cloud TTS or a 6-bit quantized 4B-parameter on-device TTS model on Apple silicon, making the speech stage fully local. We report a qualitative case study on real-time strategy footage, a cost model showing the mosaic reduces per-minute image payloads by 9x, and a candid account of observed failure modes - hallucinated game state, resolution loss from mosaicking, and prosody artifacts from time-scaling. We release the system as a reproducible baseline, with an evaluation protocol for the quantitative study a full version will report.

📄 PDF Abstract BibTeX arXiv:2608.14016

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

TeaserGen: Generating Teasers for Long Documentaries

2024-10-08 · Weihan Xu, Paul Pu Liang, Haven Kim, Julian McAuley 외

Teasers are an effective tool for promoting content in entertainment, commercial and educational fields. However, creating an effective teaser for long videos is challenging for it requires long-range multimodal modeling…

Language ModellingLarge Language Model

Narration Generation for Cartoon Videos

2021-01-17 · Nikos Papasarantopoulos, Shay B. Cohen

Research on text generation from multimodal inputs has largely focused on static images, and less on video data. In this paper, we propose a new task, narration generation, that is complementing videos with narration tex…

Text Generation

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models

2025-07-22 · Tz-Ying Wu, Tahani Trigui, Sharath Nittur Sridhar, Anand Bodas 외 arxiv

In this paper, we introduce VideoNarrator, a novel training-free pipeline designed to generate dense video captions that offer a structured snapshot of video content. These captions offer detailed narrations with precise…

Video Question AnsweringVideo Summarization

CLIP meets GamePhysics: Towards bug identification in gameplay videos using zero-shot transfer learning

2022-03-21 · Mohammad Reza Taesiri, Finlay Macklon, Cor-Paul Bezemer

Gameplay videos contain rich information about how players interact with the game and how the game responds. Sharing gameplay videos on social media platforms, such as Reddit, has become a common practice for many player…

Event DetectionTransfer Learning

Movie101v2: Improved Movie Narration Benchmark

2024-04-20 · Zihao Yue, Yepeng Zhang, Ziheng Wang, Qin Jin

Automatic movie narration aims to generate video-aligned plot descriptions to assist visually impaired audiences. Unlike standard video captioning, it involves not only describing key visual details but also inferring pl…

Video Captioning