paper-with-me

Papers

TimeSoccer: An End-to-End Multimodal Large Language Model for Soccer Commentary Generation

2025-04-24 · Ling You, Wenxuan Huang, Xinni Xie, Xiangyi Wei, Bangyan Li, Shaohui Lin, Yang Li, Changbo Wang

Soccer is a globally popular sporting event, typically characterized by long matches and distinctive highlight moments. Recent advances in Multimodal Large Language Models (MLLMs) offer promising capabilities in temporal grounding and video understanding, soccer commentary generation often requires precise temporal localization and semantically rich descriptions over long-form video. However, existing soccer MLLMs often rely on the temporal a priori for caption generation, so they cannot process the soccer video end-to-end. While some traditional approaches follow a two-step paradigm that is complex and fails to capture the global context to achieve suboptimal performance. To solve the above issues, we present TimeSoccer, the first end-to-end soccer MLLM for Single-anchor Dense Video Captioning (SDVC) in full-match soccer videos. TimeSoccer jointly predicts timestamps and generates captions in a single pass, enabling global context modeling across 45-minute matches. To support long video understanding of soccer matches, we introduce MoFA-Select, a training-free, motion-aware frame compression module that adaptively selects representative frames via a coarse-to-fine strategy, and incorporates complementary training paradigms to strengthen the model's ability to handle long temporal sequences. Extensive experiments demonstrate that our TimeSoccer achieves State-of-The-Art (SoTA) performance on the SDVC task in an end-to-end form, generating high-quality commentary with accurate temporal alignment and strong semantic relevance.

📄 PDF Abstract BibTeX arXiv:2504.17365

Code (0)

등록된 구현이 없습니다.

Tasks

Caption GenerationDense Video CaptioningLanguage ModelingLanguage ModellingLarge Language ModelMultimodal Large Language ModelTemporal LocalizationTemporal SequencesVideo CaptioningVideo Understanding

Similar Papers 제목 키워드 기반

StreamSoccer: Event-Driven Memory for Streaming Soccer Commentary

2026-08-20 · Chenxi Shao, Bozhong Wang, Jiaxin Huang, Zhao Liu 외 arxiv

Streaming video understanding requires models to causally update state as video arrives and organize growing history into semantic units that can evolve, persist, and be recalled under bounded computation and memory. Thi…

MCAD: Multimodal Context-Aware Audio Description Generation For Soccer

2025-11-12 · Lipisha Chaudhary, Trisha Mittal, Subhadra Gopalakrishnan, Ifeoma Nwogu 외 arxiv

Audio Descriptions (AD) are essential for making visual content accessible to individuals with visual impairments. Recent works have shown a promising step towards automating AD, but they have been limited to describing …

Commentary Generation for Soccer Highlights

2025-08-11 · Chidaksh Ravuru arxiv

Automated soccer commentary generation has evolved from template-based systems to advanced neural architectures, aiming to produce real-time descriptions of sports events. While frameworks like SoccerNet-Caption laid fou…

MatchTime: Towards Automatic Soccer Game Commentary Generation

2024-06-26 · Jiayuan Rao, HaoNing Wu, Chang Liu, Yanfeng Wang 외

Soccer is a globally popular sport with a vast audience, in this paper, we consider constructing an automatic soccer game commentary model to improve the audiences' viewing experience. In general, we make the following c…

Towards Automatic Soccer Commentary Generation with Knowledge-Enhanced Visual Reasoning

2026-03-31 · Zeyu Jin, Xiaoyu Qin, Songtao Zhou, Kaifeng Yun 외 arxiv

Soccer commentary plays a crucial role in enhancing the soccer game viewing experience for audiences. Previous studies in automatic soccer commentary generation typically adopt an end-to-end method to generate anonymous …

Visual Reasoning