paper-with-me

홈 › Papers

BoxComm: Benchmarking Category-Aware Commentary Generation and Narration Rhythm in Boxing

2026-04-06 · Kaiwen Wang, Kaili Zheng, Rongrong Deng, Yiming Shi, Chenyi Guo, Ji Wu arxiv

Recent multimodal large language models (MLLMs) have shown strong capabilities in general video understanding, driving growing interest in automatic sports commentary generation. However, existing benchmarks for this task focus exclusively on team sports such as soccer and basketball, leaving combat sports entirely unexplored. Notably, combat sports present distinct challenges: critical actions unfold within milliseconds with visually subtle yet semantically decisive differences, and professional commentary contains a substantially higher proportion of tactical analysis compared to team sports. In this paper, we present BoxComm, a large-scale dataset comprising 445 World Boxing Championship match videos with over 52K commentary sentences from professional broadcasts. We propose a structured commentary taxonomy that categorizes each sentence into play-by-play, tactical, or contextual, providing the first category-level annotation for sports commentary benchmarks. Building on this taxonomy, we introduce two novel and complementary evaluations tailored to sports commentary generation: (1) category-conditioned generation, which evaluates whether models can produce accurate commentary of a specified type given video context; and (2) commentary rhythm assessment, which measures whether freely generated commentary exhibits appropriate temporal pacing and type distribution over continuous video segments, capturing a dimension of commentary competence that prior benchmarks have not addressed. Experiments on multiple state-of-the-art MLLMs reveal that current models struggle on both evaluations. We further propose EIC-Gen, an improved baseline incorporating detected punch events to supply structured action cues, yielding consistent gains and highlighting the importance of perceiving fleeting and subtle events for combat sports commentary.

📄 PDF Abstract BibTeX arXiv:2604.04419

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Real-Time Generation of Game Video Commentary with Multimodal LLMs: Pause-Aware Decoding Approaches

2026-03-03 · Anum Afzal, Yuki Saito, Hiroya Takamura, Katsuhito Sudoh 외 arxiv

Real-time video commentary generation provides textual descriptions of ongoing events in videos. It supports accessibility and engagement in domains such as sports, esports, and livestreaming. Commentary generation invol…

SCBench: A Sports Commentary Benchmark for Video LLMs

2024-12-23 · Kuangzhi Ge, Lingjun Chen, Kevin Zhang, Yulin Luo 외

Recently, significant advances have been made in Video Large Language Models (Video LLMs) in both academia and industry. However, methods to evaluate and benchmark the performance of different Video LLMs, especially thei…

Benchmarking

Bridging the Gap between Expert and Language Models: Concept-guided Chess Commentary Generation and Evaluation

2024-10-28 · Jaechang Kim, Jinmin Goh, Inseok Hwang, Jaewoong Cho 외

Deep learning-based expert models have reached superhuman performance in decision-making domains such as chess and Go. However, it is under-explored to explain or comment on given decisions although it is important for h…

Decision MakingInformativeness

Towards Automatic Soccer Commentary Generation with Knowledge-Enhanced Visual Reasoning

2026-03-31 · Zeyu Jin, Xiaoyu Qin, Songtao Zhou, Kaifeng Yun 외 arxiv

Soccer commentary plays a crucial role in enhancing the soccer game viewing experience for audiences. Previous studies in automatic soccer commentary generation typically adopt an end-to-end method to generate anonymous …

Visual Reasoning

MatchTime: Towards Automatic Soccer Game Commentary Generation

2024-06-26 · Jiayuan Rao, HaoNing Wu, Chang Liu, Yanfeng Wang 외

Soccer is a globally popular sport with a vast audience, in this paper, we consider constructing an automatic soccer game commentary model to improve the audiences' viewing experience. In general, we make the following c…