paper-with-me

홈 › Papers

Large VLM-based Stylized Sports Captioning

2025-08-25 · Sauptik Dhar, Nicholas Buoncristiani, Joe Anakata, Haoyu Zhang, Michelle Munson arxiv

The advent of large (visual) language models (LLM / LVLM) have led to a deluge of automated human-like systems in several domains including social media content generation, search and recommendation, healthcare prognosis, AI assistants for cognitive tasks etc. Although these systems have been successfully integrated in production; very little focus has been placed on sports, particularly accurate identification and natural language description of the game play. Most existing LLM/LVLMs can explain generic sports activities, but lack sufficient domain-centric sports' jargon to create natural (human-like) descriptions. This work highlights the limitations of existing SoTA LLM/LVLMs for generating production-grade sports captions from images in a desired stylized format, and proposes a two-level fine-tuned LVLM pipeline to address that. The proposed pipeline yields an improvement > 8-10% in the F1, and > 2-10% in BERT score compared to alternative approaches. In addition, it has a small runtime memory footprint and fast execution time. During Super Bowl LIX the pipeline proved its practical application for live professional sports journalism; generating highly accurate and stylized captions at the rate of 6 images per 3-5 seconds for over 1000 images during the game play.

📄 PDF Abstract BibTeX arXiv:2508.19295

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Similar Scenes arouse Similar Emotions: Parallel Data Augmentation for Stylized Image Captioning

2021-08-26 · Guodun Li, Yuchen Zhai, Zehao Lin, Yin Zhang

Stylized image captioning systems aim to generate a caption not only semantically related to a given image but also consistent with a given style description. One of the biggest challenges with this task is the lack of s…

Data AugmentationImage CaptioningSentence

"Factual" or "Emotional": Stylized Image Captioning with Adaptive Learning and Attention

2018-07-10 · Tianlang Chen, Zhongping Zhang, Quanzeng You, Chen Fang 외

Generating stylized captions for an image is an emerging topic in image captioning. Given an image as input, it requires the system to generate a caption that has a specific style (e.g., humorous, romantic, positive, and…

Image Captioning

``Factual'' or ``Emotional'': Stylized Image Captioning with Adaptive Learning and Attention

2018-09-01 · ECCV 2018 9 · Tianlang Chen, Zhongping Zhang, Quanzeng You, Chen Fang 외

Generating stylized captions for an image is an emerging topic in image captioning. Given an image as input, it requires the system to generate a caption that has a specific style (e.g., humorous, romantic, positive, and…

Image Captioning

UnMA-CapSumT: Unified and Multi-Head Attention-driven Caption Summarization Transformer

2024-12-16 · Dhruv Sharma, Chhavi Dhiman, Dinesh Kumar

Image captioning is the generation of natural language descriptions of images which have increased immense popularity in the recent past. With this different deep-learning techniques are devised for the development of fa…

Image Captioning

Visual Captioning at Will: Describing Images and Videos Guided by a Few Stylized Sentences

2023-07-31 · Dingyi Yang, Hongyu Chen, Xinglin Hou, Tiezheng Ge 외

Stylized visual captioning aims to generate image or video descriptions with specific styles, making them more attractive and emotionally appropriate. One major challenge with this task is the lack of paired stylized cap…

DecoderImage CaptioningLanguage Modelling