paper-with-me

홈 › Papers

Domain Adaptation of VLM for Soccer Video Understanding

2025-05-20 · Tiancheng Jiang, Henry Wang, Md Sirajus Salekin, Parmida Atighehchian, Shinan Zhang

Vision Language Models (VLMs) have demonstrated strong performance in multi-modal tasks by effectively aligning visual and textual representations. However, most video understanding VLM research has been domain-agnostic, leaving the understanding of their transfer learning capability to specialized domains under-explored. In this work, we address this by exploring the adaptability of open-source VLMs to specific domains, and focusing on soccer as an initial case study. Our approach uses large-scale soccer datasets and LLM to create instruction-following data, and use them to iteratively fine-tune the general-domain VLM in a curriculum learning fashion (first teaching the model key soccer concepts to then question answering tasks). The final adapted model, trained using a curated dataset of 20k video clips, exhibits significant improvement in soccer-specific tasks compared to the base model, with a 37.5% relative improvement for the visual question-answering task and an accuracy improvement from 11.8% to 63.5% for the downstream soccer action classification task.

📄 PDF Abstract BibTeX arXiv:2505.13860

Code (0)

등록된 구현이 없습니다.

Tasks

Action ClassificationDomain AdaptationInstruction FollowingQuestion AnsweringTransfer LearningVideo UnderstandingVisual Question Answering

Methods 이 논문이 사용한 방법론

BASE 설명 없음

Similar Papers 제목 키워드 기반

Towards Universal Soccer Video Understanding

2024-12-02 · CVPR 2025 1 · Jiayuan Rao, HaoNing Wu, Hao Jiang, Ya zhang 외

As a globally celebrated sport, soccer has attracted widespread interest from fans all over the world. This paper aims to develop a comprehensive multi-modal framework for soccer video understanding. Specifically, we mak…

Action ClassificationSports UnderstandingVideo Understanding

SoccerLens: Grounded Soccer Video Understanding Beyond Accuracy

2026-05-10 · Ismael Elsharkawi, Ahmed Sait, Silvio Giancola, Bernard Ghanem 외 arxiv

Vision-language models (VLMs) have recently shown strong potential in soccer video understanding. However, given the high complexity of soccer videos due to large viewpoint variations, rapid shot transitions, and clutter…

Visual Grounding

Multi-Agent System for Comprehensive Soccer Understanding

2025-05-06 · Jiayuan Rao, Zifeng Li, HaoNing Wu, Ya zhang 외

Recent advancements in AI-driven soccer understanding have demonstrated rapid progress, yet existing research predominantly focuses on isolated or narrow tasks. To bridge this gap, we propose a comprehensive framework fo…

SoccerNet-v2: A Dataset and Benchmarks for Holistic Understanding of Broadcast Soccer Videos

2020-11-26 · Adrien Deliège, Anthony Cioppa, Silvio Giancola, Meisam J. Seikavandi 외

Understanding broadcast videos is a challenging task in computer vision, as it requires generic reasoning capabilities to appreciate the content offered by the video editing. In this work, we propose SoccerNet-v2, a nove…

Action SpottingBoundary DetectionCamera shot boundary detectionCamera shot segmentation+3

SoccerDB: A Large-Scale Database for Comprehensive Video Understanding

2019-12-10 · Yudong Jiang, Kaixu Cui, Leilei Chen, Canjin Wang 외

Soccer videos can serve as a perfect research object for video understanding because soccer games are played under well-defined rules while complex and intriguing enough for researchers to study. In this paper, we propos…

Action ClassificationAction DetectionAction LocalizationAction Recognition+5