paper-with-me

홈 › Papers

SoccerLens: Grounded Soccer Video Understanding Beyond Accuracy

2026-05-10 · Ismael Elsharkawi, Ahmed Sait, Silvio Giancola, Bernard Ghanem, Hossam Sharara, Abdelrahman Eldesokey arxiv

Vision-language models (VLMs) have recently shown strong potential in soccer video understanding. However, given the high complexity of soccer videos due to large viewpoint variations, rapid shot transitions, and cluttered scenes, it remains unclear on whether VLMs rely on meaningful visual evidence or exploit spurious correlations and shortcut learning. Existing evaluation protocols focus primarily on classification accuracy and do not assess visual grounding. To address this limitation, we introduce SoccerLens, a benchmark for grounded soccer video understanding. The benchmark contains annotated video segments spanning $13$ common soccer events, with structured visual cues organized into three levels of semantic relevance. We further extend the attribution method of Chefer [arXiv:2103.15679] to jointly model spatial and temporal attention, and introduce evaluation metrics that measure whether model attention aligns with annotated cues or drifts toward spurious regions. Our evaluation of state-of-the-art soccer VLMs shows that, despite strong classification accuracy, current models fail to exceed $50\%$ grounding performance even under the loosest cue definitions and consistently underutilize temporal information. These results reveal a substantial gap between predictive performance and true visual grounding, highlighting the need for grounded evaluation in complex spatio-temporal domains such as soccer.

📄 PDF Abstract BibTeX arXiv:2605.09598

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Grounding

Similar Papers 제목 키워드 기반

SoccerNet-v2: A Dataset and Benchmarks for Holistic Understanding of Broadcast Soccer Videos

2020-11-26 · Adrien Deliège, Anthony Cioppa, Silvio Giancola, Meisam J. Seikavandi 외

Understanding broadcast videos is a challenging task in computer vision, as it requires generic reasoning capabilities to appreciate the content offered by the video editing. In this work, we propose SoccerNet-v2, a nove…

Action SpottingBoundary DetectionCamera shot boundary detectionCamera shot segmentation+3

SoccerDB: A Large-Scale Database for Comprehensive Video Understanding

2019-12-10 · Yudong Jiang, Kaixu Cui, Leilei Chen, Canjin Wang 외

Soccer videos can serve as a perfect research object for video understanding because soccer games are played under well-defined rules while complex and intriguing enough for researchers to study. In this paper, we propos…

Action ClassificationAction DetectionAction LocalizationAction Recognition+5

SoccerChat: Integrating Multimodal Data for Enhanced Soccer Game Understanding

2025-05-22 · Sushant Gautam, Cise Midoglu, Vajira Thambawita, Michael A. Riegler 외

The integration of artificial intelligence in sports analytics has transformed soccer video understanding, enabling real-time, automated insights into complex game dynamics. Traditional approaches rely on isolated data s…

Action ClassificationAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Decision Making+4

Towards Universal Soccer Video Understanding

2024-12-02 · CVPR 2025 1 · Jiayuan Rao, HaoNing Wu, Hao Jiang, Ya zhang 외

As a globally celebrated sport, soccer has attracted widespread interest from fans all over the world. This paper aims to develop a comprehensive multi-modal framework for soccer video understanding. Specifically, we mak…

Action ClassificationSports UnderstandingVideo Understanding

SoccerNet 2024 Challenges Results

2024-09-16 · Anthony Cioppa, Silvio Giancola, Vladimir Somers, Victor Joos 외

The SoccerNet 2024 challenges represent the fourth annual video understanding challenges organized by the SoccerNet team. These challenges aim to advance research across multiple themes in football, including broadcast v…

Action SpottingDense Video CaptioningGame State ReconstructionVideo Captioning+1