paper-with-me

홈 › Papers

Multi-Modal Soccer Scene Analysis with Masked Pre-Training

2025-12-22 · Marc Peral, Guillem Capellera, Luis Ferraz, Antonio Rubio, Antonio Agudo arxiv

In this work we propose a multi-modal architecture for analyzing soccer scenes from tactical camera footage, with a focus on three core tasks: ball trajectory inference, ball state classification, and ball possessor identification. To this end, our solution integrates three distinct input modalities (player trajectories, player types and image crops of individual players) into a unified framework that processes spatial and temporal dynamics using a cascade of sociotemporal transformer blocks. Unlike prior methods, which rely heavily on accurate ball tracking or handcrafted heuristics, our approach infers the ball trajectory without direct access to its past or future positions, and robustly identifies the ball state and ball possessor under noisy or occluded conditions from real top league matches. We also introduce CropDrop, a modality-specific masking pre-training strategy that prevents over-reliance on image features and encourages the model to rely on cross-modal patterns during pre-training. We show the effectiveness of our approach on a large-scale dataset providing substantial improvements over state-of-the-art baselines in all tasks. Our results highlight the benefits of combining structured and visual cues in a transformer-based architecture, and the importance of realistic masking strategies in multi-modal learning.

📄 PDF Abstract BibTeX arXiv:2512.19528

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Survey of Action Recognition, Spotting and Spatio-Temporal Localization in Soccer -- Current Trends and Research Perspectives

2023-09-21 · Karolina Seweryn, Anna Wróblewska, Szymon Łukasik

Action scene understanding in soccer is a challenging task due to the complex and dynamic nature of the game, as well as the interactions between players. This article provides a comprehensive overview of this task divid…

Action LocalizationAction RecognitionScene UnderstandingSpatio-Temporal Action Localization+2

SoccerNet-v3D: Leveraging Sports Broadcast Replays for 3D Scene Understanding

2025-04-14 · Marc Gutiérrez-Pérez, Antonio Agudo

Sports video analysis is a key domain in computer vision, enabling detailed spatial understanding through multi-view correspondences. In this work, we introduce SoccerNet-v3D and ISSIA-3D, two enhanced and scalable datas…

Camera CalibrationObject LocalizationScene UnderstandingSports Analytics

SoccerChat: Integrating Multimodal Data for Enhanced Soccer Game Understanding

2025-05-22 · Sushant Gautam, Cise Midoglu, Vajira Thambawita, Michael A. Riegler 외

The integration of artificial intelligence in sports analytics has transformed soccer video understanding, enabling real-time, automated insights into complex game dynamics. Traditional approaches rely on isolated data s…

Action ClassificationAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Decision Making+4

Multi-Agent System for Comprehensive Soccer Understanding

2025-05-06 · Jiayuan Rao, Zifeng Li, HaoNing Wu, Ya zhang 외

Recent advancements in AI-driven soccer understanding have demonstrated rapid progress, yet existing research predominantly focuses on isolated or narrow tasks. To bridge this gap, we propose a comprehensive framework fo…

SoccerRAG: Multimodal Soccer Information Retrieval via Natural Queries

2024-06-03 · Aleksander Theo Strand, Sushant Gautam, Cise Midoglu, Pål Halvorsen

The rapid evolution of digital sports media necessitates sophisticated information retrieval systems that can efficiently parse extensive multimodal datasets. This paper introduces SoccerRAG, an innovative framework desi…

Information RetrievalNatural Language QueriesRAGRetrieval+2