paper-with-me

홈 › Papers

SV3.3B: A Sports Video Understanding Model for Action Recognition

2025-07-23 · Sai Varun Kodathala, Yashwanth Reddy Vutukoori, Rakesh Vunnam arxiv

This paper addresses the challenge of automated sports video analysis, which has traditionally been limited by computationally intensive models requiring server-side processing and lacking fine-grained understanding of athletic movements. Current approaches struggle to capture the nuanced biomechanical transitions essential for meaningful sports analysis, often missing critical phases like preparation, execution, and follow-through that occur within seconds. To address these limitations, we introduce SV3.3B, a lightweight 3.3B parameter video understanding model that combines novel temporal motion difference sampling with self-supervised learning for efficient on-device deployment. Our approach employs a DWT-VGG16-LDA based keyframe extraction mechanism that intelligently identifies the 16 most representative frames from sports sequences, followed by a V-DWT-JEPA2 encoder pretrained through mask-denoising objectives and an LLM decoder fine-tuned for sports action description generation. Evaluated on a subset of the NSVA basketball dataset, SV3.3B achieves superior performance across both traditional text generation metrics and sports-specific evaluation criteria, outperforming larger closed-source models including GPT-4o variants while maintaining significantly lower computational requirements. Our model demonstrates exceptional capability in generating technically detailed and analytically rich sports descriptions, achieving 29.2% improvement over GPT-4o in ground truth validation metrics, with substantial improvements in information density, action complexity, and measurement precision metrics essential for comprehensive athletic analysis. Model Available at https://huggingface.co/sportsvision/SV3.3B.

📄 PDF Abstract BibTeX arXiv:2507.17844

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised LearningAction RecognitionText Generation

Similar Papers 제목 키워드 기반

Distantly Supervised Semantic Text Detection and Recognition for Broadcast Sports Videos Understanding

2021-10-31 · Avijit Shah, Topojoy Biswas, Sathish Ramadoss, Deven Santosh Shah

Comprehensive understanding of key players and actions in multiplayer sports broadcast videos is a challenging problem. Unlike in news or finance videos, sports videos have limited text. While both action recognition for…

Action RecognitionText DetectionVideo Understanding

Video Pose Distillation for Few-Shot, Fine-Grained Sports Action Recognition

2021-09-03 · ICCV 2021 10 · James Hong, Matthew Fisher, Michaël Gharbi, Kayvon Fatahalian

Human pose is a useful feature for fine-grained sports action understanding. However, pose estimators are often unreliable when run on sports video due to domain shift and factors such as motion blur and occlusions. This…

Action RecognitionAction UnderstandingFine-grained Action RecognitionPose Estimation+1

A Survey on Video Action Recognition in Sports: Datasets, Methods and Applications

2022-06-02 · Fei Wu, Qingzhong Wang, Jian Bian, Haoyi Xiong 외

To understand human behaviors, action recognition based on videos is a common approach. Compared with image-based action recognition, videos provide much more information. Reducing the ambiguity of actions and in the las…

Action RecognitionSports AnalyticsTemporal Action Localization

ViSTec: Video Modeling for Sports Technique Recognition and Tactical Analysis

2024-02-25 · Yuchen He, Zeqing Yuan, Yihong Wu, Liqi Cheng 외

The immense popularity of racket sports has fueled substantial demand in tactical analysis with broadcast videos. However, existing manual methods require laborious annotation, and recent attempts leveraging video percep…

Action SegmentationInductive Bias

PL-NBA: A Possession-level Universal Basketball Video Dataset Supporting Multiple Visual Understanding Tasks

2026-08-20 · Yunhao Zhao, Haoying Sun, Jiarui Li, Zhuming Wang 외 arxiv

Visual understanding in sports has emerged as a hot topic in computer vision in recent years. Most existing basketball video datasets adopt single action or activity as sample, which can neither preserve the temporal con…

Temporal Action LocalizationAction AnticipationVideo Captioning