Unsupervised Temporal Feature Aggregation for Event Detection in Unstructured Sports Videos
Image-based sports analytics enable automatic retrieval of key events in a game to speed up the analytics process for human experts. However, most existing methods focus on structured television broadcast video datasets with a straight and fixed camera having minimum variability in the capturing pose. In this paper, we study the case of event detection in sports videos for unstructured environments with arbitrary camera angles. The transition from structured to unstructured video analysis produces multiple challenges that we address in our paper. Specifically, we identify and solve two major problems: unsupervised identification of players in an unstructured setting and generalization of the trained models to pose variations due to arbitrary shooting angles. For the first problem, we propose a temporal feature aggregation algorithm using person re-identification features to obtain high player retrieval precision by boosting a weak heuristic scoring method. Additionally, we propose a data augmentation technique, based on multi-modal image translation model, to reduce bias in the appearance of training samples. Experimental evaluations show that our proposed method improves precision for player retrieval from 0.78 to 0.86 for obliquely angled videos. Additionally, we obtain an improvement in F1 score for rally detection in table tennis videos from 0.79 in case of global frame-level features to 0.89 using our proposed player-level features. Please see the supplementary video submission at https://ibm.biz/BdzeZA.
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationEvent DetectionPerson Re-IdentificationRetrievalSports AnalyticsMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Rethinking Event-Based Object Detection through Representation-Level Temporal Aggregation and Model-Level Hypergraph Reasoning
Event cameras provide microsecond-level temporal resolution, low latency, and high dynamic range, offering potential for perception under fast motion and challenging illumination conditions. However, existing Event-based…
Relational ReasoningObject DetectionVISTA: Unsupervised 2D Temporal Dependency Representations for Time Series Anomaly Detection
Time Series Anomaly Detection (TSAD) is essential for uncovering rare and potentially harmful events in unlabeled time series data. Existing methods are highly dependent on clean, high-quality inputs, making them suscept…
Anomaly DetectionTime SeriesTime Series Anomaly DetectionVVitCutLER: Towards Unsupervised Object Detection and Segmentation in Videos
Unsupervised pixel-level video understanding remains challenging in real-world scenarios, where motion blur, occlusion, and fast object dynamics often cause temporal drift and flickering pseudo-labels.We propose VVitCutL…
Video Object DetectionInstance SegmentationMTGA: Multi-View Temporal Granularity Aligned Aggregation for Event-Based Lip-Reading
Lip-reading is to utilize the visual information of the speaker's lip movements to recognize words and sentences. Existing event-based lip-reading solutions integrate different frame rate branches to learn spatio-tempora…
Lip ReadingUnsupervised Event Clustering and Aggregation from Newswire and Web Articles
In this paper, we present an unsupervised pipeline approach for clustering news articles based on identified event instances in their content. We leverage press agency newswire and monolingual word alignment techniques t…
ArticlesClusteringDocument SummarizationWord Alignment