paper-with-me

홈 › Papers

Fusing Motion Patterns and Key Visual Information for Semantic Event Recognition in Basketball Videos

2020-07-13 · Lifang Wu, Zhou Yang, Qi. Wang, Meng Jian, Boxuan Zhao, Junchi Yan, Chang Wen Chen

Many semantic events in team sport activities e.g. basketball often involve both group activities and the outcome (score or not). Motion patterns can be an effective means to identify different activities. Global and local motions have their respective emphasis on different activities, which are difficult to capture from the optical flow due to the mixture of global and local motions. Hence it calls for a more effective way to separate the global and local motions. When it comes to the specific case for basketball game analysis, the successful score for each round can be reliably detected by the appearance variation around the basket. Based on the observations, we propose a scheme to fuse global and local motion patterns (MPs) and key visual information (KVI) for semantic event recognition in basketball videos. Firstly, an algorithm is proposed to estimate the global motions from the mixed motions based on the intrinsic property of camera adjustments. And the local motions could be obtained from the mixed and global motions. Secondly, a two-stream 3D CNN framework is utilized for group activity recognition over the separated global and local motion patterns. Thirdly, the basket is detected and its appearance features are extracted through a CNN structure. The features are utilized to predict the success or failure. Finally, the group activity recognition and success/failure prediction results are integrated using the kronecker product for event recognition. Experiments on NCAA dataset demonstrate that the proposed method obtains state-of-the-art performance.

📄 PDF Abstract BibTeX arXiv:2007.06288

Code (0)

등록된 구현이 없습니다.

Tasks

Activity RecognitionGroup Activity RecognitionOptical Flow Estimation

Similar Papers 제목 키워드 기반

Interpretable Multimodal Emotion Recognition using Facial Features and Physiological Signals

2023-06-05 · Puneet Kumar, Xiaobai Li

This paper aims to demonstrate the importance and feasibility of fusing multimodal information for emotion recognition. It introduces a multimodal framework for emotion understanding by fusing the information from visual…

Emotion ClassificationEmotion RecognitionFeature ImportanceMultimodal Emotion Recognition

Multimodal Emotion Recognition by Fusing Video Semantic in MOOC Learning Scenarios

2024-04-11 · Yuan Zhang, Xiaomei Tao, Hanxu Ai, Tao Chen 외

In the Massive Open Online Courses (MOOC) learning scenario, the semantic information of instructional videos has a crucial impact on learners' emotional state. Learners mainly acquire knowledge by watching instructional…

Emotion RecognitionLanguage ModellingLarge Language ModelMultimodal Emotion Recognition+1

SMILE: Infusing Spatial and Motion Semantics in Masked Video Learning

2025-04-01 · CVPR 2025 1 · Fida Mohammad Thoker, Letian Jiang, Chen Zhao, Bernard Ghanem

Masked video modeling, such as VideoMAE, is an effective paradigm for video self-supervised learning (SSL). However, they are primarily based on reconstructing pixel-level details on natural videos which have substantial…

Representation LearningSelf-Supervised Learning

Fusing Structure from Motion and Simulation-Augmented Pose Regression from Optical Flow for Challenging Indoor Environments

2023-04-14 · Felix Ott, Lucas Heublein, David Rügamer, Bernd Bischl 외

The localization of objects is a crucial task in various applications such as robotics, virtual and augmented reality, and the transportation of goods in warehouses. Recent advances in deep learning have enabled the loca…

Optical Flow EstimationPose Predictionregression

SLAck: Semantic, Location, and Appearance Aware Open-Vocabulary Tracking

2024-09-17 · Siyuan Li, Lei Ke, Yung-Hsu Yang, Luigi Piccinelli 외

Open-vocabulary Multiple Object Tracking (MOT) aims to generalize trackers to novel categories not in the training set. Currently, the best-performing methods are mainly based on pure appearance matching. Due to the comp…

Multiple Object TrackingObject Tracking