Fine-grained Action Analysis: A Multi-modality and Multi-task Dataset of Figure Skating
The fine-grained action analysis of the existing action datasets is challenged by insufficient action categories, low fine granularities, limited modalities, and tasks. In this paper, we propose a Multi-modality and Multi-task dataset of Figure Skating (MMFS) which was collected from the World Figure Skating Championships. MMFS, which possesses action recognition and action quality assessment, captures RGB, skeleton, and is collected the score of actions from 11671 clips with 256 categories including spatial and temporal labels. The key contributions of our dataset fall into three aspects as follows. (1) Independently spatial and temporal categories are first proposed to further explore fine-grained action recognition and quality assessment. (2) MMFS first introduces the skeleton modality for complex fine-grained action quality assessment. (3) Our multi-modality and multi-task dataset encourage more action analysis models. To benchmark our dataset, we adopt RGB-based and skeleton-based baseline methods for action recognition and action quality assessment.
Code (1)
Tasks
Action AnalysisAction Quality AssessmentAction RecognitionFine-grained Action RecognitionSimilar Papers 제목 키워드 기반
Graph-based Fine-grained Multimodal Attention Mechanism for Sentiment Analysis
Multimodal sentiment analysis is a popular research area in natural language processing. Mainstream multimodal learning models barely consider that the visual and acoustic behaviors often have a much higher temporal freq…
multimodal interactionMultimodal Sentiment AnalysisSentiment AnalysisPeVL: Pose-Enhanced Vision-Language Model for Fine-Grained Human Action Recognition
Recent progress in Vision-Language (VL) foundation models has revealed the great advantages of cross-modality learning. However due to a large gap between vision and text they might not be able to sufficiently utiliz…
Action RecognitionContrastive LearningLanguage ModelingLanguage Modelling+1Similarity Guided Multimodal Fusion Transformer for Semantic Location Prediction in Social Media
Semantic location prediction aims to derive meaningful location insights from multimodal social media posts, offering a more contextual understanding of daily activities than using GPS coordinates. This task faces signif…
Language ModellingOmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation
Recent advances in joint audio-video generation have been remarkable, yet real-world applications demand strong per-modality fidelity, cross-modal alignment, and fine-grained synchronization. Reinforcement Learning (RL) …
Reinforcement LearningVideo GenerationNew Benchmark Dataset and Fine-Grained Cross-Modal Fusion Framework for Vietnamese Multimodal Aspect-Category Sentiment Analysis
The emergence of multimodal data on social media platforms presents new opportunities to better understand user sentiments toward a given aspect. However, existing multimodal datasets for Aspect-Category Sentiment Analys…
Aspect Category Sentiment AnalysisMultimodal Sentiment AnalysisSentiment AnalysisVietnamese Datasets+3