paper-with-me

Papers

Fine-grained Action Analysis: A Multi-modality and Multi-task Dataset of Figure Skating

2023-07-06 · Sheng-Lan Liu, Yu-Ning Ding, Gang Yan, Si-Fan Zhang, Jin-Rong Zhang, Wen-Yue Chen, Xue-Hai Xu

The fine-grained action analysis of the existing action datasets is challenged by insufficient action categories, low fine granularities, limited modalities, and tasks. In this paper, we propose a Multi-modality and Multi-task dataset of Figure Skating (MMFS) which was collected from the World Figure Skating Championships. MMFS, which possesses action recognition and action quality assessment, captures RGB, skeleton, and is collected the score of actions from 11671 clips with 256 categories including spatial and temporal labels. The key contributions of our dataset fall into three aspects as follows. (1) Independently spatial and temporal categories are first proposed to further explore fine-grained action recognition and quality assessment. (2) MMFS first introduces the skeleton modality for complex fine-grained action quality assessment. (3) Our multi-modality and multi-task dataset encourage more action analysis models. To benchmark our dataset, we adopt RGB-based and skeleton-based baseline methods for action recognition and action quality assessment.

📄 PDF Abstract BibTeX arXiv:2307.02730

Code (1)

dingyn-reno/mmfs 공식 구현

Tasks

Action AnalysisAction Quality AssessmentAction RecognitionFine-grained Action Recognition

Similar Papers 제목 키워드 기반

Graph-based Fine-grained Multimodal Attention Mechanism for Sentiment Analysis

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Multimodal sentiment analysis is a popular research area in natural language processing. Mainstream multimodal learning models barely consider that the visual and acoustic behaviors often have a much higher temporal freq…

multimodal interactionMultimodal Sentiment AnalysisSentiment Analysis

PeVL: Pose-Enhanced Vision-Language Model for Fine-Grained Human Action Recognition

2024-01-01 · CVPR 2024 1 · Haosong Zhang, Mei Chee Leong, Liyuan Li, Weisi Lin

Recent progress in Vision-Language (VL) foundation models has revealed the great advantages of cross-modality learning. However due to a large gap between vision and text they might not be able to sufficiently utiliz…

Action RecognitionContrastive LearningLanguage ModelingLanguage Modelling+1

Similarity Guided Multimodal Fusion Transformer for Semantic Location Prediction in Social Media

2024-05-09 · Zhizhen Zhang, Ning Wang, Haojie Li, Zhihui Wang

Semantic location prediction aims to derive meaningful location insights from multimodal social media posts, offering a more contextual understanding of daily activities than using GPS coordinates. This task faces signif…

Language Modelling

OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation

2026-05-12 · Guohui Zhang, XiaoXiao Ma, Jie Huang, Hang Xu 외 arxiv

Recent advances in joint audio-video generation have been remarkable, yet real-world applications demand strong per-modality fidelity, cross-modal alignment, and fine-grained synchronization. Reinforcement Learning (RL) …

Reinforcement LearningVideo Generation

New Benchmark Dataset and Fine-Grained Cross-Modal Fusion Framework for Vietnamese Multimodal Aspect-Category Sentiment Analysis

2024-05-01 · Quy Hoang Nguyen, Minh-Van Truong Nguyen, Kiet Van Nguyen

The emergence of multimodal data on social media platforms presents new opportunities to better understand user sentiments toward a given aspect. However, existing multimodal datasets for Aspect-Category Sentiment Analys…

Aspect Category Sentiment AnalysisMultimodal Sentiment AnalysisSentiment AnalysisVietnamese Datasets+3