paper-with-me

홈 › Papers

Fine-Grained Action Recognition with Cross-Attentive Latent Sparse Experts

2026-08-13 · Imtiaz Ul Hassan, Tasweer Ahmad, Nik Bessis, Ardhendu Behera arxiv

Fine-grained human action recognition (FHAR) must distinguish visually similar actions that differ mainly in body configuration, timing, or local appearance. RGB representations retain visual context but often suppress joint-level geometry, whereas skeleton representations encode kinematics but discard dense spatial detail. We introduce FineX, which factorizes fine-grained cues into RGB appearance, pose heatmap geometry, and skeletal-graph topology. Pairwise cross-attention enables symmetric, stream-preserving information exchange, followed by a streamwise latent sparse Mixture-of-Experts that routes each representation to a content-dependent subset of shared experts, regularized by a load-balancing objective. FineX achieves state-of-the-art results on Gym99, Gym288, and Diving48. On the long-tailed Gym288, it raises mean class accuracy from 68.6% to 76.2% (+7.6 points) without textual supervision or large-scale vision-language pre-training, demonstrating the benefit of structured visual-pose-graph fusion and conditional expert refinement for FHAR.

📄 PDF Abstract BibTeX arXiv:2608.13458

Code (0)

등록된 구현이 없습니다.

Tasks

Action Recognition

Similar Papers 제목 키워드 기반

Few-shot Action Recognition with Prototype-centered Attentive Learning

2021-01-20 · Xiatian Zhu, Antoine Toisoul, Juan-Manuel Perez-Rua, Li Zhang 외

Few-shot action recognition aims to recognize action classes with few training samples. Most existing methods adopt a meta-learning approach with episodic training. In each episode, the few samples in a meta-training tas…

Action RecognitionContrastive LearningFew-Shot action recognitionFew Shot Action Recognition+2

Learning from Temporal Gradient for Semi-supervised Action Recognition

2021-11-25 · CVPR 2022 1 · Junfei Xiao, Longlong Jing, Lin Zhang, Ju He 외

Semi-supervised video action recognition tends to enable deep neural networks to achieve remarkable performance even with very limited labeled data. However, existing methods are mainly transferred from current image-bas…

Action RecognitionTemporal Action Localization

Learning Attentive Pairwise Interaction for Fine-Grained Classification

2020-02-24 · Peiqin Zhuang, Yali Wang, Yu Qiao

Fine-grained classification is a challenging problem, due to subtle differences among highly-confused categories. Most approaches address this difficulty by learning discriminative representation of individual input imag…

ClassificationFine-Grained Image ClassificationGeneral Classification

Speech Recognition Model Improves Text-to-Speech Synthesis using Fine-Grained Reward

2025-11-12 · Guansu Wang, Peijie Sun arxiv

Recent advances in text-to-speech (TTS) have enabled models to clone arbitrary unseen speakers and synthesize high-quality, natural-sounding speech. However, evaluation methods lag behind: typical mean opinion score (MOS…

Text-To-Speech SynthesisSpeech Recognition

Combined CNN Transformer Encoder for Enhanced Fine-grained Human Action Recognition

2022-08-03 · Mei Chee Leong, Haosong Zhang, Hui Li Tan, Liyuan Li 외

Fine-grained action recognition is a challenging task in computer vision. As fine-grained datasets have small inter-class variations in spatial and temporal space, fine-grained action recognition model requires good temp…

Action RecognitionAttributeFine-grained Action RecognitionTemporal Action Localization