paper-with-me

홈 › Papers

Cross-Block Fine-Grained Semantic Cascade for Skeleton-Based Sports Action Recognition

2024-04-30 · Zhendong Liu, Haifeng Xia, Tong Guo, Libo Sun, Ming Shao, Siyu Xia

Human action video recognition has recently attracted more attention in applications such as video security and sports posture correction. Popular solutions, including graph convolutional networks (GCNs) that model the human skeleton as a spatiotemporal graph, have proven very effective. GCNs-based methods with stacked blocks usually utilize top-layer semantics for classification/annotation purposes. Although the global features learned through the procedure are suitable for the general classification, they have difficulty capturing fine-grained action change across adjacent frames -- decisive factors in sports actions. In this paper, we propose a novel ``Cross-block Fine-grained Semantic Cascade (CFSC)'' module to overcome this challenge. In summary, the proposed CFSC progressively integrates shallow visual knowledge into high-level blocks to allow networks to focus on action details. In particular, the CFSC module utilizes the GCN feature maps produced at different levels, as well as aggregated features from proceeding levels to consolidate fine-grained features. In addition, a dedicated temporal convolution is applied at each level to learn short-term temporal features, which will be carried over from shallow to deep layers to maximize the leverage of low-level details. This cross-block feature aggregation methodology, capable of mitigating the loss of fine-grained information, has resulted in improved performance. Last, FD-7, a new action recognition dataset for fencing sports, was collected and will be made publicly available. Experimental results and empirical analysis on public benchmarks (FSD-10) and self-collected (FD-7) demonstrate the advantage of our CFSC module on learning discriminative patterns for action classification over others.

📄 PDF Abstract BibTeX arXiv:2404.19383

Code (0)

등록된 구현이 없습니다.

Tasks

Action ClassificationAction RecognitionGeneral ClassificationVideo Recognition

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Focus 설명 없음
GCN A Graph Convolutional Network, or GCN, is an approach for semi-supervised learning on graph-structured data. It is based on an efficient variant of [convolutional neural…

Similar Papers 제목 키워드 기반

Enhancing Fine-Grained Image Classifications via Cascaded Vision Language Models

2024-05-18 · Canshi Wei

Fine-grained image classification, particularly in zero/few-shot scenarios, presents a significant challenge for vision-language models (VLMs), such as CLIP. These models often struggle with the nuanced task of distingui…

Fine-Grained Image Classificationimage-classificationImage Classification

Fine-grained Cross-modal Fusion based Refinement for Text-to-Image Synthesis

2023-02-17 · Haoran Sun, Yang Wang, Haipeng Liu, Biao Qian

Text-to-image synthesis refers to generating visual-realistic and semantically consistent images from given textual descriptions. Previous approaches generate an initial low-resolution image and then refine it to be high…

Image Generation

Motion-Adaptive Temporal Attention for Lightweight Video Generation with Stable Diffusion

2026-03-18 · Rui Hong, Shuxue Quan arxiv

We present a motion-adaptive temporal attention mechanism for parameter-efficient video generation built upon frozen Stable Diffusion models. Rather than treating all video content uniformly, our method dynamically adjus…

Video Generation

VQA with Cascade of Self- and Co-Attention Blocks

2023-02-28 · Aakansha Mishra, Ashish Anand, Prithwijit Guha

The use of complex attention modules has improved the performance of the Visual Question Answering (VQA) task. This work aims to learn an improved multi-modal representation through dense interaction of visual and textua…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Cascading Hierarchical Networks with Multi-task Balanced Loss for Fine-grained hashing

2023-03-20 · Xianxian Zeng, Yanjun Zheng

With the explosive growth in the number of fine-grained images in the Internet era, it has become a challenging problem to perform fast and efficient retrieval from large-scale fine-grained images. Among the many retriev…

Data AugmentationFine-Grained Image ClassificationMulti-Task LearningRetrieval