paper-with-me

홈 › Papers

Semantic-aware Video Representation for Few-shot Action Recognition

2023-11-10 · Yutao Tang, Benjamin Bejar, Rene Vidal

Recent work on action recognition leverages 3D features and textual information to achieve state-of-the-art performance. However, most of the current few-shot action recognition methods still rely on 2D frame-level representations, often require additional components to model temporal relations, and employ complex distance functions to achieve accurate alignment of these representations. In addition, existing methods struggle to effectively integrate textual semantics, some resorting to concatenation or addition of textual and visual features, and some using text merely as an additional supervision without truly achieving feature fusion and information transfer from different modalities. In this work, we propose a simple yet effective Semantic-Aware Few-Shot Action Recognition (SAFSAR) model to address these issues. We show that directly leveraging a 3D feature extractor combined with an effective feature-fusion scheme, and a simple cosine similarity for classification can yield better performance without the need of extra components for temporal modeling or complex distance functions. We introduce an innovative scheme to encode the textual semantics into the video representation which adaptively fuses features from text and video, and encourages the visual encoder to extract more semantically consistent features. In this scheme, SAFSAR achieves alignment and fusion in a compact way. Experiments on five challenging few-shot action recognition benchmarks under various settings demonstrate that the proposed SAFSAR model significantly improves the state-of-the-art performance.

📄 PDF Abstract BibTeX arXiv:2311.06218

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionFew-Shot action recognitionFew Shot Action Recognition

Similar Papers 제목 키워드 기반

LEViL: Label-Efficient Video Learning via Zero-Shot Distillation over VLM-Generated Pseudo-Label Spaces

2026-06-19 · Aslı Çelik arxiv

Supervised video pretraining is a common transfer learning practice for improving downstream action recognition performance. However, it requires large-scale labeled source datasets, and the effectiveness of the learned …

Action RecognitionTransfer Learning

MHSCNet: A Multimodal Hierarchical Shot-aware Convolutional Network for Video Summarization

2022-04-18 · Wujiang Xu, Runzhong Wang, Xiaobo Guo, Shaoshuai Li 외

Video summarization intends to produce a concise video summary by effectively capturing and combining the most informative parts of the whole content. Existing approaches for video summarization regard the task as a fram…

Video Summarization

TAEN: Temporal Aware Embedding Network for Few-Shot Action Recognition

2020-04-21 · Rami Ben-Ari, Mor Shpigel, Ophir Azulai, Udi Barzelay 외

Classification of new class entities requires collecting and annotating hundreds or thousands of samples that is often prohibitively costly. Few-shot learning suggests learning to classify new classes using just a few ex…

3D Face ReconstructionAction DetectionAction RecognitionClassification+5

GranAlign: Granularity-Aware Alignment Framework for Zero-Shot Video Moment Retrieval

2026-01-02 · Mingyu Jeon, Sunjae Yoon, Jonghee Kim, Junyeoung Kim arxiv

Zero-shot video moment retrieval (ZVMR) is the task of localizing a temporal moment within an untrimmed video using a natural language query without relying on task-specific training data. The primary challenge in this s…

Moment Retrieval

Object-Aware 4D Human Motion Generation

2025-10-31 · Shurui Gui, Deep Anil Patel, Xiner Li, Martin Renqiang Min arxiv

Recent advances in video diffusion models have enabled the generation of high-quality videos. However, these videos still suffer from unrealistic deformations, semantic violations, and physical inconsistencies that are l…