paper-with-me

Papers

Motion-Modulated Temporal Fragment Alignment Network for Few-Shot Action Recognition

2022-01-01 · CVPR 2022 1 · Jiamin Wu, Tianzhu Zhang, Zhe Zhang, Feng Wu, Yongdong Zhang

While the majority of FSL models focus on image classification, the extension to action recognition is rather challenging due to the additional temporal dimension in videos. To address this issue, we propose an end-to-end Motion-modulated Temporal Fragment Alignment Network (MTFAN) by jointly exploring the task-specific motion modulation and the multi-level temporal fragment alignment for Few-Shot Action Recognition (FSAR). The proposed MTFAN model enjoys several merits. First, we design a motion modulator conditioned on the learned task-specific motion embeddings, which can activate the channels related to the task-shared motion patterns for each frame. Second, a segment attention mechanism is proposed to automatically discover the higher-level segments for multi-level temporal fragment alignment, which encompasses the frame-to-frame, segment-to-segment, and segment-to-frame alignments. To the best of our knowledge, this is the first work to exploit task-specific motion modulation for FSAR. Extensive experimental results on four standard benchmarks demonstrate that the proposed model performs favorably against the state-of-the-art FSAR methods.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionFew-Shot action recognitionFew Shot Action Recognitionimage-classificationImage Classification

Similar Papers 제목 키워드 기반

FAME: Fairness-aware Attention-modulated Video Editing

2025-10-27 · Zhangkai Wu, Xuhui Fan, Zhongyuan Xie, Kaize Shi 외 arxiv

Training-free video editing (VE) models tend to fall back on gender stereotypes when rendering profession-related prompts. We propose \textbf{FAME} for \textit{Fairness-aware Attention-modulated Video Editing} that mitig…

A Unified Conditional Flow for Motion Generation, Editing, and Intra-Structural Retargeting

2026-04-15 · Junlin Li, Xinhao Song, Siqi Wang, Haibin Huang 외 arxiv

Text-driven motion editing and intra-structural retargeting, where source and target share topology but may differ in bone lengths, are traditionally handled by fragmented pipelines with incompatible inputs and represent…

MotionShot: Adaptive Motion Transfer across Arbitrary Objects for Text-to-Video Generation

2025-07-22 · Yanchen Liu, Yanan Sun, Zhening Xing, Junyao Gao 외 arxiv

Existing text-to-video methods struggle to transfer motion smoothly from a reference object to a target object with significant differences in appearance or structure between them. To address this challenge, we introduce…

Text-to-Video Generation

HumanMM: Global Human Motion Recovery from Multi-shot Videos

2025-03-10 · CVPR 2025 1 · Yuhong Zhang, Guanlin Wu, Ling-Hao Chen, Zhuokai Zhao 외

In this paper, we present a novel framework designed to reconstruct long-sequence 3D human motion in the world coordinates from in-the-wild videos with multiple shot transitions. Such long-sequence in-the-wild motions ar…

Camera Pose EstimationMotion GenerationPose Estimation

STAR: Semantic-Temporal Adaptive Representation Learning for Few-Shot Action Recognition

2026-05-13 · Hongli Liu, Yu Wang, Shengjie Zhao arxiv

Few-shot action recognition (FSAR) requires models to generalize to novel action categories from only a handful of annotated samples. Despite progress with vision-language models, existing approaches still suffer from se…

Representation LearningAction Recognition