paper-with-me

홈 › Papers

A$^2$M$^2$-Net: Adaptively Aligned Multi-Scale Moment for Few-Shot Action Recognition

2025-09-22 · Zilin Gao, Qilong Wang, Bingbing Zhang, Qinghua Hu, Peihua Li arxiv

Thanks to capability to alleviate the cost of large-scale annotation, few-shot action recognition (FSAR) has attracted increased attention of researchers in recent years. Existing FSAR approaches typically neglect the role of individual motion pattern in comparison, and under-explore the feature statistics for video dynamics. Thereby, they struggle to handle the challenging temporal misalignment in video dynamics, particularly by using 2D backbones. To overcome these limitations, this work proposes an adaptively aligned multi-scale second-order moment network, namely A$^2$M$^2$-Net, to describe the latent video dynamics with a collection of powerful representation candidates and adaptively align them in an instance-guided manner. To this end, our A$^2$M$^2$-Net involves two core components, namely, adaptive alignment (A$^2$ module) for matching, and multi-scale second-order moment (M$^2$ block) for strong representation. Specifically, M$^2$ block develops a collection of semantic second-order descriptors at multiple spatio-temporal scales. Furthermore, A$^2$ module aims to adaptively select informative candidate descriptors while considering the individual motion pattern. By such means, our A$^2$M$^2$-Net is able to handle the challenging temporal misalignment problem by establishing an adaptive alignment protocol for strong representation. Notably, our proposed method generalizes well to various few-shot settings and diverse metrics. The experiments are conducted on five widely used FSAR benchmarks, and the results show our A$^2$M$^2$-Net achieves very competitive performance compared to state-of-the-arts, demonstrating its effectiveness and generalization.

📄 PDF Abstract BibTeX arXiv:2509.17638

Code (0)

등록된 구현이 없습니다.

Tasks

Action Recognition

Similar Papers 제목 키워드 기반

MAN: Moment Alignment Network for Natural Language Moment Retrieval via Iterative Graph Adjustment

2018-11-30 · CVPR 2019 6 · Da Zhang, Xiyang Dai, Xin Wang, Yuan-Fang Wang 외

This research strives for natural language moment retrieval in long, untrimmed video streams. The problem is not trivial especially when a video contains multiple moments of interests and the language describes complex t…

Moment RetrievalNatural Language Moment RetrievalRetrieval

Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models

2025-01-14 · Yifang Xu, Yunzhuo Sun, Benxiang Zhai, Ming Li 외

The target of video moment retrieval (VMR) is predicting temporal spans within a video that semantically match a given linguistic query. Existing VMR methods based on multimodal large language models (MLLMs) overly rely …

Moment RetrievalRetrieval

Multi-Scale 2D Temporal Adjacent Networks for Moment Localization with Natural Language

2020-12-04 · Songyang Zhang, Houwen Peng, Jianlong Fu, Yijuan Lu 외

We address the problem of retrieving a specific moment from an untrimmed video by natural language. It is a challenging problem because a target moment may take place in the context of other temporal moments in the untri…

See More, Store Less: Memory-Efficient Resolution for Video Moment Retrieval

2026-01-14 · Mingyu Jeon, Sungjin Han, Jinkwon Hwang, Minchol Kwon 외 arxiv

Recent advances in Multimodal Large Language Models (MLLMs) have improved image recognition and reasoning, but video-related tasks remain challenging due to memory constraints from dense frame processing. Existing Video …

Moment Retrieval

Knowledge-Refined Dual Context-Aware Network for Partially Relevant Video Retrieval

2026-03-25 · Junkai Yang, Qirui Wang, Yaoqing Jin, Shuai Ma 외 arxiv

Retrieving partially relevant segments from untrimmed videos remains difficult due to two persistent challenges: the mismatch in information density between text and video segments, and limited attention mechanisms that …

Partially Relevant Video Retrieval