paper-with-me

홈 › Papers

RefAtomNet++: Advancing Referring Atomic Video Action Recognition using Semantic Retrieval based Multi-Trajectory Mamba

2025-10-18 · Kunyu Peng, Di Wen, Jia Fu, Jiamin Wu, Kailun Yang, Junwei Zheng, Ruiping Liu, Yufan Chen, Yuqian Fu, Danda Pani Paudel, Luc Van Gool, Rainer Stiefelhagen arxiv

Referring Atomic Video Action Recognition (RAVAR) aims to recognize fine-grained, atomic-level actions of a specific person of interest conditioned on natural language descriptions. Distinct from conventional action recognition and detection tasks, RAVAR emphasizes precise language-guided action understanding, which is particularly critical for interactive human action analysis in complex multi-person scenarios. In this work, we extend our previously introduced RefAVA dataset to RefAVA++, which comprises >2.9 million frames and >75.1k annotated persons in total. We benchmark this dataset using baselines from multiple related domains, including atomic action localization, video question answering, and text-video retrieval, as well as our earlier model, RefAtomNet. Although RefAtomNet surpasses other baselines by incorporating agent attention to highlight salient features, its ability to align and retrieve cross-modal information remains limited, leading to suboptimal performance in localizing the target person and predicting fine-grained actions. To overcome the aforementioned limitations, we introduce RefAtomNet++, a novel framework that advances cross-modal token aggregation through a multi-hierarchical semantic-aligned cross-attention mechanism combined with multi-trajectory Mamba modeling at the partial-keyword, scene-attribute, and holistic-sentence levels. In particular, scanning trajectories are constructed by dynamically selecting the nearest visual spatial tokens at each timestep for both partial-keyword and scene-attribute levels. Moreover, we design a multi-hierarchical semantic-aligned cross-attention strategy, enabling more effective aggregation of spatial and temporal tokens across different semantic hierarchies. Experiments show that RefAtomNet++ establishes new state-of-the-art results. The dataset and code are released at https://github.com/KPeng9510/refAVA2.

📄 PDF Abstract BibTeX arXiv:2510.16444

Code (0)

등록된 구현이 없습니다.

Tasks

Video Question AnsweringAction UnderstandingAction RecognitionSemantic Retrieval

Similar Papers 제목 키워드 기반

Referring Atomic Video Action Recognition

2024-07-02 · Kunyu Peng, Jia Fu, Kailun Yang, Di Wen 외

We introduce a new task called Referring Atomic Video Action Recognition (RAVAR), aimed at identifying atomic actions of a particular person based on a textual description and the video data of this person. This task dif…

Action LocalizationAction RecognitionQuestion AnsweringReferring Expression+3

Exploring Modulated Detection Transformer as a Tool for Action Recognition in Videos

2022-09-21 · Tomás Crisol, Joel Ermantraut, Adrián Rostagno, Santiago L. Aggio 외

During recent years transformers architectures have been growing in popularity. Modulated Detection Transformer (MDETR) is an end-to-end multi-modal understanding model that performs tasks such as phase grounding, referr…

Action DetectionAction RecognitionAction Recognition In VideosQuestion Answering+5

InterRVOS: Interaction-aware Referring Video Object Segmentation

2025-06-03 · Woojeong Jin, Seongchan Kim, Seungryong Kim

Referring video object segmentation aims to segment the object in a video corresponding to a given natural language expression. While prior works have explored various referring scenarios, including motion-centric or mul…

8kObjectReferring Video Object SegmentationSemantic Segmentation+3

Cross-Modal Progressive Comprehension for Referring Segmentation

2021-05-15 · Si Liu, Tianrui Hui, Shaofei Huang, Yunchao Wei 외

Given a natural language expression and an image/video, the goal of referring segmentation is to produce the pixel-level masks of the entities described by the subject of the expression. Previous approaches tackle this p…

AttributeImage SegmentationReferring Expression SegmentationSegmentation+3

Referring Segmentation in Images and Videos with Cross-Modal Self-Attention Network

2021-02-09 · Linwei Ye, Mrigank Rochan, Zhi Liu, Xiaoqin Zhang 외

We consider the problem of referring segmentation in images and videos with natural language. Given an input image (or video) and a referring expression, the goal is to segment the entity referred by the expression in th…

Referring ExpressionReferring Expression SegmentationSegmentationVideo Segmentation+1