paper-with-me

홈 › Papers

ActAR: Actor-Driven Pose Embeddings for Video Action Recognition

2022-04-19 · Soufiane Lamghari, Guillaume-Alexandre Bilodeau, Nicolas Saunier

Human action recognition (HAR) in videos is one of the core tasks of video understanding. Based on video sequences, the goal is to recognize actions performed by humans. While HAR has received much attention in the visible spectrum, action recognition in infrared videos is little studied. Accurate recognition of human actions in the infrared domain is a highly challenging task because of the redundant and indistinguishable texture features present in the sequence. Furthermore, in some cases, challenges arise from the irrelevant information induced by the presence of multiple active persons not contributing to the actual action of interest. Therefore, most existing methods consider a standard paradigm that does not take into account these challenges, which is in some part due to the ambiguous definition of the recognition task in some cases. In this paper, we propose a new method that simultaneously learns to recognize efficiently human actions in the infrared spectrum, while automatically identifying the key-actors performing the action without using any prior knowledge or explicit annotations. Our method is composed of three stages. In the first stage, optical flow-based key-actor identification is performed. Then for each key-actor, we estimate key-poses that will guide the frame selection process. A scale-invariant encoding process along with embedded pose filtering are performed in order to enhance the quality of action representations. Experimental results on InfAR dataset show that our proposed model achieves promising recognition performance and learns useful action representations.

📄 PDF Abstract BibTeX arXiv:2204.08671

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionOptical Flow EstimationTemporal Action LocalizationVideo Understanding

Similar Papers 제목 키워드 기반

Towards Comprehensive Stage-wise Benchmarking of Large Language Models in Fact-Checking

2026-01-06 · Hongzhan Lin, Zixin Chen, Zhiqi Shen, Ziyang Luo 외 arxiv

Large Language Models (LLMs) are increasingly deployed in real-world fact-checking systems, yet existing evaluations focus predominantly on claim verification and overlook the broader fact-checking workflow, including cl…

ContactArt: Learning 3D Interaction Priors for Category-level Articulated Object and Hand Poses Estimation

2023-05-02 · Zehao Zhu, Jiashun Wang, Yuzhe Qin, Deqing Sun 외

We propose a new dataset and a novel approach to learning hand-object interaction priors for hand and articulated object pose estimation. We first collect a dataset using visual teleoperation, where the human operator ca…

Hand Pose EstimationObjectPose Estimation

Text2Performer: Text-Driven Human Video Generation

2023-04-17 · ICCV 2023 1 · Yuming Jiang, Shuai Yang, Tong Liang Koh, Wayne Wu 외

Text-driven content creation has evolved to be a transformative technique that revolutionizes creativity. Here we study the task of text-driven human video generation, where a video sequence is synthesized from texts des…

Video Generation

StableAvatar: Infinite-Length Audio-Driven Avatar Video Generation

2025-08-11 · Shuyuan Tu, Yueming Pan, Yinming Huang, Xintong Han 외 arxiv

Current diffusion models for audio-driven avatar video generation struggle to synthesize long videos with natural audio synchronization and identity consistency. This paper presents StableAvatar, the first end-to-end vid…

Video Generation

X-Streamer: Unified Human World Modeling with Audiovisual Interaction

2025-09-25 · You Xie, Tianpei Gu, Zenan Li, Chenxu Zhang 외 arxiv

We introduce X-Streamer, an end-to-end multimodal human world modeling framework for building digital human agents capable of infinite interactions across text, speech, and video within a single unified architecture. Sta…