Trear: Transformer-based RGB-D Egocentric Action Recognition
In this paper, we propose a \textbf{Tr}ansformer-based RGB-D \textbf{e}gocentric \textbf{a}ction \textbf{r}ecognition framework, called Trear. It consists of two modules, inter-frame attention encoder and mutual-attentional fusion block. Instead of using optical flow or recurrent units, we adopt self-attention mechanism to model the temporal structure of the data from different modalities. Input frames are cropped randomly to mitigate the effect of the data redundancy. Features from each modality are interacted through the proposed fusion block and combined through a simple yet effective fusion operation to produce a joint RGB-D representation. Empirical experiments on two large egocentric RGB-D datasets, THU-READ and FPHA, and one small dataset, WCVS, have shown that the proposed method outperforms the state-of-the-art results by a large margin.
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionOptical Flow EstimationSimilar Papers 제목 키워드 기반
Transformer-based Action recognition in hand-object interacting scenarios
This report describes the 2nd place solution to the ECCV 2022 Human Body, Hands, and Activities (HBHA) from Egocentric and Multi-view Cameras Challenge: Action Recognition. This challenge aims to recognize hand-object in…
Action RecognitionObjectEgoViT: Pyramid Video Transformer for Egocentric Action Recognition
Capturing interaction of hands with objects is important to autonomously detect human actions from egocentric videos. In this work, we present a pyramid video transformer with a dynamic class token generator for egocentr…
Action RecognitionHuman Action Recognition in Egocentric Perspective Using 2D Object and Hands Pose
Egocentric action recognition is essential for healthcare and assistive technology that relies on egocentric cameras because it allows for the automatic and continuous monitoring of activities of daily living (ADLs) with…
Action ClassificationAction RecognitionTemporal Action LocalizationCross-view Action Recognition Understanding From Exocentric to Egocentric Perspective
Understanding action recognition in egocentric videos has emerged as a vital research topic with numerous practical applications. With the limitation in the scale of egocentric data collection, learning robust deep learn…
Action RecognitionIn My Perspective, In My Hands: Accurate Egocentric 2D Hand Pose and Action Recognition
Action recognition is essential for egocentric video understanding, allowing automatic and continuous monitoring of Activities of Daily Living (ADLs) without user effort. Existing literature focuses on 3D hand pose input…
Action RecognitionHand Pose EstimationSkeleton Based Action RecognitionVideo Understanding