EgoDistill: Egocentric Head Motion Distillation for Efficient Video Understanding
Recent advances in egocentric video understanding models are promising, but their heavy computational expense is a barrier for many real-world applications. To address this challenge, we propose EgoDistill, a distillation-based approach that learns to reconstruct heavy egocentric video clip features by combining the semantics from a sparse set of video frames with the head motion from lightweight IMU readings. We further devise a novel self-supervised training strategy for IMU feature learning. Our method leads to significant improvements in efficiency, requiring 200x fewer GFLOPs than equivalent video models. We demonstrate its effectiveness on the Ego4D and EPICKitchens datasets, where our method outperforms state-of-the-art efficient video understanding methods.
Code (0)
등록된 구현이 없습니다.
Tasks
Video UnderstandingMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Ego-Body Pose Estimation via Ego-Head Pose Estimation
Estimating 3D human motion from an egocentric video sequence plays a critical role in human behavior understanding and has various applications in VR/AR. However, naively learning a mapping between egocentric videos and …
BenchmarkingDisentanglementHead Pose EstimationPose EstimationTemporal Segmentation of Egocentric Videos
The use of wearable cameras makes it possible to record life logging egocentric videos. Browsing such long unstructured videos is time consuming and tedious. Segmentation into meaningful chapters is an important first st…
SegmentationVideo SegmentationVideo Semantic SegmentationComputing Egomotion with Local Loop Closures for Egocentric Videos
Finding the camera pose is an important step in many egocentric video applications. It has been widely reported that, state of the art SLAM algorithms fail on egocentric videos. In this paper, we propose a robust method …
Camera Pose EstimationDepth EstimationPose EstimationFirst Person Action Recognition Using Deep Learned Descriptors
We focus on the problem of wearer's action recognition in first person a.k.a. egocentric videos. This problem is more challenging than third person activity recognition due to unavailability of wearer's pose and sharp mo…
Action RecognitionActivity RecognitionTemporal Action LocalizationGenerative Adversarial Network for Future Hand Segmentation from Egocentric Video
We introduce the novel problem of anticipating a time series of future hand masks from egocentric video. A key challenge is to model the stochasticity of future head motions, which globally impact the head-worn camera vi…
Generative Adversarial NetworkHand SegmentationImage SegmentationSemantic Segmentation+2