paper-with-me

홈 › Papers

How You Move Your Head Tells What You Do: Self-supervised Video Representation Learning with Egocentric Cameras and IMU Sensors

2021-10-04 · Satoshi Tsutsui, Ruta Desai, Karl Ridgeway

Understanding users' activities from head-mounted cameras is a fundamental task for Augmented and Virtual Reality (AR/VR) applications. A typical approach is to train a classifier in a supervised manner using data labeled by humans. This approach has limitations due to the expensive annotation cost and the closed coverage of activity labels. A potential way to address these limitations is to use self-supervised learning (SSL). Instead of relying on human annotations, SSL leverages intrinsic properties of data to learn representations. We are particularly interested in learning egocentric video representations benefiting from the head-motion generated by users' daily activities, which can be easily obtained from IMU sensors embedded in AR/VR devices. Towards this goal, we propose a simple but effective approach to learn video representation by learning to tell the corresponding pairs of video clip and head-motion. We demonstrate the effectiveness of our learned representation for recognizing egocentric activities of people and dogs.

📄 PDF Abstract BibTeX arXiv:2110.01680

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningSelf-Supervised Learning

Similar Papers 제목 키워드 기반

What My Motion tells me about Your Pose: A Self-Supervised Monocular 3D Vehicle Detector

2020-07-29 · Cédric Picron, Punarjay Chakravarty, Tom Roussel, Tinne Tuytelaars

The estimation of the orientation of an observed vehicle relative to an Autonomous Vehicle (AV) from monocular camera data is an important building block in estimating its 6 DoF pose. Current Deep Learning based solution…

Autonomous VehiclesDomain AdaptationMonocular Visual Odometryvehicle detection+1

How You Move Tells What You'll Do: Trajectory-Conditioned Egocentric Prediction

2026-05-19 · Sejoon Jun, Hai Nguyen-Truong, Luigi Seminara, Lorenzo Torresani arxiv

Predicting how a person's first-person view will evolve (what action will follow, what plan completes a task, whether an in-progress shot will score) is fundamentally under-specified: the same context admits many plausib…

Pose Estimation

As You Are, So Shall You Move Your Head: A System-Level Analysis between Head Movements and Corresponding Traits and Emotions

2019-10-11 · Sharmin Akther Purabi, Rayhan Rashed, Md. Mirajul Islam, Md. Nahiyan Uddin 외

Identifying physical traits and emotions based on system-sensed physical activities is a challenging problem in the realm of human-computer interaction. Our work contributes in this context by investigating an underlying…

Your Pretrained Model Tells the Difficulty Itself: A Self-Adaptive Curriculum Learning Paradigm for Natural Language Understanding

2025-07-13 · Qi Feng, Yihong Liu, Hinrich Schütze arxiv

Curriculum learning is a widely adopted training strategy in natural language processing (NLP), where models are exposed to examples organized by increasing difficulty to enhance learning efficiency and performance. Howe…

Natural Language UnderstandingMulti-class Classification

What Time Tells Us? An Explorative Study of Time Awareness Learned from Static Images

2025-03-23 · Dongheng Lin, Han Hu, Jianbo Jiao

Time becomes visible through illumination changes in what we see. Inspired by this, in this paper we explore the potential to learn time awareness from static images, trying to answer: what time tells us? To this end, we…

Contrastive LearningImage RetrievalScene Classification