paper-with-me

홈 › Papers

EgoLM: Multi-Modal Language Model of Egocentric Motions

2024-09-26 · CVPR 2025 1 · Fangzhou Hong, Vladimir Guzov, Hyo Jin Kim, Yuting Ye, Richard Newcombe, Ziwei Liu, Lingni Ma

As the prevalence of wearable devices, learning egocentric motions becomes essential to develop contextual AI. In this work, we present EgoLM, a versatile framework that tracks and understands egocentric motions from multi-modal inputs, e.g., egocentric videos and motion sensors. EgoLM exploits rich contexts for the disambiguation of egomotion tracking and understanding, which are ill-posed under single modality conditions. To facilitate the versatile and multi-modal framework, our key insight is to model the joint distribution of egocentric motions and natural languages using large language models (LLM). Multi-modal sensor inputs are encoded and projected to the joint latent space of language models, and used to prompt motion generation or text generation for egomotion tracking or understanding, respectively. Extensive experiments on large-scale multi-modal human motion dataset validate the effectiveness of EgoLM as a generalist model for universal egocentric learning.

📄 PDF Abstract BibTeX arXiv:2409.18127

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingmodelMotion GenerationText Generation

Similar Papers 제목 키워드 기반

EgoPriMo: Egocentric Motion Generation for Interactive Humanoid Control

2026-06-07 · Haoyang Ge, Peng Ren, Yukun Shi, Cong Huang 외 arxiv

Humanoid robots require whole-body motions that adapt to scene context, task requirements, and user intent. Motion tracking reproduces specified trajectories, and humanoid vision-language-action systems provide semantic …

Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions

2025-08-06 · Liang Xu, Chengqun Yang, Zili Lin, Fei Xu 외 arxiv

Learning action models from real-world human-centric interaction datasets is important towards building general-purpose intelligent assistants with efficiency. However, most existing datasets only offer specialist intera…

Towards Continual Egocentric Activity Recognition: A Multi-modal Egocentric Activity Dataset for Continual Learning

2023-01-26 · Linfeng Xu, Qingbo Wu, Lili Pan, Fanman Meng 외

With the rapid development of wearable cameras, a massive collection of egocentric video for first-person visual perception becomes available. Using egocentric videos to predict first-person activity faces many challenge…

Activity RecognitionContinual LearningEgocentric Activity RecognitionHuman Activity Recognition

Immersive Social Interaction with VR and LLM-Assisted Humanoids

2026-07-08 · Niraj Pudasaini, Geeta Chandra Raju Bethala, Pranav Doma, Anthony Tzes 외 arxiv

Humanoid robots can extend human presence to remote, constrained, or hazardous environments, but existing teleoperation interfaces often require physically demanding motion tracking or cognitively demanding low-level con…

EgoCogNav: Cognition-aware Human Egocentric Navigation

2025-11-15 · Zhiwen Qiu, Ziang Liu, Wenqian Niu, Tapomayukh Bhattacharjee 외 arxiv

Modeling the cognitive and experiential factors of human navigation is central to deepening our understanding of human-environment interaction and to enabling safe social navigation and effective assistive wayfinding. Mo…

Motion Forecasting