paper-with-me

홈 › Papers

EGO-TOPO: Environment Affordances from Egocentric Video

2020-01-14 · CVPR 2020 6 · Tushar Nagarajan, Yanghao Li, Christoph Feichtenhofer, Kristen Grauman

First-person video naturally brings the use of a physical environment to the forefront, since it shows the camera wearer interacting fluidly in a space based on his intentions. However, current methods largely separate the observed actions from the persistent space itself. We introduce a model for environment affordances that is learned directly from egocentric video. The main idea is to gain a human-centric model of a physical space (such as a kitchen) that captures (1) the primary spatial zones of interaction and (2) the likely activities they support. Our approach decomposes a space into a topological map derived from first-person activity, organizing an ego-video into a series of visits to the different zones. Further, we show how to link zones across multiple related environments (e.g., from videos of multiple kitchens) to obtain a consolidated representation of environment functionality. On EPIC-Kitchens and EGTEA+, we demonstrate our approach for learning scene affordances and anticipating future actions in long-form video.

📄 PDF Abstract BibTeX arXiv:2001.04583

Code (1)

facebookresearch/ego-topo 공식 구현 pytorch

Similar Papers 제목 키워드 기반

What to Do Next? Memorizing skills from Egocentric Instructional Video

2025-07-01 · Jing Bi, Chenliang Xu arxiv

Learning to perform activities through demonstration requires extracting meaningful information about the environment from observations. In this research, we investigate the challenge of planning high-level goal-oriented…

DIV-FF: Dynamic Image-Video Feature Fields For Environment Understanding in Egocentric Videos

2025-03-11 · CVPR 2025 1 · Lorenzo Mur-Labadia, Josechu Guerrero, Ruben Martinez-Cantin

Environment understanding in egocentric videos is an important step for applications like robotics, augmented reality and assistive technologies. These videos are characterized by dynamic interactions and a strong depend…

Scene Understanding

Look Ma, No Hands! Agent-Environment Factorization of Egocentric Videos

2023-05-25 · NeurIPS 2023 11

The analysis and use of egocentric videos for robotic tasks is made challenging by occlusion due to the hand and the visual mismatch between the human hand and a robot end-effector. In this sense, the human hand presents…

3D ReconstructionObjectobject-detectionObject Detection+1

Integrating Affordances and Attention models for Short-Term Object Interaction Anticipation

2026-02-16 · Lorenzo Mur Labadia, Ruben Martinez-Cantin, Jose J. Guerrero, Giovanni M. Farinella 외 arxiv

Short Term object-interaction Anticipation consists in detecting the location of the next active objects, the noun and verb categories of the interaction, as well as the time to contact from the observation of egocentric…

Short-term Object Interaction Anticipation

Hi-Dyna Graph: Hierarchical Dynamic Scene Graph for Robotic Autonomy in Human-Centric Environments

2025-05-30 · Jiawei Hou, xiangyang xue, Taiping Zeng

Autonomous operation of service robotics in human-centric scenes remains challenging due to the need for understanding of changing environments and context-aware decision-making. While existing approaches like topologica…

Graph GenerationHuman-Object Interaction DetectionNeRFScene Graph Generation+1