paper-with-me

홈 › Papers

Learning to Evolve Scenes: Reasoning about Human Activities with Scene Graphs

2026-07-02 · Francesca Pistilli, Simone Alberto Peirone, Giuseppe Averta arxiv

Understanding human behavior while interacting with the surrounding world is crucial for many applications of embodied AI. First-person videos are particularly informative for this problem, as they well capture how activities reshape the scene over time. However, existing approaches often rely on implicit visual or language-aligned representations, disregarding structured reasoning over the scene dynamic. We argue that explicit, compositional and editable representations of human-environment interactions can play a crucial role for rich grounded activity understanding. To this end, we introduce SG-Ego, a large scale annotation set extending Ego4D with spatio-temporal scene graphs, where relations triplets are consolidated over time into explicit time-evolving descriptions of the scene state. To reason over this representation, we propose GLEN, a graph-based model that operates over scene graph sequences to both align them with textual actions and model their temporal evolution. In addition, we formulate the activity-driven graph-edit forecasting (A-GEF) problem, a novel task that casts scene dynamics as a sequence of structured transformations conditioned on ongoing actions, enabling explicit reasoning about how scenes change over time. We validate our approach across multiple downstream tasks, spanning retrieval benchmarks as EgoMCQ and EgoCVR, as well as long-horizon reasoning benchmarks as EXPLORE-Bench and the newly introduced A-GEF. GLEN achieves strong results compared to raw video baselines and it excels in reasoning settings, typically addressed only with MLLMs, while enabling controllable and structured predictions of scene dynamics driven by human activities. We believe our results establish spatio-temporal scene graphs, together with models that reason over them, as strong compositional and interpretable representations for video understanding and potentially beyond.

📄 PDF Abstract BibTeX arXiv:2607.02425

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

HiERO: understanding the hierarchy of human behavior enhances reasoning on egocentric videos

2025-05-19 · Simone Alberto Peirone, Francesca Pistilli, Giuseppe Averta

Human activities are particularly complex and variable, and this makes challenging for deep learning models to reason about them. However, we note that such variability does have an underlying structure, composed of a hi…

Procedure Learning

A cookbook of translating English to Xapi

2013-03-31 · Ladislau Bölöni

The Xapagy cognitive architecture had been designed to perform narrative reasoning: to model and mimic the activities performed by humans when witnessing, reading, recalling, narrating and talking about stories. Xapagy c…

COLD: Causal reasOning in cLosed Daily activities

2024-11-29 · Abhinav Joshi, Areeb Ahmad, Ashutosh Modi

Large Language Models (LLMs) have shown state-of-the-art performance in a variety of tasks, including arithmetic and reasoning; however, to gauge the intellectual capabilities of LLMs, causal reasoning has become a relia…

Causal InferenceCommonsense Causal ReasoningEvent Causality IdentificationQuestion Answering

JRDB-Pose3D: A Multi-person 3D Human Pose and Shape Estimation Dataset for Robotics

2026-02-03 · Sandika Biswas, Kian Izadpanah, Hamid Rezatofighi arxiv

Real-world scenes are inherently crowded. Hence, estimating 3D poses of all nearby humans, tracking their movements over time, and understanding their activities within social and environmental contexts are essential for…

3D human pose and shape estimation3D Human Pose EstimationAutonomous DrivingRobot Navigation

Rearrange Indoor Scenes for Human-Robot Co-Activity

2023-03-10 · Weiqi Wang, Zihang Zhao, Ziyuan Jiao, Yixin Zhu 외

We present an optimization-based framework for rearranging indoor furniture to accommodate human-robot co-activities better. The rearrangement aims to afford sufficient accessible space for robot activities without compr…