paper-with-me

홈 › Papers

Ego2World: Compiling Egocentric Cooking Videos into Executable Worlds for Belief-State Planning

2026-05-13 · Qinchuan Cheng, Zhantao Gong, Pengzhan Sun, Angela Yao, Xulei Yang, Shijie Li arxiv

Embodied agents in household environments must plan under partial observation: they need to remember objects, track state changes, and recover when actions fail. Existing benchmarks only partially test this ability. Egocentric video datasets capture realistic human activities but remain passive, while interactive simulators support execution but rely on synthetic scenes and hand-crafted dynamics, introducing a sim-to-real gap and often assuming fully observable state. We introduce Ego2World, an executable benchmark that turns egocentric cooking videos into executable symbolic worlds governed by graph-transition rules. Built on HD-EPIC, Ego2World derives reusable transition rules from video annotations and executes them in a hidden symbolic world graph. During evaluation, the simulator maintains the hidden world graph, while the agent plans over its own partial belief graph using only local observations and execution feedback. This separation forces agents to update memory and replan without observing the true world state. Experiments show that action-overlap scores overestimate physical-state success, and that persistent belief memory improves task completion while reducing repeated visual exploration -- suggesting that belief maintenance should be a first-class target of embodied-agent evaluation.

📄 PDF Abstract BibTeX arXiv:2605.13335

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

V-Nutri: Dish-Level Nutrition Estimation from Egocentric Cooking Videos

2026-04-13 · Chengkun Yue, Chuanzhi Xu, Jiangpeng He arxiv

Nutrition estimation of meals from visual data is an important problem for dietary monitoring and computational health, but existing approaches largely rely on single images of the finally completed dish. This setting is…

Temporal Saliency Adaptation in Egocentric Videos

2018-08-28 · Panagiotis Linardos, Eva Mohedano, Monica Cherto, Cathal Gurrin 외

This work adapts a deep neural model for image saliency prediction to the temporal domain of egocentric video. We compute the saliency map for each video frame, firstly with an off-the-shelf model trained from static ima…

Saliency PredictionVideo Saliency Prediction

Generating Dialogues from Egocentric Instructional Videos for Task Assistance: Dataset, Method and Benchmark

2025-08-15 · Lavisha Aggarwal, Vikas Bahirwani, Lin Li, Andrea Colaco arxiv

Many everyday tasks ranging from fixing appliances, cooking recipes to car maintenance require expert knowledge, especially when tasks are complex and multi-step. Despite growing interest in AI agents, there is a scarcit…

OSCAR: Object Status and Contextual Awareness for Recipes to Support Non-Visual Cooking

2025-03-07 · Franklin Mingzhe Li, Kaitlyn Ng, Bin Zhu, Patrick Carrington

Following recipes while cooking is an important but difficult task for visually impaired individuals. We developed OSCAR (Object Status Context Awareness for Recipes), a novel approach that provides recipe progress track…

Object

From Videos to Conversations: Egocentric Instructions for Task Assistance

2026-02-01 · Lavisha Aggarwal, Vikas Bahirwani, Andrea Colaco arxiv

Many everyday tasks, ranging from appliance repair and cooking to car maintenance, require expert knowledge, particularly for complex, multi-step procedures. Despite growing interest in AI agents for augmented reality (A…