paper-with-me

홈 › Papers

LEGO: Learning EGOcentric Action Frame Generation via Visual Instruction Tuning

2023-12-06 · Bolin Lai, Xiaoliang Dai, Lawrence Chen, Guan Pang, James M. Rehg, Miao Liu

Generating instructional images of human daily actions from an egocentric viewpoint serves as a key step towards efficient skill transfer. In this paper, we introduce a novel problem -- egocentric action frame generation. The goal is to synthesize an image depicting an action in the user's context (i.e., action frame) by conditioning on a user prompt and an input egocentric image. Notably, existing egocentric action datasets lack the detailed annotations that describe the execution of actions. Additionally, existing diffusion-based image manipulation models are sub-optimal in controlling the state transition of an action in egocentric image pixel space because of the domain gap. To this end, we propose to Learn EGOcentric (LEGO) action frame generation via visual instruction tuning. First, we introduce a prompt enhancement scheme to generate enriched action descriptions from a visual large language model (VLLM) by visual instruction tuning. Then we propose a novel method to leverage image and text embeddings from the VLLM as additional conditioning to improve the performance of a diffusion model. We validate our model on two egocentric datasets -- Ego4D and Epic-Kitchens. Our experiments show substantial improvement over prior image manipulation models in both quantitative and qualitative evaluation. We also conduct detailed ablation studies and analysis to provide insights in our method. More details of the dataset and code are available on the website (https://bolinlai.github.io/Lego_EgoActGen/).

📄 PDF Abstract BibTeX arXiv:2312.03849

Code (0)

등록된 구현이 없습니다.

Tasks

Image ManipulationLanguage ModellingLarge Language Model

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

3D Human Pose Perception from Egocentric Stereo Videos

2023-12-30 · CVPR 2024 1 · Hiroyasu Akada, Jian Wang, Vladislav Golyanik, Christian Theobalt

While head-mounted devices are becoming more compact, they provide egocentric views with significant self-occlusions of the device user. Hence, existing methods often fail to accurately estimate complex 3D poses from ego…

3D Human Pose Estimation3D Scene ReconstructionEgocentric Pose EstimationPose Estimation

UnrealEgo: A New Dataset for Robust Egocentric 3D Human Motion Capture

2022-08-02 · Hiroyasu Akada, Jian Wang, Soshi Shimada, Masaki Takahashi 외

We present UnrealEgo, i.e., a new large-scale naturalistic dataset for egocentric 3D human pose estimation. UnrealEgo is based on an advanced concept of eyeglasses equipped with two fisheye cameras that can be used in un…

3D Human Pose EstimationEgocentric Pose EstimationKeypoint EstimationPose Estimation

Interact with me: Joint Egocentric Forecasting of Intent to Interact, Attitude and Social Actions

2024-12-21 · Tongfei Bian, Yiming Ma, Mathieu Chollet, Victor Sanchez 외

For efficient human-agent interaction, an agent should proactively recognize their target user and prepare for upcoming interactions. We formulate this challenging problem as the novel task of jointly forecasting a perso…

TSR-Ego: Temporally Guided Stereo Refinement Framework for Egocentric 3D Human Pose Estimation

2026-07-10 · Md Mushfiqur Azam, John Quarles, Kevin Desai arxiv

Egocentric 3D human pose estimation from head-mounted stereo cameras is challenging due to fisheye distortion, severe self-occlusion, and frequent truncation of body joints outside the camera field of view. Recent stereo…

3D Human Pose EstimationPose Prediction

EgoReAct: Egocentric Video-Driven 3D Human Reaction Generation

2025-12-28 · Libo Zhang, Zekun Li, Tianyu Li, Zeyu Cao 외 arxiv

Humans exhibit adaptive, context-sensitive responses to egocentric visual input. However, faithfully modeling such reactions from egocentric video remains challenging due to the dual requirements of strictly causal gener…