paper-with-me

홈 › Papers

Agent-to-Sim: Learning Interactive Behavior Models from Casual Longitudinal Videos

2024-10-21 · Gengshan Yang, Andrea Bajcsy, Shunsuke Saito, Angjoo Kanazawa

We present Agent-to-Sim (ATS), a framework for learning interactive behavior models of 3D agents from casual longitudinal video collections. Different from prior works that rely on marker-based tracking and multiview cameras, ATS learns natural behaviors of animal and human agents non-invasively through video observations recorded over a long time-span (e.g., a month) in a single environment. Modeling 3D behavior of an agent requires persistent 3D tracking (e.g., knowing which point corresponds to which) over a long time period. To obtain such data, we develop a coarse-to-fine registration method that tracks the agent and the camera over time through a canonical 3D space, resulting in a complete and persistent spacetime 4D representation. We then train a generative model of agent behaviors using paired data of perception and motion of an agent queried from the 4D reconstruction. ATS enables real-to-sim transfer from video recordings of an agent to an interactive behavior simulator. We demonstrate results on pets (e.g., cat, dog, bunny) and human given monocular RGBD videos captured by a smartphone.

📄 PDF Abstract BibTeX arXiv:2410.16259

Code (0)

등록된 구현이 없습니다.

Tasks

4D reconstruction

Similar Papers 제목 키워드 기반

Matrix-game 2.0: An open-source real-time and streaming interactive world model

2025-08-18 · Xianglong He, Chunli Peng, Zexiang Liu, Boyang Wang 외 arxiv

Recent advances in interactive video generations have demonstrated diffusion model's potential as world models by capturing complex physical dynamics and interactive behaviors. However, existing interactive world models …

Video Generation

Hierarchically Structured Neural Bones for Reconstructing Animatable Objects from Casual Videos

2024-08-01 · Subin Jeon, In Cho, Minsu Kim, Woong Oh Cho 외

We propose a new framework for creating and easily manipulating 3D models of arbitrary objects using casually captured videos. Our core ingredient is a novel hierarchy deformation model, which captures motions of objects…

MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos

2024-12-05 · Zhengqi Li, Richard Tucker, Forrester Cole, Qianqian Wang 외

We present a system that allows for accurate, fast, and robust estimation of camera parameters and depth maps from casual monocular videos of dynamic scenes. Most conventional structure from motion and monocular SLAM tec…

Depth Estimation

Interactive Surveillance Technologies for Dense Crowds

2018-09-27 · Bera Aniket, Manocha Dinesh

We present an algorithm for realtime anomaly detection in low to medium density crowd videos using trajectory-level behavior learning. Our formulation combines online tracking algorithms from computer vision, non-linear …

Anomaly Detection

Genie: Generative Interactive Environments

2024-02-23 · Jake Bruce, Michael Dennis, Ashley Edwards, Jack Parker-Holder 외

We introduce Genie, the first generative interactive environment trained in an unsupervised manner from unlabelled Internet videos. The model can be prompted to generate an endless variety of action-controllable virtual …