paper-with-me

Papers

InHabit: Leveraging Image Foundation Models for Scalable 3D Human Placement

2026-04-21 · Nikita Kister, Pradyumna YM, István Sárándi, Jiayi Wang, Anna Khoreva, Gerard Pons-Moll arxiv

Training embodied agents to understand 3D scenes as humans do requires large-scale data of people meaningfully interacting with diverse environments, yet such data is scarce. Real-world capture is costly and limited to controlled settings, while existing synthetic datasets rely on simple geometric heuristics, ignoring rich scene context. In contrast, 2D foundation models trained at internet scale have acquired commonsense knowledge of human-environment interactions. To transfer this knowledge to 3D, we introduce InHabit, an automatic and scalable data generator for populating 3D scenes with interacting humans. InHabit follows a render-generate-lift principle: given a rendered 3D scene, a vision-language model proposes contextually meaningful actions, an image-editing model inserts a human, and an optimization procedure lifts the edited result into physically plausible SMPL-X bodies aligned with the scene geometry. Applied to Habitat-Matterport3D, InHabit produces InHabitants, the first large-scale photorealistic 3D human-scene interaction dataset, with 78K samples across $\sim$800 building-scale scenes with complete 3D geometry, SMPL-X bodies, and images. Augmenting standard training data with InHabitants improves RGB-based 3D human-scene reconstruction and contact estimation, and in a perceptual user study our data is preferred in 78% of cases over prior art.

📄 PDF Abstract BibTeX arXiv:2604.19673

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Perceptive Behavior Foundation Model: Adapting Human Motion Priors to Robot-Centric Terrain

2026-06-06 · Zifan Wang, Yizhao Li, Teli Ma, Qiang Zhang 외 arxiv

Humanoid behavior foundation models aim to acquire reusable whole-body control policies from broad human motion priors, enabling a single controller to produce diverse and expressive behaviors. However, existing motion-c…

CrabOS: An Operating System for Human-AI Co-inhabitation

2026-08-28 · Qi Yang, Yun Ma arxiv

AI agents are evolving into long-running computational entities that can invoke tools, maintain memory, and complete complex tasks across applications. In real-world settings, completing a task often requires humans and …

A Survey on Human Interaction Motion Generation

2025-03-17 · Kewei Sui, Anindita Ghosh, Inwoo Hwang, Jian Wang 외

Humans inhabit a world defined by interactions -- with other humans, objects, and environments. These interactive movements not only convey our relationships with our surroundings but also demonstrate how we perceive and…

Human DynamicsMotion GenerationSurvey

Leveraging ChatGPT's Multimodal Vision Capabilities to Rank Satellite Images by Poverty Level: Advancing Tools for Social Science Research

2025-01-24 · Hamid Sarmadi, Ola Hall, Thorsteinn Rögnvaldsson, Mattias Ohlsson

This paper investigates the novel application of Large Language Models (LLMs) with vision capabilities to analyze satellite imagery for village-level poverty prediction. Although LLMs were originally designed for natural…

Natural Language Understanding

Discriminating sensor activation in activity recognition within multi-occupancy environments based on nearby interaction

2022-11-03 · Aurora Polo-Rodriguez, Javier Medina-Quero

This work presents a computer model to discriminate sensor activation in multi-occupancy environments based on proximity interaction. Current proximity-based and indoor location methods allow the estimation of the positi…

Activity Recognition