paper-with-me

홈 › Papers

AnyHome: Open-Vocabulary Generation of Structured and Textured 3D Homes

2023-12-11 · Rao Fu, Zehao Wen, Zichen Liu, Srinath Sridhar

Inspired by cognitive theories, we introduce AnyHome, a framework that translates any text into well-structured and textured indoor scenes at a house-scale. By prompting Large Language Models (LLMs) with designed templates, our approach converts provided textual narratives into amodal structured representations. These representations guarantee consistent and realistic spatial layouts by directing the synthesis of a geometry mesh within defined constraints. A Score Distillation Sampling process is then employed to refine the geometry, followed by an egocentric inpainting process that adds lifelike textures to it. AnyHome stands out with its editability, customizability, diversity, and realism. The structured representations for scenes allow for extensive editing at varying levels of granularity. Capable of interpreting texts ranging from simple labels to detailed narratives, AnyHome generates detailed geometries and textures that outperform existing methods in both quantitative and qualitative measures.

📄 PDF Abstract BibTeX arXiv:2312.06644

Code (0)

등록된 구현이 없습니다.

Tasks

Diversity

Methods 이 논문이 사용한 방법론

Inpainting Train a convolutional neural network to generate the contents of an arbitrary image region conditioned on its surroundings.
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

PartNeXt: A Next-Generation Dataset for Fine-Grained and Hierarchical 3D Part Understanding

2025-10-23 · Penghao Wang, Yiyang He, Xin Lv, Yukai Zhou 외 arxiv

Understanding objects at the level of their constituent parts is fundamental to advancing computer vision, graphics, and robotics. While datasets like PartNet have driven progress in 3D part understanding, their reliance…

Question Answering

TAPS3D: Text-Guided 3D Textured Shape Generation from Pseudo Supervision

2023-03-23 · CVPR 2023 1 · Jiacheng Wei, Hao Wang, Jiashi Feng, Guosheng Lin 외

In this paper, we investigate an open research task of generating controllable 3D textured shapes from the given textual descriptions. Previous works either require ground truth caption labeling or extensive optimization…

Diversity

RWT-SLAM: Robust Visual SLAM for Highly Weak-textured Environments

2022-07-07 · Qihao Peng, Zhiyu Xiang, YuanGang Fan, Tengqi Zhao 외

As a fundamental task for intelligent robots, visual SLAM has made great progress over the past decades. However, robust SLAM under highly weak-textured environments still remains very challenging. In this paper, we prop…

Expanding Scene Graph Boundaries: Fully Open-vocabulary Scene Graph Generation via Visual-Concept Alignment and Retention

2023-11-18 · Zuyao Chen, Jinlin Wu, Zhen Lei, Zhaoxiang Zhang 외

Scene Graph Generation (SGG) offers a structured representation critical in many computer vision applications. Traditional SGG approaches, however, are limited by a closed-set assumption, restricting their ability to rec…

Concept AlignmentGraph GenerationKnowledge DistillationObject+6

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations

2024-12-09 · Mingjie Xu, Mengyang Wu, Yuzhi Zhao, Jason Chun Lok Li 외

Scene Graph Generation (SGG) converts visual scenes into structured graph representations, providing deeper scene understanding for complex vision tasks. However, existing SGG models often overlook essential spatial rela…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+3