paper-with-me

홈 › Papers

Text-to-Scene with Large Reasoning Models

2025-09-30 · Frédéric Berdoz, Luca A. Lanzendörfer, Nick Tuninga, Roger Wattenhofer arxiv

Prompt-driven scene synthesis allows users to generate complete 3D environments from textual descriptions. Current text-to-scene methods often struggle with complex geometries and object transformations, and tend to show weak adherence to complex instructions. We address these limitations by introducing Reason-3D, a text-to-scene model powered by large reasoning models (LRMs). Reason-3D integrates object retrieval using captions covering physical, functional, and contextual attributes. Reason-3D then places the selected objects based on implicit and explicit layout constraints, and refines their positions with collision-aware spatial reasoning. Evaluated on instructions ranging from simple to complex indoor configurations, Reason-3D significantly outperforms previous methods in human-rated visual fidelity, adherence to constraints, and asset retrieval quality. Beyond its contribution to the field of text-to-scene generation, our work showcases the advanced spatial reasoning abilities of modern LRMs. Additionally, we release the codebase to further the research in object retrieval and placement with LRMs.

📄 PDF Abstract BibTeX arXiv:2509.26091

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial ReasoningScene Generation

Similar Papers 제목 키워드 기반

TextOCR: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text

2021-05-12 · CVPR 2021 1 · Amanpreet Singh, Guan Pang, Mandy Toh, Jing Huang 외

A crucial component for the scene text based reasoning required for TextVQA and TextCaps datasets involve detecting and recognizing text present in the images using an optical character recognition (OCR) system. The curr…

Optical Character RecognitionOptical Character Recognition (OCR)Scene Text DetectionText Detection+1

MLLM-For3D: Adapting Multimodal Large Language Model for 3D Reasoning Segmentation

2025-03-23 · Jiaxin Huang, Runnan Chen, Ziwen Li, Zhengqing Gao 외

Reasoning segmentation aims to segment target objects in complex scenes based on human intent and spatial reasoning. While recent multimodal large language models (MLLMs) have demonstrated impressive 2D image reasoning s…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+4

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM

2026-03-29 · Haifeng Huang, Yilun Chen, Zehan Wang, Jiangmiao Pang 외 arxiv

Recent advancements in multi-modal large language models (MLLMs) have shown strong potential for 3D scene understanding. However, existing methods struggle with fine-grained object grounding and contextual reasoning, lim…

Scene UnderstandingSpatial Reasoning

SceneAlign: Aligning Multimodal Reasoning to Scene Graphs in Complex Visual Scenes

2026-01-09 · Chuhan Wang, Xintong Li, Jennifer Yuntong Zhang, Junda Wu 외 arxiv

Multimodal large language models often struggle with faithful reasoning in complex visual scenes, where intricate entities and relations require precise visual grounding at each step. This reasoning unfaithfulness freque…

Multimodal ReasoningVisual ReasoningVisual Grounding

VRSA: Jailbreaking Multimodal Large Language Models through Visual Reasoning Sequential Attack

2025-12-05 · Shiji Zhao, Shukun Xiong, Yao Huang, Yan Jin 외 arxiv

Multimodal Large Language Models (MLLMs) are widely used in various fields due to their powerful cross-modal comprehension and generation capabilities. However, more modalities bring more vulnerabilities to being utilize…

Visual Reasoning