paper-with-me

홈 › Papers

FirePlace: Geometric Refinements of LLM Common Sense Reasoning for 3D Object Placement

2025-01-01 · CVPR 2025 1 · IAn Huang, Yanan Bao, Karen Truong, Howard Zhou, Cordelia Schmid, Leonidas Guibas, Alireza Fathi

Scene generation with 3D assets presents a complex challenge, requiring both high-level semantic understanding and low-level geometric reasoning. While Multimodal Large Language Models (MLLMs) excel at semantic tasks, their application to 3D scene generation is hindered by their limited grounding on 3D geometry. In this paper, we investigate how to best work with MLLMs in an object placement task. Towards this goal, we introduce a novel framework, FirePlace, that applies existing MLLMs in (1) 3D geometric reasoning and the extraction of relevant geometric details from the 3D scene, (2) constructing and solving geometric constraints on the extracted low-level geometry, and (3) pruning for final placements that conform to common sense. By combining geometric reasoning with real-world understanding of MLLMs, our method can propose object placements that satisfy both geometric constraints as well as high-level semantic common-sense considerations. Our experiments show that these capabilities allow our method to place objects more effectively in complex scenes with intricate geometry, surpassing the quality of prior work.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

3D geometryCommon Sense ReasoningScene Generation

Methods 이 논문이 사용한 방법론

Pruning 설명 없음

Similar Papers 제목 키워드 기반

Visual Commonsense Driven Knowledge Refinements for Scene Graph Generation

2026-06-04 · Maëlic Neau, Salim Baloch, Jakob Suchan, Zoe Falomir 외 arxiv

Learning-driven Scene Graph Generation (SGG) models excel on frequent relation types but degrade sharply under annotation sparsity, failing to capture reliable visual commonsense knowledge. We propose a model-agnostic, s…

Visual Commonsense ReasoningScene Graph Generation

Truth as a Trajectory: What Internal Representations Reveal About Large Language Model Reasoning

2026-03-01 · Hamed Damirchi, Ignacio Meza De la Jara, Ehsan Abbasnejad, Afshar Shamsi 외 arxiv

Existing explainability methods for Large Language Models (LLMs) typically treat hidden states as static points in activation space, assuming that correct and incorrect inferences can be separated using representations f…

Question Answering

GeoSense: Evaluating Identification and Application of Geometric Principles in Multimodal Reasoning

2025-04-17 · Liangyu Xu, Yingxiu Zhao, Jingyun Wang, Yingyao Wang 외

Geometry problem-solving (GPS), a challenging task requiring both visual comprehension and symbolic reasoning, effectively measures the reasoning capabilities of multimodal large language models (MLLMs). Humans exhibit s…

Geometry Problem SolvingMultimodal Reasoning

Evaluate Confidence Instead of Perplexity for Zero-shot Commonsense Reasoning

2022-08-23 · Letian Peng, Zuchao Li, Hai Zhao

Commonsense reasoning is an appealing topic in natural language processing (NLP) as it plays a fundamental role in supporting the human-like actions of NLP systems. With large-scale language models as the backbone, unsup…

Language ModelingLanguage ModellingQuestion AnsweringUnsupervised Pre-training

Leveraging Explicit Reasoning for Inference Integration in Commonsense-Augmented Dialogue Models

2024-06-13 · Sarah E. Finch, Jinho D. Choi

Open-domain dialogue systems need to grasp social commonsense to understand and respond effectively to human users. Commonsense-augmented dialogue models have been proposed that aim to infer commonsense knowledge from di…

Response GenerationSpecificity