paper-with-me

Papers

Tokenization Allows Multimodal Large Language Models to Understand, Generate and Edit Architectural Floor Plans

2026-03-12 · Sizhong Qin, Ramon Elias Weber, Xinzheng Lu arxiv

Architectural floor plan design demands joint reasoning over geometry, semantics, and spatial hierarchy, which remains a major challenge for current AI systems. Although recent diffusion and language models improve visual fidelity, they still struggle with coherent spatial reasoning and controllable generation. We present HouseMind, a multimodal large language model that unifies floor plan understanding, generation, and editing in one framework. We introduce discrete room-instance tokens to construct a unified vocabulary that bridges layouts and symbolic reasoning. With multimodal alignment and instruction tuning, the model synthesizes coherent, controllable layouts from text instructions. Experiments show how the framework achieves superior geometric validity and controllability while remaining efficient and locally deployable.

📄 PDF Abstract BibTeX arXiv:2603.11640

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models

2024-04-19 · Chuofan Ma, Yi Jiang, Jiannan Wu, Zehuan Yuan 외

We introduce Groma, a Multimodal Large Language Model (MLLM) with grounded and fine-grained visual perception ability. Beyond holistic image understanding, Groma is adept at region-level tasks such as region captioning a…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+3

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

2025-02-07 · Yue Zhao, Fuzhao Xue, Scott Reed, Linxi Fan 외

We introduce Quantized Language-Image Pretraining (QLIP), a visual tokenization method that combines state-of-the-art reconstruction quality with state-of-the-art zero-shot image understanding. QLIP trains a binary-spher…

Image GenerationQuantization

Fine-tuning Multimodal Large Language Models for Product Bundling

2024-07-16 · Xiaohao Liu, Jie Wu, Zhulin Tao, Yunshan Ma 외

Recent advances in product bundling have leveraged multimodal information through sophisticated encoders, but remain constrained by limited semantic understanding and a narrow scope of knowledge. Therefore, some attempts…

In-Context LearningMultiple-choice

ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement

2025-04-02 · Runhui Huang, Chunwei Wang, Junwei Yang, Guansong Lu 외

We present ILLUME+ that leverages dual visual tokenization and a diffusion decoder to improve both deep semantic understanding and high-fidelity image generation. Existing unified models have struggled to simultaneously …

DecoderImage GenerationImage ReconstructionSuper-Resolution+2

Grounding Multimodal Large Language Models in Actions

2024-06-12 · Andrew Szot, Bogdan Mazoure, Harsh Agrawal, Devon Hjelm 외

Multimodal Large Language Models (MLLMs) have demonstrated a wide range of capabilities across many domains, including Embodied AI. In this work, we study how to best ground a MLLM into different embodiments and their as…

World Knowledge