paper-with-me

홈 › Papers

BIM-Edit: Benchmarking Large Language Models for IFC-Based Building Information Modeling

2026-06-18 · Bharathi Kannan Nithyanantham, Clemens Kujat, Tobias Sesterhenn, Stefan Telgmann, Ashwin Nedungadi, Jörn Plönnigs, Christian Bartelt, Stefan Lüdtke arxiv

Large language models (LLMs) are increasingly applied to computer-aided design (CAD) to generate design artifacts from textual instructions. In engineering practice, this requires more than creating new geometry, models must also understand existing scenes, edit them correctly, and preserve semantics and relations. However, many CAD benchmarks focus on creating new models rather than editing existing ones, and mostly evaluate geometric correctness. We introduce BIM-Edit, a benchmark for evaluating LLMs on natural-language editing of Building Information Models (BIM) represented in the Industry Foundation Classes (IFC) format. BIM provides a challenging testbed because building models encode geometry together with semantic and relational structure. BIM-Edit contains 324 editing tasks spanning 11 realistic building models and 36 synthetic scenes. Tasks are expressed using three instruction categories - direct, spatial, and topological - covering both explicit and scene-grounded edits. We evaluate outputs along three dimensions: geometric accuracy, semantic validity, and topological consistency. Across evaluated LLMs, the best-performing model achieves only 49.5% average score across the three metrics, and no model fully solves more than 3.4% of tasks. These results demonstrate a substantial gap between current LLM capabilities and the requirements of structured engineering design workflows.

📄 PDF Abstract BibTeX arXiv:2606.20146

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RAGAPHENE: A RAG Annotation Platform with Human Enhancements and Edits

2025-08-22 · Kshitij Fadnis, Sara Rosenthal, Maeda Hanafi, Yannis Katsis 외 arxiv

Retrieval Augmented Generation (RAG) is an important aspect of conversing with Large Language Models (LLMs) when factually correct information is important. LLMs may provide answers that appear correct, but could contain…

MANSION: Multi-floor lANguage-to-3D Scene generatIOn for loNg-horizon tasks

2026-03-12 · Lirong Che, Shuo Wen, Shan Huang, Chuang Wang 외 arxiv

Real-world robotic tasks are long-horizon and often span multiple floors, demanding rich spatial reasoning. However, existing embodied benchmarks are largely confined to single-floor in-house environments, failing to ref…

Spatial ReasoningScene Generation

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs

2025-05-26 · Juntong Wang, Jiarui Wang, Huiyu Duan, Guangtao Zhai 외

Text-driven video editing is rapidly advancing, yet its rigorous evaluation remains challenging due to the absence of dedicated video quality assessment (VQA) models capable of discerning the nuances of editing quality. …

BenchmarkingLarge Language ModelVideo EditingVideo Quality Assessment+1

ChartM$^3$: Benchmarking Chart Editing with Multimodal Instructions

2025-07-25 · Donglu Yang, Liang Zhang, Zihao Yue, Liangyu Chen 외 arxiv

Charts are a fundamental visualization format widely used in data analysis across research and industry. While enabling users to edit charts based on high-level intentions is of great practical value, existing methods pr…

Neuro-Symbolic AI for LEED compliance: Document-Centric Benchmarking, Deterministic Numeric Checking, and When Multimodal Hurts

2026-07-17 · Aritro De, Juliana Felkner arxiv

LEED v4.1 BD+C certification remains a document-intensive process that requires reviewers to read hundreds of pages of project evidence and apply credit-specific threshold logic by hand. This paper investigates whether s…