paper-with-me

홈 › Papers

DreamCAD: Scaling Multi-modal CAD Generation using Differentiable Parametric Surfaces

2026-03-05 · Mohammad Sadil Khan, Muhammad Usama, Rolandos Alexandros Potamias, Didier Stricker, Muhammad Zeshan Afzal, Jiankang Deng, Ismail Elezi arxiv

Computer-Aided Design (CAD) relies on structured and editable geometric representations, yet existing generative methods are constrained by small annotated datasets with explicit design histories or boundary representation (BRep) labels. Meanwhile, millions of unannotated 3D meshes remain untapped, limiting progress in scalable CAD generation. To address this, we propose DreamCAD, a multi-modal generative framework that directly produces editable BReps from point-level supervision, without CAD-specific annotations. DreamCAD represents each BRep as a set of parametric patches (e.g., Bézier surfaces) and uses a differentiable tessellation method to generate meshes. This enables large-scale training on 3D datasets while reconstructing connected and editable surfaces. Furthermore, we introduce CADCap-1M, the largest CAD captioning dataset to date, with 1M+ descriptions generated using GPT-5 for advancing text-to-CAD research. DreamCAD achieves state-of-the-art performance on ABC and Objaverse benchmarks across text, image, and point modalities, improving geometric fidelity and surpassing 75% user preference. Code and dataset will be publicly available.

📄 PDF Abstract BibTeX arXiv:2603.05607

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

UniT: Unified Multimodal Chain-of-Thought Test-time Scaling

2026-02-12 · Leon Liangyu Chen, Haoyu Ma, Zhipeng Fan, Ziqi Huang 외 arxiv

Unified models can handle both multimodal understanding and generation within a single architecture, yet they typically operate in a single pass without iteratively refining their outputs. Many multimodal tasks, especial…

Visual Reasoning

ImAgent: A Unified Multimodal Agent Framework for Test-Time Scalable Image Generation

2025-11-14 · Kaishen Wang, Ruibo Chen, Tong Zheng, Heng Huang arxiv

Recent text-to-image (T2I) models have made remarkable progress in generating visually realistic and semantically coherent images. However, they still suffer from randomness and inconsistency with the given prompts, part…

Image Generation

Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

2025-01-29 · Xiaokang Chen, Zhiyu Wu, Xingchao Liu, Zizheng Pan 외

In this work, we introduce Janus-Pro, an advanced version of the previous work Janus. Specifically, Janus-Pro incorporates (1) an optimized training strategy, (2) expanded training data, and (3) scaling to larger model s…

Image GenerationInstruction FollowingText to Image Generation+2

Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning

2026-06-06 · Cong Wan, Ying He, Zhongzhan Huang, Hefeng Wu arxiv

Test-time Scaling (TTS) has emerged as a pivotal research direction for enhancing model performance by dynamically allocating computational resources during inference. Recent advancements have adapted this paradigm to Mu…

Multimodal Reasoning

SIC3D: Style Image Conditioned Text-to-3D Gaussian Splatting Generation

2026-04-09 · Ming He, Zhixiang Chen, Steve Maddock arxiv

Recent progress in text-to-3D object generation enables the synthesis of detailed geometry from text input by leveraging 2D diffusion models and differentiable 3D representations. However, the approaches often suffer fro…

3D Generation