paper-with-me

홈 › Papers

Mira-Scene: Pixel-Aligned Layouts for Generative 3D Scene

2026-09-20 · Yang-Tian Sun, Tianjia Liu, Zehuan Huang, Yi-Hua Huang, Xiaoyang Lyu, Ziyi Yang, Zi-Xin Zou, Yuan-Chen Guo, Yan-Pei Cao, Xiaojuan Qi hf

Single-image 3D object generation can now produce high-fidelity assets, yet accurately placing them into a coherent scene layout remains an open challenge. A central difficulty lies in how object layout is represented. Holistic methods absorb placement into a scene-level generation process, sacrificing object-level detail. Compositional methods preserve object fidelity by decoupling geometry from layout, but typically parameterize layout as sparse, unbounded pose variables that are difficult to learn and generalize poorly under scarce scene-level supervision.We present Mira-Scene, a compositional 3D scene reconstruction framework that replaces sparse pose regression with dense, bounded correspondence recovery. At its core is the Canonical Coordinate Map (CCM), a pixel-aligned field that maps each visible object pixel to a surface coordinate in the object's bounded canonical space. When paired with a scene-space Point Cloud Map (PCM) from monocular geometry estimation, CCM induces dense canonical-to-scene correspondences from which object transformations are recovered through robust geometric alignment. Because CCM operates in bounded canonical space, it provides a stable prediction target that can be trained from scalable object-level 3D data without requiring scene-level layout annotations. Mira-Scene further introduces a multimodal diffusion transformer that jointly generates object geometry and CCMs, using modality-specific expert streams with shared attention and positional encoding to promote geometry-layout consistency. Experiments on indoor, outdoor, synthetic, and in-the-wild scenes show that Mira-Scene substantially outperforms strong baselines in layout accuracy, achieving relative gains of 39.8% in 3D-IoU and 16.5% in 2D-IoU over SAM3D, using limited open-source training data.

📄 PDF Abstract BibTeX arXiv:2609.23796

Code (2)

VAST-AI-Research/Mira-Scene ★ 46
Valiant-Cat/hfpaper

Similar Papers 제목 키워드 기반

InfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View Synthesis

2026-08-03 · Jiawei Wang, Hao Yu, Yongzhen Hu, Xinyi Yang 외 arxiv

Single-image feed-forward 3D Gaussian Splatting (3DGS) aims to directly generate a renderable 3D scene representation from one input image, avoiding the cost of multi-view capture and per-scene optimization. However, exi…

Zero-shot Generalization

Systematic Discovery of Semantic Attacks in Online Map Construction through Conditional Diffusion

2026-05-14 · Chenyi Wang, Ruoyu Song, Raymond Muller, Jean-Philippe Monteuuis 외 arxiv

Autonomous vehicles depend on online HD map construction to perceive lane boundaries, dividers, and pedestrian crossings -- safety-critical road elements that directly govern motion planning. While existing pixel perturb…

Autonomous VehiclesMotion Planning

FlyMirage: A Fully Automated Generation Pipeline for Diverse and Scalable UAV Flight Data via Generative World Model

2026-05-19 · Jinhan Li, Xijie Huang, Zhaoqi Wang, Yijin Wang 외 arxiv

In the field of Vision-Language Navigation (VLN), aerial datasets remain limited in their ability to combine scale, diversity, and realism, often relying on either costly real-world scenes or visually limited simulations…

Vision-Language Navigation

Palmira: A Deep Deformable Network for Instance Segmentation of Dense and Uneven Layouts in Handwritten Manuscripts

2021-08-21 · Prema Satish Sharan, Sowmya Aitha, Amandeep Kumar, Abhishek Trivedi 외

Handwritten documents are often characterized by dense and uneven layout. Despite advances, standard deep network based approaches for semantic layout segmentation are not robust to complex deformations seen across seman…

Instance SegmentationSegmentationSemantic Segmentation

World Tracing: Generative Pixel-Aligned Geometry Beyond the Visible

2026-06-11 · Hao Zhang, Mohamed El Banani, Jen-Hao Cheng, Paul Zhang 외 arxiv

Image-to-3D methods often trade off faithfulness and completeness: depth estimators are anchored to input pixels but stop at the visible surface, while image-to-3D models generate complete shapes that are often misaligne…

3D scene Editing