paper-with-me

홈 › Papers

Aligned Novel View Image and Geometry Synthesis via Cross-modal Attention Instillation

2025-06-13 · Min-Seop Kwak, Junho Kim, Sangdoo Yun, Dongyoon Han, Taekyoung Kim, Seungryong Kim, Jin-Hwa Kim

We introduce a diffusion-based framework that performs aligned novel view image and geometry generation via a warping-and-inpainting methodology. Unlike prior methods that require dense posed images or pose-embedded generative models limited to in-domain views, our method leverages off-the-shelf geometry predictors to predict partial geometries viewed from reference images, and formulates novel-view synthesis as an inpainting task for both image and geometry. To ensure accurate alignment between generated images and geometry, we propose cross-modal attention distillation, where attention maps from the image diffusion branch are injected into a parallel geometry diffusion branch during both training and inference. This multi-task approach achieves synergistic effects, facilitating geometrically robust image synthesis as well as well-defined geometry prediction. We further introduce proximity-based mesh conditioning to integrate depth and normal cues, interpolating between point cloud and filtering erroneously predicted geometry from influencing the generation process. Empirically, our method achieves high-fidelity extrapolative view synthesis on both image and geometry across a range of unseen scenes, delivers competitive reconstruction quality under interpolation settings, and produces geometrically aligned colored point clouds for comprehensive 3D completion. Project page is available at https://cvlab-kaist.github.io/MoAI.

📄 PDF Abstract BibTeX arXiv:2506.11924

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationNovel View Synthesis

Methods 이 논문이 사용한 방법론

Inpainting Train a convolutional neural network to generate the contents of an arbitrary image region conditioned on its surroundings.
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

PVSeRF: Joint Pixel-, Voxel- and Surface-Aligned Radiance Field for Single-Image Novel View Synthesis

2022-02-10 · Xianggang Yu, Jiapeng Tang, Yipeng Qin, Chenghong Li 외

We present PVSeRF, a learning framework that reconstructs neural radiance fields from single-view RGB images, for novel view synthesis. Previous solutions, such as pixelNeRF, rely only on pixel-aligned features and suffe…

DisentanglementNovel View Synthesis

GeoQuery: Geometry-Query Diffusion for Sparse-View Reconstruction

2026-05-12 · Xiao Cao, Yuze Li, Youmin Zhang, Jiayu Song 외 arxiv

3D Gaussian Splatting (3DGS) has emerged as a prominent paradigm for 3D reconstruction and novel view synthesis. However, it remains vulnerable to severe artifacts when trained under sparse-view constraints. While recent…

Novel View Synthesis3D Reconstruction

Geometry-Aware Reference Synthesis for Multi-View Image Super-Resolution

2022-07-18 · Ri Cheng, Yuqi Sun, Bo Yan, Weimin Tan 외

Recent multi-view multimedia applications struggle between high-resolution (HR) visual experience and storage or bandwidth constraints. Therefore, this paper proposes a Multi-View Image Super-Resolution (MVISR) task. It …

Image Super-ResolutionSuper-ResolutionVideo Super-Resolution

AlignMiF: Geometry-Aligned Multimodal Implicit Field for LiDAR-Camera Joint Synthesis

2024-02-27 · CVPR 2024 1 · Tao Tang, Guangrun Wang, Yixing Lao, Peng Chen 외

Neural implicit fields have been a de facto standard in novel view synthesis. Recently, there exist some methods exploring fusing multiple modalities within a single field, aiming to share implicit features from differen…

Novel View Synthesis

TexTailor: Customized Text-aligned Texturing via Effective Resampling

2025-06-12 · SuIn Lee, Dae-shik Kim

We present TexTailor, a novel method for generating consistent object textures from textual descriptions. Existing text-to-texture synthesis approaches utilize depth-aware diffusion models to progressively generate image…

Texture Synthesis