paper-with-me

홈 › Papers

SPAD: Spatially Aware Multi-View Diffusers

2024-01-01 · CVPR 2024 1 · Yash Kant, Aliaksandr Siarohin, Ziyi Wu, Michael Vasilkovsky, Guocheng Qian, Jian Ren, Riza Alp Guler, Bernard Ghanem, Sergey Tulyakov, Igor Gilitschenski

We present SPAD a novel approach for creating consistent multi-view images from text prompts or single images. To enable multi-view generation we repurpose a pretrained 2D diffusion model by extending its self-attention layers with cross-view interactions and fine-tune it on a high quality subset of Objaverse. We find that a naive extension of the self-attention proposed in prior work (e.g. MVDream) leads to content copying between views. Therefore we explicitly constrain the cross-view attention based on epipolar geometry. To further enhance 3D consistency we utilize Pl ?ucker coordinates derived from camera rays and inject them as positional encoding. This enables SPAD to reason over spatial proximity in 3D well. Compared to concurrent works that can only generate views at fixed azimuth and elevation (e.g. MVDream SyncDreamer) SPAD offers full camera control and achieves state-of-the-art results in novel view synthesis on unseen objects from the Objaverse and Google Scanned Objects datasets. Finally we demonstrate that text-to-3D generation using SPAD prevents the multi-face Janus issue.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

3D GenerationNovel View SynthesisText to 3D

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

SPAD : Spatially Aware Multiview Diffusers

2024-02-07 · Yash Kant, Ziyi Wu, Michael Vasilkovsky, Guocheng Qian 외

We present SPAD, a novel approach for creating consistent multi-view images from text prompts or single images. To enable multi-view generation, we repurpose a pretrained 2D diffusion model by extending its self-attentio…

3D GenerationNovel View SynthesisText to 3D

Rethinking Spatially-Adaptive Normalization

2020-04-06 · Zhentao Tan, Dongdong Chen, Qi Chu, Menglei Chai 외

Spatially-adaptive normalization is remarkably successful recently in conditional semantic image synthesis, which modulates the normalized activation with spatially-varying transformations learned from semantic layouts, …

Image Generation

Efficient Semantic Image Synthesis via Class-Adaptive Normalization

2020-12-08 · Zhentao Tan, Dongdong Chen, Qi Chu, Menglei Chai 외

Spatially-adaptive normalization (SPADE) is remarkably successful recently in conditional semantic image synthesis \cite{park2019semantic}, which modulates the normalized activation with spatially-varying transformations…

Image Generation

ESPADA: Execution Speedup via Semantics Aware Demonstration Data Downsampling for Imitation Learning

2025-12-08 · Byung-ju Kim, Jinu Pahk, Chungwoo Lee, Jaejoon Kim 외 arxiv

Behavior-cloning based visuomotor policies enable precise manipulation but often inherit the slow, cautious tempo of human demonstrations, limiting practical deployment. However, prior studies on acceleration methods mai…

Leveraging generative adversarial networks with spatially adaptive denormalization for multivariate stochastic seismic data inversion

2025-12-02 · Roberto Miele, Leonardo Azevedo arxiv

Probabilistic seismic inverse modeling often requires the prediction of both spatially correlated geological heterogeneities (e.g., facies) and continuous parameters (e.g., rock and elastic properties). Generative advers…