paper-with-me

Papers

PRISM: A Unified Framework for Photorealistic Reconstruction and Intrinsic Scene Modeling

2025-04-19 · Alara Dirik, Tuanfeng Wang, Duygu Ceylan, Stefanos Zafeiriou, Anna Frühstück

We present PRISM, a unified framework that enables multiple image generation and editing tasks in a single foundational model. Starting from a pre-trained text-to-image diffusion model, PRISM proposes an effective fine-tuning strategy to produce RGB images along with intrinsic maps (referred to as X layers) simultaneously. Unlike previous approaches, which infer intrinsic properties individually or require separate models for decomposition and conditional generation, PRISM maintains consistency across modalities by generating all intrinsic layers jointly. It supports diverse tasks, including text-to-RGBX generation, RGB-to-X decomposition, and X-to-RGBX conditional generation. Additionally, PRISM enables both global and local image editing through conditioning on selected intrinsic layers and text prompts. Extensive experiments demonstrate the competitive performance of PRISM both for intrinsic image decomposition and conditional image generation while preserving the base model's text-to-image generation capability.

📄 PDF Abstract BibTeX arXiv:2504.14219

Code (0)

등록된 구현이 없습니다.

Tasks

Conditional Image GenerationImage GenerationIntrinsic Image DecompositionText to Image GenerationText-to-Image Generation

Methods 이 논문이 사용한 방법론

BASE 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

PRISM: Rethinking Atmospheric Scattering Reconstruction as a Unified Understanding and Restoration Model for Real-world Dehazing

2026-04-08 · Chengyu Fang, Chunming He, Yuelin Zhang, Chubin Chen 외 arxiv

Real-world image dehazing (RID) aims to remove haze-induced degradation from real scenes. This task remains challenging due to non-uniform haze distribution, spatially varying color shifts, and the scarcity of paired rea…

Image Dehazing

Real-Time Human Frontal View Synthesis from a Single Image

2026-03-16 · Fangyu Lin, Yingdong Hu, Lunjie Zhu, Zhening Liu 외 arxiv

Photorealistic human novel view synthesis from a single image is crucial for democratizing immersive 3D telepresence, eliminating the need for complex multi-camera setups. However, current rendering-centric methods prior…

Novel View SynthesisPoint Clouds

Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video

2025-04-28 · Sonia Joseph, Praneet Suresh, Lorenz Hufe, Edward Stevinson 외

Robust tooling and publicly available pre-trained models have helped drive recent advances in mechanistic interpretability for language models. However, similar progress in vision mechanistic interpretability has been hi…

PrismGS: Physically-Grounded Anti-Aliasing for High-Fidelity Large-Scale 3D Gaussian Splatting

2025-10-09 · Houqiang Zhong, Zhenglong Wu, Sihua Fu, Zihan Zheng 외 arxiv

3D Gaussian Splatting (3DGS) has recently enabled real-time photorealistic rendering in compact scenes, but scaling to large urban environments introduces severe aliasing artifacts and optimization instability, especiall…

PRISM3D: Probabilistic Refinement and Robust Initialization for Physically Consistent Scene Modeling under Extreme Motion Blur

2026-07-04 · Gopi Raju Matta, Reddypalli Trisha, Vemunuri Divya Madhuri, Kaushik Mitra arxiv

We address the inverse problem of blind 3D scene reconstruction from extremely motion-blurred images, a scenario where traditional Structure-from-Motion (SfM) pipelines fail. Existing approaches typically circumvent this…