paper-with-me

Papers

PRISM: Latent Composition Consistency for Single-Image Reflection Removal

2026-06-30 · Junseong Shin, Tae Hyun Kim arxiv

Single-image reflection removal (SIRR) seeks to recover the transmission layer from a mixture corrupted by reflections -- a severely ill-posed problem. Existing methods operate in pixel space, where the nonlinear sRGB formation model entangles the two layers and limits generalization. We observe that pretrained VAE latent spaces exhibit substantially lower coherence between image layers compared to pixel space, providing a more favorable working space for decomposition. Building on this finding, we propose \textbf{PRISM} (Pretrained-latent Reflection Image Separation Model), which reinterprets SIRR as a latent linear separation problem. Under an approximate additive formulation in latent space, PRISM learns a flow matching velocity field on a pretrained FLUX backbone that recovers both transmission and reflection in a single forward pass. To enforce robust disentanglement, we introduce a Latent Composition Consistency (LCC) strategy that constructs synthetic mixtures by swapping reflection latents across samples and enforces consistent decomposition via a cycle loss. We further propose a Layer Contrastive Separation (LCS) loss that promotes semantic separation between layers through patch-level contrastive learning, without requiring explicit reflection targets. Experiments on six benchmarks demonstrate that PRISM consistently outperforms state-of-the-art methods by significant margins, with strong generalization to in-the-wild images.

📄 PDF Abstract BibTeX arXiv:2606.31513

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningReflection Removal

Similar Papers 제목 키워드 기반

PRISM: A Unified Framework for Photorealistic Reconstruction and Intrinsic Scene Modeling

2025-04-19 · Alara Dirik, Tuanfeng Wang, Duygu Ceylan, Stefanos Zafeiriou 외

We present PRISM, a unified framework that enables multiple image generation and editing tasks in a single foundational model. Starting from a pre-trained text-to-image diffusion model, PRISM proposes an effective fine-t…

Conditional Image GenerationImage GenerationIntrinsic Image DecompositionText to Image Generation+1

PRISM: Predictive Recomposition via Semantic Latent Decomposition for View-invariant Video Representation Learning

2026-08-31 · Youngchae Chee, Hosu Lee, Sungjune Park, Junho Kim 외 arxiv

Cross-view video representation learning aims to capture viewpoint-invariant action semantics despite substantial appearance changes across egocentric and exocentric videos. However, existing methods encode each video as…

Representation Learning

Long-Text-to-Image Generation via Compositional Prompt Decomposition

2026-04-20 · Jen-Yuan Huang, Tong Lin, Yilun Du arxiv

While modern text-to-image (T2I) models excel at generating images from intricate prompts, they struggle to capture the key details when the inputs are descriptive paragraphs. This limitation stems from the prevalence of…

Text-to-Image Generation

Compositional Reward Models for Conditional Medical Image Generation

2026-09-04 · Aayush Kumar Tyagi, Prathosh A. P., Mausam arxiv

Acquiring high quality annotated medical image data is critical for training deep learning models; however, annotation is expensive, time consuming, and requires domain expertise. Conditional diffusion models, such as Co…

Skin Lesion ClassificationMedical Image GenerationReinforcement LearningCell Segmentation

Weakly Supervised Learning of Multi-Object 3D Scene Decompositions Using Deep Shape Priors

2020-10-08 · Cathrin Elich, Martin R. Oswald, Marc Pollefeys, Joerg Stueckler

Representing scenes at the granularity of objects is a prerequisite for scene understanding and decision making. We propose PriSMONet, a novel approach based on Prior Shape knowledge for learning Multi-Object 3D scene de…

Decision MakingScene UnderstandingWeakly-supervised Learning