paper-with-me

Papers

Text2Stereo: Repurposing Stable Diffusion for Stereo Generation with Consistency Rewards

2025-05-27 · Aakash Garg, Libing Zeng, Andrii Tsarov, Nima Khademi Kalantari

In this paper, we propose a novel diffusion-based approach to generate stereo images given a text prompt. Since stereo image datasets with large baselines are scarce, training a diffusion model from scratch is not feasible. Therefore, we propose leveraging the strong priors learned by Stable Diffusion and fine-tuning it on stereo image datasets to adapt it to the task of stereo generation. To improve stereo consistency and text-to-image alignment, we further tune the model using prompt alignment and our proposed stereo consistency reward functions. Comprehensive experiments demonstrate the superiority of our approach in generating high-quality stereo images across diverse scenarios, outperforming existing methods.

📄 PDF Abstract BibTeX arXiv:2506.05367

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

'Person' == Light-skinned, Western Man, and Sexualization of Women of Color: Stereotypes in Stable Diffusion

2023-10-30 · Sourojit Ghosh, Aylin Caliskan

We study stereotypes embedded within one of the most popular text-to-image generators: Stable Diffusion. We examine what stereotypes of gender and nationality/continental identity does Stable Diffusion display in the abs…

Language Modelling

Fast Timing-Conditioned Latent Audio Diffusion

2024-02-07 · Zach Evans, CJ Carr, Josiah Taylor, Scott H. Hawley 외

Generating long-form 44.1kHz stereo audio from text prompts can be computationally demanding. Further, most previous works do not tackle that music and sound effects naturally vary in their duration. Our research focuses…

Audio GenerationGPUText-to-Music Generation

RS-Corrector: Correcting the Racial Stereotypes in Latent Diffusion Models

2023-12-08 · Yue Jiang, Yueming Lyu, Tianxiang Ma, Bo Peng 외

Recent text-conditioned image generation models have demonstrated an exceptional capacity to produce diverse and creative imagery with high visual quality. However, when pre-trained on billion-sized datasets randomly col…

Image Generation

StereoDiffusion: Training-Free Stereo Image Generation Using Latent Diffusion Models

2024-03-08 · Lezhong Wang, Jeppe Revall Frisvad, Mark Bo Jensen, Siavash Arjomand Bigdeli

The demand for stereo images increases as manufacturers launch more XR devices. To meet this demand, we introduce StereoDiffusion, a method that, unlike traditional inpainting pipelines, is trainning free, remarkably str…

Image Generation

AI-generated faces influence gender stereotypes and racial homogenization

2024-02-01 · Nouar AlDahoul, Talal Rahwan, Yasir Zaki

Text-to-image generative AI models such as Stable Diffusion are used daily by millions worldwide. However, the extent to which these models exhibit racial and gender stereotypes is not yet fully understood. Here, we docu…