paper-with-me

홈 › Papers

PoGDiff: Product-of-Gaussians Diffusion Models for Imbalanced Text-to-Image Generation

2025-02-12 · Ziyan Wang, Sizhe Wei, Xiaoming Huo, Hao Wang

Diffusion models have made significant advancements in recent years. However, their performance often deteriorates when trained or fine-tuned on imbalanced datasets. This degradation is largely due to the disproportionate representation of majority and minority data in image-text pairs. In this paper, we propose a general fine-tuning approach, dubbed PoGDiff, to address this challenge. Rather than directly minimizing the KL divergence between the predicted and ground-truth distributions, PoGDiff replaces the ground-truth distribution with a Product of Gaussians (PoG), which is constructed by combining the original ground-truth targets with the predicted distribution conditioned on a neighboring text embedding. Experiments on real-world datasets demonstrate that our method effectively addresses the imbalance problem in diffusion models, improving both generation accuracy and quality.

📄 PDF Abstract BibTeX arXiv:2502.08106

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationText to Image GenerationText-to-Image Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

GaussianBlender: Instant Stylization of 3D Gaussians with Disentangled Latent Spaces

2025-12-03 · Melis Ocal, Xiaoyan Xing, Yue Li, Ngo Anh Vien 외 arxiv

3D stylization is central to game development, virtual reality, and digital arts, where the demand for diverse assets calls for scalable methods that support fast, high-fidelity manipulation. Existing text-to-3D stylizat…

GaussianEditor: Editing 3D Gaussians Delicately with Text Instructions

2023-11-27 · CVPR 2024 1 · Junjie Wang, Jiemin Fang, Xiaopeng Zhang, Lingxi Xie 외

Recently, impressive results have been achieved in 3D scene editing with text instructions based on a 2D diffusion model. However, current diffusion models primarily generate images by predicting noise in the latent spac…

3D scene EditingGPU

Align Your Gaussians: Text-to-4D with Dynamic 3D Gaussians and Composed Diffusion Models

2023-12-21 · CVPR 2024 1 · Huan Ling, Seung Wook Kim, Antonio Torralba, Sanja Fidler 외

Text-guided diffusion models have revolutionized image and video generation and have also been successfully used for optimization-based 3D object synthesis. Here, we instead focus on the underexplored text-to-4D setting …

Synthetic Data GenerationVideo Generation

Atlas Gaussians Diffusion for 3D Generation

2024-08-23 · Haitao Yang, Yuan Dong, Hanwen Jiang, Dejia Xu 외

Using the latent diffusion model has proven effective in developing novel 3D generation techniques. To harness the latent diffusion model, a key challenge is designing a high-fidelity and efficient representation that li…

3D Generation

ScalingGaussian: Enhancing 3D Content Creation with Generative Gaussian Splatting

2024-07-26 · Shen Chen, Jiale Zhou, Zhongyu Jiang, Tianfang Zhang 외

The creation of high-quality 3D assets is paramount for applications in digital heritage preservation, entertainment, and robotics. Traditionally, this process necessitates skilled professionals and specialized software …

Image to 3D