paper-with-me

홈 › Papers

VMDiff: Visual Mixing Diffusion for Limitless Cross-Object Synthesis

2025-09-28 · Zeren Xiong, Yue Yu, Zedong Zhang, Shuo Chen, Jian Yang, Jun Li arxiv

Creating novel images by fusing visual cues from multiple sources is a fundamental yet underexplored problem in image-to-image generation, with broad applications in artistic creation, virtual reality and visual media. Existing methods often face two key challenges: coexistent generation, where multiple objects are simply juxtaposed without true integration, and bias generation, where one object dominates the output due to semantic imbalance. To address these issues, we propose Visual Mixing Diffusion (VMDiff), a simple yet effective diffusion-based framework that synthesizes a single, coherent object by integrating two input images at both noise and latent levels. Our approach comprises: (1) a hybrid sampling process that combines guided denoising, inversion, and spherical interpolation with adjustable parameters to achieve structure-aware fusion, mitigating coexistent generation; and (2) an efficient adaptive adjustment module, which introduces a novel similarity-based score to automatically and adaptively search for optimal parameters, countering semantic bias. Experiments on a curated benchmark of 780 concept pairs demonstrate that our method outperforms strong baselines in visual quality, semantic consistency, and human-rated creativity.

📄 PDF Abstract BibTeX arXiv:2509.23605

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generation

Similar Papers 제목 키워드 기반

Why does music source separation benefit from cacophony?

2024-02-28 · Chang-Bin Jeon, Gordon Wichern, François G. Germain, Jonathan Le Roux

In music source separation, a standard training data augmentation procedure is to create new training samples by randomly combining instrument stems from different songs. These random mixes have mismatched characteristic…

Data AugmentationMusic Source Separation

VMix: Improving Text-to-Image Diffusion Model with Cross-Attention Mixing Control

2024-12-30 · Shaojin Wu, Fei Ding, Mengqi Huang, Wei Liu 외

While diffusion models show extraordinary talents in text-to-image generation, they may still fail to generate highly aesthetic images. More specifically, there is still a gap between the generated images and the real-wo…

DenoisingImage GenerationText to Image GenerationText-to-Image Generation

World in a Frame: Understanding Culture Mixing as a New Challenge for Vision-Language Models

2025-11-27 · Eunsu Kim, Junyeong Park, Na Min An, Junseong Kim 외 arxiv

In a globalized world, cultural elements from diverse origins frequently appear together within a single visual scene. We refer to these as culture mixing scenarios, yet how Large Vision-Language Models (LVLMs) perceive …

Visual Question Answering

Dataset Augmentation by Mixing Visual Concepts

2024-12-19 · Abdullah Al Rahat, Hemanth Venkateswara

This paper proposes a dataset augmentation method by fine-tuning pre-trained diffusion models. Generating images using a pre-trained diffusion model with textual conditioning often results in domain discrepancy between r…

Image Captioning

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects

2024-11-28 · CVPR 2025 1 · Weimin Qiu, Jieke Wang, Meng Tang

Diffusion models have achieved unprecedented fidelity and diversity for synthesizing image, video, 3D assets, etc. However, subject mixing is a known and unresolved issue for diffusion-based image synthesis, particularly…

Image Generation