paper-with-me

홈 › Papers

Align, Adapt and Inject: Sound-guided Unified Image Generation

2023-06-20 · Yue Yang, Kaipeng Zhang, Yuying Ge, Wenqi Shao, Zeyue Xue, Yu Qiao, Ping Luo

Text-guided image generation has witnessed unprecedented progress due to the development of diffusion models. Beyond text and image, sound is a vital element within the sphere of human perception, offering vivid representations and naturally coinciding with corresponding scenes. Taking advantage of sound therefore presents a promising avenue for exploration within image generation research. However, the relationship between audio and image supervision remains significantly underdeveloped, and the scarcity of related, high-quality datasets brings further obstacles. In this paper, we propose a unified framework 'Align, Adapt, and Inject' (AAI) for sound-guided image generation, editing, and stylization. In particular, our method adapts input sound into a sound token, like an ordinary word, which can plug and play with existing powerful diffusion-based Text-to-Image (T2I) models. Specifically, we first train a multi-modal encoder to align audio representation with the pre-trained textual manifold and visual manifold, respectively. Then, we propose the audio adapter to adapt audio representation into an audio token enriched with specific semantics, which can be injected into a frozen T2I model flexibly. In this way, we are able to extract the dynamic information of varied sounds, while utilizing the formidable capability of existing T2I models to facilitate sound-guided image generation, editing, and stylization in a convenient and cost-effective manner. The experiment results confirm that our proposed AAI outperforms other text and sound-guided state-of-the-art methods. And our aligned multi-modal encoder is also competitive with other approaches in the audio-visual retrieval and audio-text retrieval tasks.

📄 PDF Abstract BibTeX arXiv:2306.11504

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationRetrievalText Retrieval

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
Adapter 설명 없음
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Cross-modal Cognitive Consensus guided Audio-Visual Segmentation

2023-10-10 · Zhaofeng Shi, Qingbo Wu, Fanman Meng, Linfeng Xu 외

Audio-Visual Segmentation (AVS) aims to extract the sounding object from a video frame, which is represented by a pixel-wise segmentation mask for application scenarios such as multi-modal video editing, augmented realit…

ObjectSegmentationVideo Editing

Do You Have Freestyle? Expressive Humanoid Locomotion via Audio Control

2025-12-29 · Zhe Li, Cheng Chi, Yangyang Wei, Boan Zhu 외 arxiv

Humans intuitively move to sound, but current humanoid robots lack expressive improvisational capabilities, confined to predefined motions or sparse commands. Generating motion from audio and then retargeting it to robot…

DICT: Data Injection and Contrastive Trajectory Refinement for Conditional Image Generation with Diffusion Models

2026-07-04 · Chunnan Shang, Xin Zhang, Zhizhong Wang, Hongwei Wang arxiv

Diffusion models have become a dominant paradigm for conditional image generation, yet existing approaches generally follow two directions: task-specific designs that can improve performance but limit generalization, and…

Conditional Image GenerationImage Super-ResolutionImage DeblurringStyle Transfer

Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation

2025-10-14 · Xiao He, Huangxuan Zhao, Guojia Wan, Wei Zhou 외 arxiv

Recent medical vision-language models have shown promise on tasks such as VQA, report generation, and anomaly detection. However, most are adapted to structured adult imaging and underperform in fetal ultrasound, which p…

Reinforcement LearningAnomaly Detection

Boosting Ultrasound Image Classification via Attribute-Guided Dual-Branch Framework

2026-07-02 · Bo Zhao, Yapeng Li, Juhua Liu, Bo Du arxiv

Ultrasound image classification is essential for computer-aided diagnosis. However, current methods often neglect clinical priors, leading to poor generalization in challenging scenarios and a lack of interpretability th…

Image Classification