paper-with-me

홈 › Papers

Seeing What You Say: Expressive Image Generation from Speech

2025-11-05 · Jiyoung Lee, Song Park, Sanghyuk Chun, Soo-Whan Chung arxiv

This paper proposes VoxStudio, the first unified and end-to-end speech-to-image model that generates expressive images directly from spoken descriptions by jointly aligning linguistic and paralinguistic information. At its core is a speech information bottleneck (SIB) module, which compresses raw speech into compact semantic tokens, preserving prosody and emotional nuance. By operating directly on these tokens, VoxStudio eliminates the need for an additional speech-to-text system, which often ignores the hidden details beyond text, e.g., tone or emotion. We also release VoxEmoset, a large-scale paired emotional speech-image dataset built via an advanced TTS engine to affordably generate richly expressive utterances. Comprehensive experiments on the SpokenCOCO, Flickr8kAudio, and VoxEmoset benchmarks demonstrate the feasibility of our method and highlight key challenges, including emotional consistency and linguistic ambiguity, paving the way for future research.

📄 PDF Abstract BibTeX arXiv:2511.03423

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generation

Similar Papers 제목 키워드 기반

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation

2025-08-22 · Weiting Tan, Jiachen Lian, Hirofumi Inaguma, Paden Tomasello 외 arxiv

We present an Audio-Visual Language Model (AVLM) for expressive speech generation by integrating full-face visual cues into a pre-trained expressive speech model. We explore multiple visual encoders and multimodal fusion…

Emotion Recognition

Seeing Voices: Generating A-Roll Video from Audio with Mirage

2025-06-09 · Aditi Sundararaman, Amogh Adishesha, Andrew Jaegle, Dan Bigioi 외

From professional filmmaking to user-generated content, creators and consumers have long recognized that the power of video depends on the harmonious integration of what we hear (the video's audio track) with what we see…

Speech Synthesistext-to-speechText to SpeechVideo Generation

Bridging What the Model Thinks and How It Speaks: Expressive Speech Generation via Self-Aware Intent-Realization Alignment

2026-04-13 · Kuang Wang, Lai Wei, Ping Lin, Qibing Bai 외 arxiv

Speech Language Models (SLMs) exhibit strong semantic understanding, yet often fail to translate this capacity into expressive acoustic realization, producing speech with flattened prosody and misaligned emotion. We iden…

Seeing What You Said: Talking Face Generation Guided by a Lip Reading Expert

2023-03-29 · CVPR 2023 1 · Jiadong Wang, Xinyuan Qian, Malu Zhang, Robby T. Tan 외

Talking face generation, also known as speech-to-lip generation, reconstructs facial motions concerning lips given coherent speech input. The previous studies revealed the importance of lip-speech synchronization and vis…

Contrastive LearningFace GenerationLip ReadingTalking Face Generation

Facial Expression-Enhanced TTS: Combining Face Representation and Emotion Intensity for Adaptive Speech

2024-09-24 · Yunji Chu, Yunseob Shim, Unsang Park

We propose FEIM-TTS, an innovative zero-shot text-to-speech (TTS) model that synthesizes emotionally expressive speech, aligned with facial images and modulated by emotion intensity. Leveraging deep learning, FEIM-TTS tr…

Emotional Speech SynthesisSpeech Synthesistext-to-speechText to Speech