paper-with-me

홈 › Papers

Show and Speak: Directly Synthesize Spoken Description of Images

2020-10-23 · Xinsheng Wang, Siyuan Feng, Jihua Zhu, Mark Hasegawa-Johnson, Odette Scharenborg

This paper proposes a new model, referred to as the show and speak (SAS) model that, for the first time, is able to directly synthesize spoken descriptions of images, bypassing the need for any text or phonemes. The basic structure of SAS is an encoder-decoder architecture that takes an image as input and predicts the spectrogram of speech that describes this image. The final speech audio is obtained from the predicted spectrogram via WaveNet. Extensive experiments on the public benchmark database Flickr8k demonstrate that the proposed SAS is able to synthesize natural spoken descriptions for images, indicating that synthesizing spoken descriptions for images while bypassing text and phonemes is feasible.

📄 PDF Abstract BibTeX arXiv:2010.12267

Code (1)

xinshengwang/Show-and-Speak 공식 구현 pytorch

Tasks

Decoder

Methods 이 논문이 사용한 방법론

Dilated Causal Convolution A Dilated Causal Convolution is a causal convolution where the filter is applied over an area larger than its length by…
Mixture of Logistic Distributions 설명 없음
WaveNet WaveNet is an audio generative model based on the PixelCNN architecture. In order to deal with long-range temporal dependencies…

Similar Papers 제목 키워드 기반

RUSLAN: Russian Spoken Language Corpus for Speech Synthesis

2019-06-26 · Lenar Gabdrakhmanov, Rustem Garaev, Evgenii Razinkov

We present RUSLAN -- a new open Russian spoken language corpus for the text-to-speech task. RUSLAN contains 22200 audio samples with text annotations -- more than 31 hours of high-quality speech of one person -- being th…

Speech Synthesistext-to-speechText to Speech

Investigating the Effects of Large-Scale Pseudo-Stereo Data and Different Speech Foundation Model on Dialogue Generative Spoken Language Model

2024-07-02 · Yu-Kuan Fu, Cheng-Kuang Lee, Hsiu-Hsuan Wang, Hung-Yi Lee

Recent efforts in Spoken Dialogue Modeling aim to synthesize spoken dialogue without the need for direct transcription, thereby preserving the wealth of non-textual information inherent in speech. However, this approach …

Dialogue GenerationDiversityLanguage ModelingLanguage Modelling

Spoken DialogSum: An Emotion-Rich Conversational Dataset for Spoken Dialogue Summarization

2025-12-16 · Yen-Ju Lu, Kunxiao Gao, Mingrui Liang, Helin Wang 외 arxiv

Recent audio language models can follow long conversations. However, research on emotion-aware or spoken dialogue summarization is constrained by the lack of data that links speech, summaries, and paralinguistic cues. We…

When Misinformation Speaks and Converses: Rethinking Fact-Checking in Audio Platforms

2026-04-18 · Chaewan Chun, Delvin Ce Zhang, Dongwon Lee arxiv

Audio platforms have evolved beyond entertainment. They have become central to public discourse, from podcasts and radio to WhatsApp voice notes and live streams. With millions of shows and hundreds of millions of listen…

SpMis: An Investigation of Synthetic Spoken Misinformation Detection

2024-09-17 · Peizhuo Liu, Li Wang, Renqiang He, Haorui He 외

In recent years, speech generation technology has advanced rapidly, fueled by generative models and large-scale training techniques. While these developments have enabled the production of high-quality synthetic speech, …

Misinformationtext-to-speechText to Speech