paper-with-me

홈 › Papers

SaSLaW: Dialogue Speech Corpus with Audio-visual Egocentric Information Toward Environment-adaptive Dialogue Speech Synthesis

2024-08-13 · Osamu Take, Shinnosuke Takamichi, Kentaro Seki, Yoshiaki Bando, Hiroshi Saruwatari

This paper presents SaSLaW, a spontaneous dialogue speech corpus containing synchronous recordings of what speakers speak, listen to, and watch. Humans consider the diverse environmental factors and then control the features of their utterances in face-to-face voice communications. Spoken dialogue systems capable of this adaptation to these audio environments enable natural and seamless communications. SaSLaW was developed to model human-speech adjustment for audio environments via first-person audio-visual perceptions in spontaneous dialogues. We propose the construction methodology of SaSLaW and display the analysis result of the corpus. We additionally conducted an experiment to develop text-to-speech models using SaSLaW and evaluate their performance of adaptations to audio environments. The results indicate that models incorporating hearing-audio data output more plausible speech tailored to diverse audio environments than the vanilla text-to-speech model.

📄 PDF Abstract BibTeX arXiv:2408.06858

Code (1)

sarulab-speech/saslaw 공식 구현 pytorch

Tasks

Speech SynthesisSpoken Dialogue Systemstext-to-speechText to Speech

Similar Papers 제목 키워드 기반

Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation

2024-06-12 · Se Jin Park, Chae Won Kim, Hyeongseop Rha, Minsu Kim 외

In this paper, we introduce a novel Face-to-Face spoken dialogue model. It processes audio-visual speech from user input and generates audio-visual speech as the response, marking the initial step towards creating an ava…

ChatbotLanguage ModelingLanguage ModellingLarge Language Model

A Multimodal Corpus of Rapid Dialogue Games

2014-05-01 · LREC 2014 5 · Maike Paetzel, David Nicolas Racca, David DeVault

This paper presents a multimodal corpus of spoken human-human dialogues collected as participants played a series of Rapid Dialogue Games (RDGs). The corpus consists of a collection of about 11 hours of spoken audio, vid…

Dialogue ManagementManagementNatural Language UnderstandingQuestion Answering+4

Audiobook Dialogues as Training Data for Conversational Style Synthetic Voices

2022-06-01 · LREC 2022 6 · Liisi Piits, Hille Pajupuu, Heete Sahkai, Rene Altrov 외

Synthetic voices are increasingly used in applications that require a conversational speaking style, raising the question as to which type of training data yields the most suitable speaking style for such applications. T…

Sentencetext-to-speechText to Speech

Candor-LR: A Dyadic Conversational Dataset for Audio-Visual Speech Recognition

2026-09-09 · Rishabh Jain, Aristeidis Papadopoulos, Zhaofeng Lin, Naomi Harte arxiv

Current audio-visual speech recognition (AVSR) benchmarks, like LRS3, rely heavily on clean, scripted and rehearsed speech. They fail to reflect the complexity of natural conversation, which involves overlapping speech, …

Audio-Visual Speech Recognition

MuAViC: A Multilingual Audio-Visual Corpus for Robust Speech Recognition and Robust Speech-to-Text Translation

2023-03-01 · Mohamed Anwar, Bowen Shi, Vedanuj Goswami, Wei-Ning Hsu 외

We introduce MuAViC, a multilingual audio-visual corpus for robust speech recognition and robust speech-to-text translation providing 1200 hours of audio-visual speech in 9 languages. It is fully transcribed and covers 6…

Audio-Visual Speech RecognitionRobust Speech Recognitionspeech-recognitionSpeech Recognition+4