paper-with-me

Papers

CoPrUS: Consistency Preserving Utterance Synthesis towards more realistic benchmark dialogues

2024-12-10 · Sebastian Steindl, Ulrich Schäfer, Bernd Ludwig

Large-scale Wizard-Of-Oz dialogue datasets have enabled the training of deep learning-based dialogue systems. While they are successful as benchmark datasets, they lack certain types of utterances, which would make them more realistic. In this work, we investigate the creation of synthetic communication errors in an automatic pipeline. Based on linguistic theory, we propose and follow a simple error taxonomy. We focus on three types of miscommunications that could happen in real-world dialogues but are underrepresented in the benchmark dataset: misunderstandings, non-understandings and vaguely related questions. Our two-step approach uses a state-of-the-art Large Language Model (LLM) to first create the error and secondly the repairing utterance. We perform Language Model-based evaluation to ensure the quality of the generated utterances. We apply the method to the MultiWOZ dataset and evaluate it both qualitatively and empirically as well as with human judges. Our results indicate that current LLMs can aid in adding post-hoc miscommunications to benchmark datasets as a form of data augmentation. We publish the resulting dataset, in which nearly 1900 dialogues have been modified, as CoPrUS-MultiWOZ to facilitate future work on dialogue systems.

📄 PDF Abstract BibTeX arXiv:2412.07515

Code (1)

sebastian-steindl/CoPrUS_data 공식 구현

Tasks

Data AugmentationLanguage ModelingLanguage ModellingLarge Language Model

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Location-Relative Attention Mechanisms For Robust Long-Form Speech Synthesis

2019-10-23 · Eric Battenberg, RJ Skerry-Ryan, Soroosh Mariooryad, Daisy Stanton 외

Despite the ability to produce human-level speech for in-domain text, attention-based end-to-end text-to-speech (TTS) systems suffer from text alignment failures that increase in frequency for out-of-domain text. We show…

FormSpeech Synthesistext-to-speechText to Speech

Zero-Shot Long-Form Voice Cloning with Dynamic Convolution Attention

2022-01-25 · Artem Gorodetskii, Ivan Ozhiganov

With recent advancements in voice cloning, the performance of speech synthesis for a target speaker has been rendered similar to the human level. However, autoregressive voice cloning systems still suffer from text align…

FormSpeech Synthesistext-to-speechText to Speech+1

Do Prosody Transfer Models Transfer Prosody?

2023-03-07 · Atli Thor Sigurgeirsson, Simon King

Some recent models for Text-to-Speech synthesis aim to transfer the prosody of a reference utterance to the generated target synthetic speech. This is done by using a learned embedding of the reference utterance, which i…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Building a mixed-lingual neural TTS system with only monolingual data

2019-04-12 · Liumeng Xue, Wei Song, Guanghui Xu, Lei Xie 외

When deploying a Chinese neural text-to-speech (TTS) synthesis system, one of the challenges is to synthesize Chinese utterances with English phrases or words embedded. This paper looks into the problem in the encoder-de…

Decodertext-to-speechText to Speech

Generative Moment Matching Network-based Random Modulation Post-filter for DNN-based Singing Voice Synthesis and Neural Double-tracking

2019-02-09 · Hiroki Tamaru, Yuki Saito, Shinnosuke Takamichi, Tomoki Koriyama 외

This paper proposes a generative moment matching network (GMMN)-based post-filter that provides inter-utterance pitch variation for deep neural network (DNN)-based singing voice synthesis. The natural pitch variation of …

Singing Voice Synthesis