paper-with-me

홈 › Papers

Automatic Prosody Annotation with Pre-Trained Text-Speech Model

2022-06-16 · Ziqian Dai, Jianwei Yu, Yan Wang, Nuo Chen, Yanyao Bian, Guangzhi Li, Deng Cai, Dong Yu

Prosodic boundary plays an important role in text-to-speech synthesis (TTS) in terms of naturalness and readability. However, the acquisition of prosodic boundary labels relies on manual annotation, which is costly and time-consuming. In this paper, we propose to automatically extract prosodic boundary labels from text-audio data via a neural text-speech model with pre-trained audio encoders. This model is pre-trained on text and speech data separately and jointly fine-tuned on TTS data in a triplet format: {speech, text, prosody}. The experimental results on both automatic evaluation and human evaluation demonstrate that: 1) the proposed text-speech prosody annotation framework significantly outperforms text-only baselines; 2) the quality of automatic prosodic boundary annotations is comparable to human annotations; 3) TTS systems trained with model-annotated boundaries are slightly better than systems that use manual ones.

📄 PDF Abstract BibTeX arXiv:2206.07956

Code (1)

daisyqk/automatic-prosody-annotation 공식 구현 pytorch

Tasks

Speech Synthesistext-to-speechText to SpeechText-To-Speech SynthesisTriplet

Similar Papers 제목 키워드 기반

Multi-Modal Automatic Prosody Annotation with Contrastive Pretraining of SSWP

2023-09-11 · Jinzuomu Zhong, Yang Li, Hui Huang, Korin Richmond 외

In expressive and controllable Text-to-Speech (TTS), explicit prosodic features significantly improve the naturalness and controllability of synthesised speech. However, manual prosody annotation is labor-intensive and i…

text-to-speechText to Speech

An Automatic Prosody Tagger for Spontaneous Speech

2016-12-01 · COLING 2016 12 · M{\'o}nica Dom{\'\i}nguez, Mireia Farr{\'u}s, Leo Wanner

Speech prosody is known to be central in advanced communication technologies. However, despite the advances of theoretical studies in speech prosody, so far, no large scale prosody annotated resources that would facilita…

Descriptive

Voice Conversion by Cascading Automatic Speech Recognition and Text-to-Speech Synthesis with Prosody Transfer

2020-09-03 · Jing-Xuan Zhang, Li-Juan Liu, Yan-Nian Chen, Ya-Jun Hu 외

With the development of automatic speech recognition (ASR) and text-to-speech synthesis (TTS) technique, it's intuitive to construct a voice conversion system by cascading an ASR and TTS system. In this paper, we present…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+5

Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech

2025-07-17 · Kirill Borodin, Nikita Vasiliev, Vasiliy Kudryavtsev, Maxim Maslov 외 arxiv

We introduce Balalaika, an open-source, data-centric pipeline for processing audio and producing prosody-aware annotations. It combines semantic VAD for context-preserving segmentation, multi-ASR ensembling with ROVER co…

Speech Denoising

Prominence-aware automatic speech recognition for conversational speech

2025-09-12 · Julian Linke, Barbara Schuppler arxiv

This paper investigates prominence-aware automatic speech recognition (ASR) by combining prominence detection and speech recognition for conversational Austrian German. First, prominence detectors were developed by fine-…

Speech Recognition