paper-with-me

홈 › Papers

Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech

2025-07-17 · Kirill Borodin, Nikita Vasiliev, Vasiliy Kudryavtsev, Maxim Maslov, Mikhail Gorodnichev, Grach Mkrtchian arxiv

We introduce Balalaika, an open-source, data-centric pipeline for processing audio and producing prosody-aware annotations. It combines semantic VAD for context-preserving segmentation, multi-ASR ensembling with ROVER consensus decoding, while retaining optional word-level timestamps, followed by automatic quality and speaker-purity filtering. The text is further enriched with punctuation restoration, lexical stress and "\textipa{e}/\textipa{He}" normalization, and IPA phonemes. Using Balalaika, we build a 5.1k-hour multi-source Russian corpus with rich annotations, and show consistent gains under equalized training budgets for both speech denoising and TTS; ablations confirm complementary benefits of stress and punctuation and improved synthesis with stricter MOS filtering. The datasets are publicly available at \href{https://huggingface.co/collections/lab260/balalaika-dataset}{\underline{\textbf{HuggingFace}}}

📄 PDF Abstract BibTeX arXiv:2507.13563

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Denoising

Similar Papers 제목 키워드 기반

An Automatic Prosody Tagger for Spontaneous Speech

2016-12-01 · COLING 2016 12 · M{\'o}nica Dom{\'\i}nguez, Mireia Farr{\'u}s, Leo Wanner

Speech prosody is known to be central in advanced communication technologies. However, despite the advances of theoretical studies in speech prosody, so far, no large scale prosody annotated resources that would facilita…

Descriptive

Prominence-aware automatic speech recognition for conversational speech

2025-09-12 · Julian Linke, Barbara Schuppler arxiv

This paper investigates prominence-aware automatic speech recognition (ASR) by combining prominence detection and speech recognition for conversational Austrian German. First, prominence detectors were developed by fine-…

Speech Recognition

Automatic Prosody Annotation with Pre-Trained Text-Speech Model

2022-06-16 · Ziqian Dai, Jianwei Yu, Yan Wang, Nuo Chen 외

Prosodic boundary plays an important role in text-to-speech synthesis (TTS) in terms of naturalness and readability. However, the acquisition of prosodic boundary labels relies on manual annotation, which is costly and t…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis+1

PAVITS: Exploring Prosody-aware VITS for End-to-End Emotional Voice Conversion

2024-03-03 · Tianhua Qi, Wenming Zheng, Cheng Lu, Yuan Zong 외

In this paper, we propose Prosody-aware VITS (PAVITS) for emotional voice conversion (EVC), aiming to achieve two major objectives of EVC: high content naturalness and high emotional naturalness, which are crucial for me…

Voice Conversion

Multi-Modal Automatic Prosody Annotation with Contrastive Pretraining of SSWP

2023-09-11 · Jinzuomu Zhong, Yang Li, Hui Huang, Korin Richmond 외

In expressive and controllable Text-to-Speech (TTS), explicit prosodic features significantly improve the naturalness and controllability of synthesised speech. However, manual prosody annotation is labor-intensive and i…

text-to-speechText to Speech