Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech
We introduce Balalaika, an open-source, data-centric pipeline for processing audio and producing prosody-aware annotations. It combines semantic VAD for context-preserving segmentation, multi-ASR ensembling with ROVER consensus decoding, while retaining optional word-level timestamps, followed by automatic quality and speaker-purity filtering. The text is further enriched with punctuation restoration, lexical stress and "\textipa{e}/\textipa{He}" normalization, and IPA phonemes. Using Balalaika, we build a 5.1k-hour multi-source Russian corpus with rich annotations, and show consistent gains under equalized training budgets for both speech denoising and TTS; ablations confirm complementary benefits of stress and punctuation and improved synthesis with stricter MOS filtering. The datasets are publicly available at \href{https://huggingface.co/collections/lab260/balalaika-dataset}{\underline{\textbf{HuggingFace}}}
Code (0)
등록된 구현이 없습니다.
Tasks
Speech DenoisingSimilar Papers 제목 키워드 기반
An Automatic Prosody Tagger for Spontaneous Speech
Speech prosody is known to be central in advanced communication technologies. However, despite the advances of theoretical studies in speech prosody, so far, no large scale prosody annotated resources that would facilita…
DescriptiveProminence-aware automatic speech recognition for conversational speech
This paper investigates prominence-aware automatic speech recognition (ASR) by combining prominence detection and speech recognition for conversational Austrian German. First, prominence detectors were developed by fine-…
Speech RecognitionAutomatic Prosody Annotation with Pre-Trained Text-Speech Model
Prosodic boundary plays an important role in text-to-speech synthesis (TTS) in terms of naturalness and readability. However, the acquisition of prosodic boundary labels relies on manual annotation, which is costly and t…
Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis+1PAVITS: Exploring Prosody-aware VITS for End-to-End Emotional Voice Conversion
In this paper, we propose Prosody-aware VITS (PAVITS) for emotional voice conversion (EVC), aiming to achieve two major objectives of EVC: high content naturalness and high emotional naturalness, which are crucial for me…
Voice ConversionMulti-Modal Automatic Prosody Annotation with Contrastive Pretraining of SSWP
In expressive and controllable Text-to-Speech (TTS), explicit prosodic features significantly improve the naturalness and controllability of synthesised speech. However, manual prosody annotation is labor-intensive and i…
text-to-speechText to Speech