paper-with-me

Papers

ReadAlong Studio: Practical Zero-Shot Text-Speech Alignment for Indigenous Language Audiobooks

2022-06-01 · SIGUL (LREC) 2022 6 · Patrick Littell, Eric Joanis, Aidan Pine, Marc Tessier, David Huggins Daines, Delasie Torkornoo

While the alignment of audio recordings and text (often termed “forced alignment”) is often treated as a solved problem, in practice the process of adapting an alignment system to a new, under-resourced language comes with significant challenges, requiring experience and expertise that many outside of the speech community lack. This puts otherwise “solvable” problems, like the alignment of Indigenous language audiobooks, out of reach for many real-world Indigenous language organizations. In this paper, we detail ReadAlong Studio, a suite of tools for creating and visualizing aligned audiobooks, including educational features like time-aligned highlighting, playing single words in isolation, and variable-speed playback. It is intended to be accessible to creators without an extensive background in speech or NLP, by automating or making optional many of the specialist steps in an alignment pipeline. It is well documented at a beginner-technologist level, has already been adapted to 30 languages, and can work out-of-the-box on many more languages without adaptation.

📄 PDF Abstract BibTeX

Code (2)

readalongs/studio 공식 구현
readalongs/web-component 공식 구현

Similar Papers 제목 키워드 기반

i-Code Studio: A Configurable and Composable Framework for Integrative AI

2023-05-23 · Yuwei Fang, Mahmoud Khademi, Chenguang Zhu, ZiYi Yang 외

Artificial General Intelligence (AGI) requires comprehensive understanding and generation capabilities for a variety of tasks spanning different modalities and functionalities. Integrative AI is one important direction t…

Question AnsweringRetrievalSpeech-to-Speech TranslationText Retrieval+2

DialogStudio: Towards Richest and Most Diverse Unified Dataset Collection for Conversational AI

2023-07-19 · JianGuo Zhang, Kun Qian, Zhiwei Liu, Shelby Heinecke 외

Despite advancements in conversational AI, language models encounter challenges to handle diverse conversational tasks, and existing dialogue dataset collections often lack diversity and comprehensiveness. To tackle thes…

Conversational RecommendationDiversityFew-Shot LearningLanguage Modeling+2

Learning to Speak from Text: Zero-Shot Multilingual Text-to-Speech with Unsupervised Text Pretraining

2023-01-30 · Takaaki Saeki, Soumi Maiti, Xinjian Li, Shinji Watanabe 외

While neural text-to-speech (TTS) has achieved human-like natural synthetic speech, multilingual TTS systems are limited to resource-rich languages due to the need for paired text and studio-quality audio data. This pape…

Language ModelingLanguage Modellingtext-to-speechText to Speech

LoRP-TTS: Low-Rank Personalized Text-To-Speech

2025-02-11 · Łukasz Bondaruk, Jakub Kubiak

Speech synthesis models convert written text into natural-sounding audio. While earlier models were limited to a single speaker, recent advancements have led to the development of zero-shot systems that generate realisti…

Speech Synthesistext-to-speechText to Speech

What Changes Can Large-scale Language Models Bring? Intensive Study on HyperCLOVA: Billions-scale Korean Generative Pretrained Transformers

2021-09-10 · EMNLP 2021 11 · Boseop Kim, HyoungSeok Kim, Sang-Woo Lee, Gichang Lee 외

GPT-3 shows remarkable in-context learning ability of large-scale language models (LMs) trained on hundreds of billion scale data. Here we address some remaining issues less reported by the GPT-3 paper, such as a non-Eng…

Few-Shot LearningIn-Context LearningPrompt Engineering