paper-with-me

Papers

CRAFT: A multifunction online platform for speech prosody visualisation

2019-03-18 · Dafydd Gibbon

There are many research tools which are also used for teaching the acoustic phonetics of speech rhythm and speech melody. But they were not purpose-designed for teaching-learning situations, and some have a steep learning curve. CRAFT (Creation and Recovery of Amplitude and Frequency Tracks) is custom-designed as a novel flexible online tool for visualisation and critical comparison of functions and transforms, with implementations of the Reaper, RAPT, PyRapt, YAAPT, YIN and PySWIPE F0 estimators, three Praat configurations, and two purpose-built estimators, PyAMDF, S0FT. Visualisations of amplitude and frequency envelope spectra, spectral edge detection of rhythm zones, and a parametrised spectrogram are included. A selection of audio clips from tone and intonation languages is provided for demonstration purposes. The main advantages of online tools are consistency (users have the same version and the same data selection), interoperability over different platforms, and ease of maintenance. The code is available on GitHub.

📄 PDF Abstract BibTeX arXiv:1903.08718

Code (0)

등록된 구현이 없습니다.

Tasks

Edge DetectionRhythm

Similar Papers 제목 키워드 기반

Speech prosody and remote experiments: a technical report

2021-06-21 · Giuseppe Magistro

The aim of this paper is twofold. First, we present a review of different recording options for gathering prosodic data in the event that fieldwork is impracticable (e.g. due to pandemics). Under this light, we mimic a l…

PRESENT: Zero-Shot Text-to-Prosody Control

2024-08-13 · Perry Lam, Huayun Zhang, Nancy F. Chen, Berrak Sisman 외

Current strategies for achieving fine-grained prosody control in speech synthesis entail extracting additional style embeddings or adopting more complex architectures. To enable zero-shot application of pretrained text-t…

Prosody PredictionSpeech Synthesistext-to-speechText to Speech

ProsoSpeech: Enhancing Prosody With Quantized Vector Pre-training in Text-to-Speech

2022-02-16 · Yi Ren, Ming Lei, Zhiying Huang, Shiliang Zhang 외

Expressive text-to-speech (TTS) has become a hot research topic recently, mainly focusing on modeling prosody in speech. Prosody modeling has several challenges: 1) the extracted pitch used in previous prosody modeling w…

text-to-speechText to Speech

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models

2025-07-27 · Kaizhi Qian, Xulin Fan, Junrui Ni, Slava Shechtman 외 arxiv

Speech language models refer to language models with speech processing and understanding capabilities. One key desirable capability for speech language models is the ability to capture the intricate interdependency betwe…

DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training

2023-07-31 · Hyung-Seok Oh, Sang-Hoon Lee, Seong-Whan Lee

Expressive text-to-speech systems have undergone significant advancements owing to prosody modeling, but conventional methods can still be improved. Traditional approaches have relied on the autoregressive method to pred…

DenoisingExpressive Speech SynthesisSpeech Synthesistext-to-speech+1