paper-with-me

홈 › Papers

Speech earthquakes: scaling and universality in human voice

2014-08-05 · Jordi Luque, Bartolo Luque, Lucas Lacasa

Speech is a distinctive complex feature of human capabilities. In order to understand the physics underlying speech production, in this work we empirically analyse the statistics of large human speech datasets ranging several languages. We first show that during speech the energy is unevenly released and power-law distributed, reporting a universal robust Gutenberg-Richter-like law in speech. We further show that such earthquakes in speech show temporal correlations, as the interevent statistics are again power-law distributed. Since this feature takes place in the intra-phoneme range, we conjecture that the responsible for this complex phenomenon is not cognitive, but it resides on the physiological speech production mechanism. Moreover, we show that these waiting time distributions are scale invariant under a renormalisation group transformation, suggesting that the process of speech generation is indeed operating close to a critical point. These results are put in contrast with current paradigms in speech processing, which point towards low dimensional deterministic chaos as the origin of nonlinear traits in speech fluctuations. As these latter fluctuations are indeed the aspects that humanize synthetic speech, these findings may have an impact in future speech synthesis technologies. Results are robust and independent of the communication language or the number of speakers, pointing towards an universal pattern and yet another hint of complexity in human speech.

📄 PDF Abstract BibTeX arXiv:1408.0985

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesis

Similar Papers 제목 키워드 기반

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

2025-05-23 · Zhihao Du, Changfeng Gao, Yuxuan Wang, Fan Yu 외

In our prior works, we introduced a scalable streaming speech synthesis model, CosyVoice 2, which integrates a large language model (LLM) and a chunk-aware flow matching (FM) model, and achieves low-latency bi-streaming …

Automatic Speech RecognitionEmotion RecognitionEvent DetectionLanguage Identification+5

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

2024-12-03 · Aohan Zeng, Zhengxiao Du, Mingdao Liu, Kedong Wang 외

We introduce GLM-4-Voice, an intelligent and human-like end-to-end spoken chatbot. It supports both Chinese and English, engages in real-time voice conversations, and varies vocal nuances such as emotion, intonation, spe…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)ChatbotLanguage Modeling+6

The Voice Behind the Words: Quantifying Intersectional Bias in SpeechLLMs

2026-03-15 · Shree Harsha Bokkahalli Satish, Christoph Minixhofer, Maria Teleki, James Caverlee 외 arxiv

Speech Large Language Models (SpeechLLMs) process spoken input directly, retaining cues such as accent and perceived gender that were previously removed in cascaded pipelines. This introduces speaker identity dependent v…

UltraVoice: Scaling Fine-Grained Style-Controlled Speech Conversations for Spoken Dialogue Models

2025-10-26 · Wenming Tu, Guanrou Yang, Ruiqi Yan, Wenxi Chen 외 arxiv

Spoken dialogue models currently lack the ability for fine-grained speech style control, a critical capability for human-like interaction that is often overlooked in favor of purely functional capabilities like reasoning…

Instruction FollowingQuestion AnsweringSpeech Synthesis

Deep Voice 3: Scaling Text-to-Speech with Convolutional Sequence Learning

2017-10-20 · ICLR 2018 1 · Wei Ping, Kainan Peng, Andrew Gibiansky, Sercan O. Arik 외

We present Deep Voice 3, a fully-convolutional attention-based neural text-to-speech (TTS) system. Deep Voice 3 matches state-of-the-art neural speech synthesis systems in naturalness while training ten times faster. We …

GPUSpeech Synthesistext-to-speechText to Speech