paper-with-me

Papers

Investigating the Utility of Surprisal from Large Language Models for Speech Synthesis Prosody

2023-06-16 · Sofoklis Kakouros, Juraj Šimko, Martti Vainio, Antti Suni

This paper investigates the use of word surprisal, a measure of the predictability of a word in a given context, as a feature to aid speech synthesis prosody. We explore how word surprisal extracted from large language models (LLMs) correlates with word prominence, a signal-based measure of the salience of a word in a given discourse. We also examine how context length and LLM size affect the results, and how a speech synthesizer conditioned with surprisal values compares with a baseline system. To evaluate these factors, we conducted experiments using a large corpus of English text and LLMs of varying sizes. Our results show that word surprisal and word prominence are moderately correlated, suggesting that they capture related but distinct aspects of language use. We find that length of context and size of the LLM impact the correlations, but not in the direction anticipated, with longer contexts and larger LLMs generally underpredicting prominent words in a nearly linear manner. We demonstrate that, in line with these findings, a speech synthesizer conditioned with surprisal values provides a minimal improvement over the baseline with the results suggesting a limited effect of using surprisal values for eliciting appropriate prominence patterns.

📄 PDF Abstract BibTeX arXiv:2306.09814

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesis

Similar Papers 제목 키워드 기반

Testing the Predictions of Surprisal Theory in 11 Languages

2023-07-07 · Ethan Gotlieb Wilcox, Tiago Pimentel, Clara Meister, Ryan Cotterell 외

A fundamental result in psycholinguistics is that less predictable words take a longer time to process. One theoretical explanation for this finding is Surprisal Theory (Hale, 2001; Levy, 2008), which quantifies a word's…

Light-weight Pronunciation Assessment via Discrete Speech Token Surprisal

2026-06-18 · Syeda Faiza Ahmed Sara, Shammur Absar Chowdhury arxiv

Training automated pronunciation assessment often relies on labeled learner errors or non-native corpora that are costly to collect. We propose a lightweight framework trained only on native speech resources, operating u…

Expectation and Locality Effects in the Prediction of Disfluent Fillers and Repairs in English Speech

2019-06-01 · NAACL 2019 6 · Samvit Dammalapati, Rajakrishnan Rajkumar, Sumeet Agarwal

This study examines the role of three influential theories of language processing, \textit{viz.}, Surprisal Theory, Uniform Information Density (UID) hypothesis and Dependency Locality Theory (DLT), in predicting disflue…

A Computational Approach to Analyzing Disrupted Language in Schizophrenia: Integrating Surprisal and Coherence Measures

2025-11-05 · Gowtham Premananth, Carol Espy-Wilson arxiv

Language disruptions are one of the well-known effects of schizophrenia symptoms. They are often manifested as disorganized speech and impaired discourse coherence. These abnormalities in spontaneous language production …

Effects of Duration, Locality, and Surprisal in Speech Disfluency Prediction in English Spontaneous Speech

2021-02-01 · SCiL 2021 2 · Samvit Dammalapati, Rajakrishnan Rajkumar, Sumeet Agarwal