paper-with-me

Papers

Computational Induction of Prosodic Structure

2019-12-15 · Dafydd Gibbon

The present study has two goals relating to the grammar of prosody, understood as the rhythms and melodies of speech. First, an overview is provided of the computable grammatical and phonetic approaches to prosody analysis which use hypothetico-deductive methods and are based on learned hermeneutic intuitions about language. Second, a proposal is presented for an inductive grounding in the physical signal, in which prosodic structure is inferred using a language-independent method from the low-frequency spectrum of the speech signal. The overview includes a discussion of computational aspects of standard generative and post-generative models, and suggestions for reformulating these to form inductive approaches. Also included is a discussion of linguistic phonetic approaches to analysis of annotations (pairs of speech unit labels with time-stamps) of recorded spoken utterances. The proposal introduces the inductive approach of Rhythm Formant Theory (RFT) and the associated Rhythm Formant Analysis (RFA) method are introduced, with the aim of completing a gap in the linguistic hypothetico-deductive cycle by grounding in a language-independent inductive procedure of speech signal analysis. The validity of the method is demonstrated and applied to rhythm patterns in read-aloud Mandarin Chinese, finding differences from English which are related to lexical and grammatical differences between the languages, as well as individual variation. The overall conclusions are (1) that normative language-to-language phonological or phonetic comparisons of rhythm, for example of Mandarin and English, are too simplistic, in view of diverse language-internal factors due to genre and style differences as well as utterance dynamics, and (2) that language-independent empirical grounding of rhythm in the physical signal is called for.

📄 PDF Abstract BibTeX arXiv:1912.07050

Code (0)

등록된 구현이 없습니다.

Tasks

Rhythm

Similar Papers 제목 키워드 기반

Double Articulation Analyzer with Prosody for Unsupervised Word and Phoneme Discovery

2021-03-15 · Yasuaki Okuda, Ryo Ozaki, Tadahiro Taniguchi

Infants acquire words and phonemes from unsegmented speech signals using segmentation cues, such as distributional, prosodic, and co-occurrence cues. Many pre-existing computational models that represent the process tend…

Language ModellingTime SeriesTime Series Analysis

The Future of Prosody: It's about Time

2018-04-23 · Dafydd Gibbon

Prosody is usually defined in terms of the three distinct but interacting domains of pitch, intensity and duration patterning, or, more generally, as phonological and phonetic properties of 'suprasegmentals', speech segm…

A Character-level Span-based Model for Mandarin Prosodic Structure Prediction

2022-03-31 · Xueyuan Chen, Changhe Song, Yixuan Zhou, Zhiyong Wu 외

The accuracy of prosodic structure prediction is crucial to the naturalness of synthesized speech in Mandarin text-to-speech system, but now is limited by widely-used sequence-to-sequence framework and error accumulation…

Sentencetext-to-speechText to Speech

Logical Transductions for the Typology of Ditransitive Prosody

2022-07-01 · NAACL (SIGMORPHON) 2022 7 · Mai Ha Vu, Aniello De Santo, Hossep Dolatian

Given the empirical landscape of possible prosodic parses, this paper examines the computations required to formalize the mapping from syntactic structure to prosodic structure. In particular, we use logical tree transdu…

Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework

2025-10-08 · Zhu Li, Yuqing Zhang, Xiyuan Gao, Shekhar Nayak 외 arxiv

Sarcasm is a pragmatic phenomenon in which speakers convey meanings that diverge from literal content, relying on an interaction between semantics and prosodic expression. However, how these cues jointly contribute to th…

Speech Synthesis