paper-with-me

Papers

Modelling prosodic structure using Artificial Neural Networks

2017-06-13 · Jean-Philippe Bernardy, Charalambos Themistocleous

The ability to accurately perceive whether a speaker is asking a question or is making a statement is crucial for any successful interaction. However, learning and classifying tonal patterns has been a challenging task for automatic speech recognition and for models of tonal representation, as tonal contours are characterized by significant variation. This paper provides a classification model of Cypriot Greek questions and statements. We evaluate two state-of-the-art network architectures: a Long Short-Term Memory (LSTM) network and a convolutional network (ConvNet). The ConvNet outperforms the LSTM in the classification task and exhibited an excellent performance with 95% classification accuracy.

📄 PDF Abstract BibTeX arXiv:1706.03952

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)ClassificationGeneral Classificationspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Sigmoid Activation 설명 없음
Tanh Activation 설명 없음
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…

Similar Papers 제목 키워드 기반

A Variational Prosody Model for the decomposition and synthesis of speech prosody

2018-06-22 · Branislav Gerazov, Gérard Bailly, Omar Mohammed, Yi Xu 외

The quest for comprehensive generative models of intonation that link linguistic and paralinguistic functions to prosodic forms has been a longstanding challenge of speech communication research. More traditional intonat…

Speech Synthesis

The Prosody of Emojis

2025-08-01 · Giulio Zhou, Tsz Kin Lam, Alexandra Birch, Barry Haddow arxiv

Prosodic features such as pitch, timing, and intonation are central to spoken communication, conveying emotion, intent, and discourse structure. In text-based settings, where these cues are absent, emojis act as visual s…

ProSarc: Prosody-Aware Sarcasm Recognition Framework via Temporal Prosodic Incongruity

2026-06-04 · Prathamjyot Singh, Ashima Sood, Sahil Sharma, Jasmeet Singh arxiv

We present ProSarc, an audio-only framework that detects sarcasm by modelling temporal prosodic incongruity, that is, the mismatch between local prosodic dynamics and the utterance-level emotional baseline. Dual encoding…

Hierarchical Representation of Prosody for Statistical Speech Synthesis

2015-10-07 · Antti Suni, Daniel Aalto, Martti Vainio

Prominences and boundaries are the essential constituents of prosodic structure in speech. They provide for means to chunk the speech stream into linguistically relevant units by providing them with relative saliences an…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

The Role of Prosody and Speech Register in Word Segmentation: A Computational Modelling Perspective

2017-07-01 · ACL 2017 7 · Bogdan Ludusan, Reiko Mazuka, Mathieu Bernard, Alej Cristia 외

This study explores the role of speech register and prosody for the task of word segmentation. Since these two factors are thought to play an important role in early language acquisition, we aim to quantify their contrib…

Language AcquisitionSegmentation