paper-with-me

홈 › Papers

CNN Encoding of Acoustic Parameters for Prominence Detection

2021-04-12 · Kamini Sabu, Mithilesh Vaidya, Preeti Rao

Expressive reading, considered the defining attribute of oral reading fluency, comprises the prosodic realization of phrasing and prominence. In the context of evaluating oral reading, it helps to establish the speaker's comprehension of the text. We consider a labeled dataset of children's reading recordings for the speaker-independent detection of prominent words using acoustic-prosodic and lexico-syntactic features. A previous well-tuned random forest ensemble predictor is replaced by an RNN sequence classifier to exploit potential context dependency across the longer utterance. Further, deep learning is applied to obtain word-level features from low-level acoustic contours of fundamental frequency, intensity and spectral shape in an end-to-end fashion. Performance comparisons are presented across the different feature types and across different feature learning architectures for prominent word prediction to draw insights wherever possible.

📄 PDF Abstract BibTeX arXiv:2104.05488

Code (0)

등록된 구현이 없습니다.

Tasks

Attribute

Similar Papers 제목 키워드 기반

Deep Learning For Prominence Detection In Children's Read Speech

2021-10-27 · Mithilesh Vaidya, Kamini Sabu, Preeti Rao

The detection of perceived prominence in speech has attracted approaches ranging from the design of linguistic knowledge-based acoustic features to the automatic feature learning from suprasegmental attributes such as pi…

Deep Learning

BERT, can HE predict contrastive focus? Predicting and controlling prominence in neural TTS using a language model

2022-07-04 · Brooke Stephenson, Laurent Besacier, Laurent Girin, Thomas Hueber

Several recent studies have tested the use of transformer language model representations to infer prosodic features for text-to-speech synthesis (TTS). While these studies have explored prosody in general, in this work, …

Language ModelingLanguage ModellingSpeech Synthesistext-to-speech+2

Crocodile perception of distress in hominid baby cries

2023-10-02 · Julie Thévenet, Léo Papet, Gérard Coureaud, Nicolas Boyer 외

It is generally argued that distress vocalizations, a common modality for alerting conspecifics across a wide range of terrestrial vertebrates, share acoustic features that allow heterospecific communication. Yet studies…

Prosody leaks into the memories of words

2020-05-29 · Kevin Tang, Jason A. Shaw

The average predictability (aka informativity) of a word in context has been shown to condition word duration (Seyfarth, 2014). All else being equal, words that tend to occur in more predictable environments are shorter …

Prominence-aware automatic speech recognition for conversational speech

2025-09-12 · Julian Linke, Barbara Schuppler arxiv

This paper investigates prominence-aware automatic speech recognition (ASR) by combining prominence detection and speech recognition for conversational Austrian German. First, prominence detectors were developed by fine-…

Speech Recognition