paper-with-me

Papers

Prosodic Event Recognition using Convolutional Neural Networks with Context Information

2017-06-02 · Sabrina Stehwien, Ngoc Thang Vu

This paper demonstrates the potential of convolutional neural networks (CNN) for detecting and classifying prosodic events on words, specifically pitch accents and phrase boundary tones, from frame-based acoustic features. Typical approaches use not only feature representations of the word in question but also its surrounding context. We show that adding position features indicating the current word benefits the CNN. In addition, this paper discusses the generalization from a speaker-dependent modelling approach to a speaker-independent setup. The proposed method is simple and efficient and yields strong results not only in speaker-dependent but also speaker-independent cases.

📄 PDF Abstract BibTeX arXiv:1706.00741

Code (0)

등록된 구현이 없습니다.

Tasks

Position

Similar Papers 제목 키워드 기반

Voice Quality and Pitch Features in Transformer-Based Speech Recognition

2021-12-21 · Guillermo Cámbara, Jordi Luque, Mireia Farrús

Jitter and shimmer measurements have shown to be carriers of voice quality and prosodic information which enhance the performance of tasks like speaker recognition, diarization or automatic speech recognition (ASR). Howe…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speaker Recognitionspeech-recognition+1

Non-verbal information in spontaneous speech -- towards a new framework of analysis

2024-03-06 · Tirza Biron, Moshe Barboy, Eran Ben-Artzy, Alona Golubchik 외

Non-verbal signals in speech are encoded by prosody and carry information that ranges from conversation action to attitude and emotion. Despite its importance, the principles that govern prosodic structure are not yet ad…

speech-recognitionSpeech Recognition

VowelPrompt: Hearing Speech Emotions from Text via Vowel-level Prosodic Augmentation

2026-02-06 · Yancheng Wang, Osama Hanna, Ruiming Xie, Xianfeng Rui 외 arxiv

Emotion recognition in speech presents a complex multimodal challenge, requiring comprehension of both linguistic content and vocal expressivity, particularly prosodic features such as fundamental frequency, intensity, a…

Reinforcement LearningEmotion Recognition

Improving coreference resolution with automatically predicted prosodic information

2017-07-28 · WS 2017 9 · Ina Rösiger, Sabrina Stehwien, Arndt Riester, Ngoc Thang Vu

Adding manually annotated prosodic information, specifically pitch accents and phrasing, to the typical text-based feature set for coreference resolution has previously been shown to have a positive effect on German data…

coreference-resolutionCoreference Resolution

Improving Mandarin Prosodic Structure Prediction with Multi-level Contextual Information

2023-08-31 · Jie Chen, Changhe Song, Deyi Tuo, Xixin Wu 외

For text-to-speech (TTS) synthesis, prosodic structure prediction (PSP) plays an important role in producing natural and intelligible speech. Although inter-utterance linguistic information can influence the speech inter…

DecoderMulti-Task Learningtext-to-speechText to Speech