paper-with-me

Papers

EMNS /Imz/ Corpus: An emotive single-speaker dataset for narrative storytelling in games, television and graphic novels

2023-05-22 · Kari Ali Noriy, Xiaosong Yang, Jian Jun Zhang

The increasing adoption of text-to-speech technologies has led to a growing demand for natural and emotive voices that adapt to a conversation's context and emotional tone. The Emotive Narrative Storytelling (EMNS) corpus is a unique speech dataset created to enhance conversations' expressiveness and emotive quality in interactive narrative-driven systems. The corpus consists of a 2.3-hour recording featuring a female speaker delivering labelled utterances. It encompasses eight acted emotional states, evenly distributed with a variance of 0.68%, along with expressiveness levels and natural language descriptions with word emphasis labels. The evaluation of audio samples from different datasets revealed that the EMNS corpus achieved the highest average scores in accurately conveying emotions and demonstrating expressiveness. It outperformed other datasets in conveying shared emotions and achieved comparable levels of genuineness. A classification task confirmed the accurate representation of intended emotions in the corpus, with participants recognising the recordings as genuine and expressive. Additionally, the availability of the dataset collection tool under the Apache 2.0 License simplifies remote speech data collection for researchers.

📄 PDF Abstract BibTeX arXiv:2305.13137

Code (1)

knoriy/emns-dct 공식 구현

Tasks

Expressive Speech SynthesisSpeech Synthesistext-to-speechText to Speech

Similar Papers 제목 키워드 기반

Syntactic characteristics of emotive predicates in Bulgarian: A corpus-based study

2022-09-01 · CLIB 2022 9 · Yovka Tisheva, Marina Dzhonova

The paper presents a corpus-based study of emotive predicates (verbs and predicative constructions with adjectival, adverbial or noun phrases) in Bulgarian with respect to their syntactic characteristics. The sources of …

Expanding the Workspace of Electromagnetic Navigation Systems Using Dynamic Feedback for Single- and Multi-agent Control

2025-11-23 · Jasan Zughaibi, Denis von Arx, Maurus Derungs, Florian Heemeyer 외 arxiv

Electromagnetic navigation systems (eMNS) enable a number of magnetically guided surgical procedures. A challenge in magnetically manipulating surgical tools is that the effective workspace of an eMNS is often severely c…

Pose Estimation

Modeling Electromagnetic Navigation Systems for Medical Applications using Random Forests and Artificial Neural Networks

2019-09-26 · Ruoxi Yu, Samuel L. Charreyron, Quentin Boehler, Cameron Weibel 외

Electromagnetic Navigation Systems (eMNS) can be used to control a variety of multiscale devices within the human body for remote surgery. Accurate modeling of the magnetic fields generated by the electromagnets of an eM…

BIG-bench Machine Learning

Mellotron: Multispeaker expressive voice synthesis by conditioning on rhythm, pitch and global style tokens

2019-10-26 · Rafael Valle, Jason Li, Ryan Prenger, Bryan Catanzaro

Mellotron is a multispeaker voice synthesis model based on Tacotron 2 GST that can make a voice emote and sing without emotive or singing training data. By explicitly conditioning on rhythm and continuous pitch contours …

RhythmStyle Transfer

Transfer Learning Framework for Low-Resource Text-to-Speech using a Large-Scale Unlabeled Speech Corpus

2022-03-29 · Minchan Kim, Myeonghun Jeong, Byoung Jin Choi, Sunghwan Ahn 외

Training a text-to-speech (TTS) model requires a large scale text labeled speech corpus, which is troublesome to collect. In this paper, we propose a transfer learning framework for TTS that utilizes a large amount of un…

text-to-speechText to SpeechTransfer LearningZero-Shot Multi-Speaker TTS