paper-with-me

홈 › Papers

The I3MEDIA speech database: a trilingual annotated corpus for the analysis and synthesis of emotional speech

2012-05-01 · LREC 2012 5 · Juan Mar{\'\i}a Garrido, Yesika Laplaza, Montse Marquina, Andrea Pearman, Jos{\'e} Gregorio Escalada, Miguel {\'A}ngel Rodr{\'\i}guez, Ana Armenta

In this article the I3Media corpus is presented, a trilingual (Catalan, English, Spanish) speech database of neutral and emotional material collected for analysis and synthesis purposes. The corpus is actually made up of six different subsets of material: a neutral subcorpus, containing emotionless utterances; a ‘dialog' subcorpus, containing typical call center utterances; an ‘emotional' corpus, a set of sentences representative of pure emotional states; a ‘football' subcorpus, including utterances imitating a football broadcasting situation; a ‘SMS' subcorpus, including readings of SMS texts; and a ‘paralinguistic elements' corpus, including recordings of interjections and paralinguistic sounds uttered in isolation. The corpus was read by professional speakers (male, in the case of Spanish and Catalan; female, in the case of the English corpus), carefully selected to meet criteria of language competence, voice quality and acting conditions. It is the result of a collaboration between the Speech Technology Group at Telef{\'o}nica Investigaci{\'o}n y Desarrollo (TID) and the Speech and Language Group at Barcelona Media Centre d'Innovaci{\'o} (BM), as part of the I3Media project.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Trilingual Topic Modeling of Sri Lankan Parliamentary Debates

2026-06-18 · Himath Dhanapala, Haren Daishika, Himandhi Kuruppu, Sithija Seneviratne 외 arxiv

Sri Lankan parliamentary debates (Hansards) constitute a trilingual corpus of speeches in Sinhala, Tamil, and English, including code-mixed content, yet remain inaccessible to standard NLP pipelines due to layout-complex…

The POTUS Corpus, a Database of Weekly Addresses for the Study of Stance in Politics and Virtual Agents

2020-05-01 · LREC 2020 5 · Thomas Janssoone, K{\'e}vin Bailly, Ga{\"e}l Richard, Chlo{\'e} Clavel

One of the main challenges in the field of Embodied Conversational Agent (ECA) is to generate socially believable agents. The common strategy for agent behaviour synthesis is to rely on dedicated corpus analysis. Such a …

Versatile Speech Databases for High Quality Synthesis for Basque

2012-05-01 · LREC 2012 5 · I{\~n}aki Sainz, Daniel Erro, Eva Navas, Inma Hern{\'a}ez 외

This paper presents three new speech databases for standard Basque. They are designed primarily for corpus-based synthesis but each database has its specific purpose: 1) AhoSyn: high quality speech synthesis (recorded al…

Emotional Speech SynthesisSpeech SynthesisVocal Bursts Intensity PredictionVoice Conversion

The Trilingual ALLEGRA Corpus: Presentation and Possible Use for Lexicon Induction

2012-05-01 · LREC 2012 5 · Yves Scherrer, Bruno Cartoni

In this paper, we present a trilingual parallel corpus for German, Italian and Romansh, a Swiss minority language spoken in the canton of Grisons. The corpus called ALLEGRA contains press releases automatically gathered …

Sentence

HateBR: Large expert annotated corpus of Brazilian Instagram comments for abusive language detection

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Due to the severity of the social media abusive comments in Brazil, and the lack of research in Portuguese, this paper provides the first large-scale annotated corpus of Brazilian Instagram comments for hate speech and o…

Abusive LanguageBinary Classification