paper-with-me

홈 › Papers

An Annotated Corpus of Direct Speech

2016-05-01 · LREC 2016 5 · John Lee, Chak Yan Yeung

We propose a scheme for annotating direct speech in literary texts, based on the Text Encoding Initiative (TEI) and the coreference annotation guidelines from the Message Understanding Conference (MUC). The scheme encodes the speakers and listeners of utterances in a text, as well as the quotative verbs that reports the utterances. We measure inter-annotator agreement on this annotation task. We then present statistics on a manually annotated corpus that consists of books from the New Testament. Finally, we visualize the corpus as a conversational network.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

QUEMDISSE? Reported speech in Portuguese

2016-05-01 · LREC 2016 5 · Cl{\'a}udia Freitas, Bianca Freitas, Diana Santos

This paper presents some work on direct and indirect speech in Portuguese using corpus-based methods: we report on a study whose aim was to identify (i) Portuguese verbs used to introduce reported speech and (ii) syntact…

The IFCASL Corpus of French and German Non-native and Native Read Speech

2016-05-01 · LREC 2016 5 · Juergen Trouvain, Anne Bonneau, Vincent Colotte, Camille Fauth 외

The IFCASL corpus is a French-German bilingual phonetic learner corpus designed, recorded and annotated in a project on individualized feedback in computer-assisted spoken language learning. The motivation for setting up…

Preparation of Bangla Speech Corpus from Publicly Available Audio \& Text

2020-05-01 · LREC 2020 5 · Shafayat Ahmed, Nafis Sadeq, Sudipta Saha Shubha, Md. Nahidul Islam 외

Automatic speech recognition systems require large annotated speech corpus. The manual annotation of a large corpus is very difficult. In this paper, we focus on the automatic preparation of a speech corpus for Banglades…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speaker-diarizationSpeaker Diarization+2

Designing the Latvian Speech Recognition Corpus

2014-05-01 · LREC 2014 5 · M{\=a}rcis Pinnis, Ilze Auzi{\c{n}}a, K{\=a}rlis Goba

In this paper the authors present the first Latvian speech corpus designed specifically for speech recognition purposes. The paper outlines the decisions made in the corpus designing process through analysis of related w…

speech-recognitionSpeech RecognitionSpeech Synthesis

MERLIon CCS Challenge: A English-Mandarin code-switching child-directed speech corpus for language identification and diarization

2023-05-30 · Victoria Y. H. Chua, Hexin Liu, Leibny Paola Garcia Perera, Fei Ting Woon 외

To enhance the reliability and robustness of language identification (LID) and language diarization (LD) systems for heterogeneous populations and scenarios, there is a need for speech processing models to be trained on …

Language Identification