paper-with-me

홈 › Papers

VOLIP: a corpus of spoken Italian and a virtuous example of reuse of linguistic resources

2014-05-01 · LREC 2014 5 · Iol Alfano, a, Francesco Cutugno, Aurelio De Rosa, Claudio Iacobini, Renata Savy, Miriam Voghera

The corpus VoLIP (The Voice of LIP) is an Italian speech resource which associates the audio signals to the orthographic transcriptions of the LIP Corpus. The LIP Corpus was designed to represent diaphasic, diatopic and diamesic variation. The Corpus was collected in the early ‘90s to compile a frequency lexicon of spoken Italian and its size was tailored to produce a reliable frequency lexicon for the first 3,000 lemmas. Therefore, it consists of about 500,000 word tokens for 60 hours of recording. The speech materials belong to five different text registers and they were collected in four different cities. Thanks to a modern technological approach VoLIP web service allows users to search the LIP corpus using IMDI metadata, lexical or morpho-syntactic entry keys, receiving as result the audio portions aligned to the corresponding required entry. The VoLIP corpus is freely available at the URL http://www.parlaritaliano.it.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards the first UD Treebank of Spoken Italian: the KIParla forest

2024-10-06 · Ludovica Pannitto

The present project endeavors to enrich the linguistic resources available for Italian by constructing a Universal Dependencies treebank for the KIParla corpus (Mauri et al., 2019, Ballar\`e et al., 2020), an existing an…

The KIPARLA Forest treebank of spoken Italian: an overview of initial design choices

2024-11-10 · Ludovica Pannitto, Caterina Mauri

The paper presents an overview of initial design choices discussed towards the creation of a treebank for the Italian KIParla corpus

Leveraging study of robustness and portability of spoken language understanding systems across languages and domains: the PORTMEDIA corpora

2012-05-01 · LREC 2012 5 · Fabrice Lef{\`e}vre, Djamel Mostefa, Laurent Besacier, Yannick Est{\`e}ve 외

The PORTMEDIA project is intended to develop new corpora for the evaluation of spoken language understanding systems. The newly collected data are in the field of human-machine dialogue systems for tourist information in…

Semantic CompositionSpeech RecognitionSpoken Language Understanding

Cross-corpora experiments of automatic proficiency assessment and error detection for spoken English

2022-07-01 · NAACL (BEA) 2022 7 · Stefano Bannò, Marco Matassoni

The growing demand for learning English as a second language has led to an increasing interest in automatic approaches for assessing spoken language proficiency. One of the most significant challenges in this field is th…

The Trilingual ALLEGRA Corpus: Presentation and Possible Use for Lexicon Induction

2012-05-01 · LREC 2012 5 · Yves Scherrer, Bruno Cartoni

In this paper, we present a trilingual parallel corpus for German, Italian and Romansh, a Swiss minority language spoken in the canton of Grisons. The corpus called ALLEGRA contains press releases automatically gathered …

Sentence