paper-with-me

Papers

LaVA – Latvian Language Learner corpus

2022-06-01 · LREC 2022 6 · Roberts Darģis, Ilze Auziņa, Inga Kaija, Kristīne Levāne-Petrova, Kristīne Pokratniece

This paper presents the Latvian Language Learner Corpus (LaVA) developed at the Institute of Mathematics and Computer Science, University of Latvia. LaVA corpus contains 1015 essays (190k tokens and 790k characters excluding whitespaces) from foreigners studying at Latvian higher education institutions and who are learning Latvian as a foreign language in the first or second semester, reaching the A1 (possibly A2) Latvian language proficiency level. The corpus has morphological and error annotations. Error analysis and the statistics of the LaVA corpus are also provided in the paper. The corpus is publicly available at: http://www.korpuss.lv/id/LaVA.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Quality Focused Approach to a Learner Corpus Development

2020-05-01 · LREC 2020 5 · Roberts Dar{\c{g}}is, Ilze Auzi{\c{n}}a, Krist{\=\i}ne Lev{\=a}ne-Petrova, Inga Kaija

The paper presents quality focused approach to a learner corpus development. The methodology was developed with multiple design considerations put in place to make the annotation process easier and at the same time reduc…

Morphological Analysis

The Use of Text Alignment in Semi-Automatic Error Analysis: Use Case in the Development of the Corpus of the Latvian Language Learners

2018-05-01 · LREC 2018 5 · Roberts Dar{\c{g}}is, Ilze Auzi{\c{n}}a, Krist{\=\i}ne Lev{\=a}ne-Petrova
Language AcquisitionLemmatizationMorphological AnalysisPart-Of-Speech Tagging+1

Designing the Latvian Speech Recognition Corpus

2014-05-01 · LREC 2014 5 · M{\=a}rcis Pinnis, Ilze Auzi{\c{n}}a, K{\=a}rlis Goba

In this paper the authors present the first Latvian speech corpus designed specifically for speech recognition purposes. The paper outlines the decisions made in the corpus designing process through analysis of related w…

speech-recognitionSpeech RecognitionSpeech Synthesis

Deriving a PropBank Corpus from Parallel FrameNet and UD Corpora

2020-05-01 · LREC 2020 5 · Normunds Gruzitis, Roberts Dar{\c{g}}is, Laura Rituma, Gunta Ne{\v{s}}pore-B{\=e}rzkalne 외

We propose an approach for generating an accurate and consistent PropBank-annotated corpus, given a FrameNet-annotated corpus which has an underlying dependency annotation layer, namely, a parallel Universal Dependencies…

Development and Evaluation of Speech Synthesis Corpora for Latvian

2020-05-01 · LREC 2020 5 · Roberts Dar{\c{g}}is, Peteris Paikens, Normunds Gruzitis, Ilze Auzina 외

Text to speech (TTS) systems are necessary for all languages to ensure accessibility and availability of digital language services. Recent advances in neural speech synthesis have eText to speech (TTS) systems are necess…

speech-recognitionSpeech RecognitionSpeech Synthesistext-to-speech+1