paper-with-me

홈 › Papers

The Corpus of Finnish Sign Language

2020-05-01 · LREC 2020 5 · Juhana Salonen, Antti Kronqvist, Tommi Jantunen

This paper presents the Corpus of Finnish Sign Language (Corpus FinSL), a structured and annotated collection of Finnish Sign Language (FinSL) videos published in May 2019 in FIN-CLARIN{'}s Language Bank of Finland. The corpus is divided into two subcorpora, one of which comprises elicited narratives and the other conversations. All of the FinSL material has been annotated using ELAN and the lexical database Finnish Signbank. Basic annotation includes ID-glosses and translations into Finnish. The anonymized metadata of Corpus FinSL has been organized in accordance with the IMDI standard. Altogether, Corpus FinSL contains nearly 15 hours of video material from 21 FinSL users. Corpus FinSL has already been exploited in FinSL research and teaching, and it is predicted that in the future it will have a significant positive impact on these fields as well as on the status of the sign language community in Finland. Keywords: Corpus of Finnish Sign Language, Language Bank of Finland, Finnish Signbank, annotation, metadata, research, teaching

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Dialect Text Normalization to Normative Standard Finnish

2019-11-01 · WS 2019 11 · Niko Partanen, Mika H{\"a}m{\"a}l{\"a}inen, Khalid Alnajjar

We compare different LSTMs and transformer models in terms of their effectiveness in normalizing dialectal Finnish into the normative standard Finnish. As dialect is the common way of communication for people online in F…

Text Normalization

Fine-grained Named Entity Annotation for Finnish

2021-05-01 · NoDaLiDa 2021 5 · Jouni Luoma, Li-Hsin Chang, Filip Ginter, Sampo Pyysalo

We introduce a corpus with fine-grained named entity annotation for Finnish, following the OntoNotes guidelines to create a resource that is cross-lingually compatible with existing annotations for other languages. We co…

NER

A Finnish News Corpus for Named Entity Recognition

2019-08-12 · Teemu Ruokolainen, Pekka Kauppinen, Miikka Silfverberg, Krister Lindén

We present a corpus of Finnish news articles with a manually prepared named entity annotation. The corpus consists of 953 articles (193,742 word tokens) with six named entity classes (organization, location, person, prod…

Articlesnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)

Development of the Siberian Ingrian Finnish Speech Corpus

2022-05-01 · ComputEL (ACL) 2022 5 · Ivan Ubaleht, Taisto-Kalevi Raudalainen

In this paper we present the speech corpus for the Siberian Ingrian Finnish language. The speech corpus includes audio data, annotations, software tools for data-processing, two databases and a web application. We have p…

Specifying Treebanks, Outsourcing Parsebanks: FinnTreeBank 3

2012-05-01 · LREC 2012 5 · Atro Voutilainen, Kristiina Muhonen, Tanja Purtonen, Krister Lind{\'e}n

Corpus-based treebank annotation is known to result in incomplete coverage of mid- and low-frequency linguistic constructions: the linguistic representation and corpus annotation quality are sometimes suboptimal. Large d…

Descriptive