paper-with-me

홈 › Papers

Development of the Siberian Ingrian Finnish Speech Corpus

2022-05-01 · ComputEL (ACL) 2022 5 · Ivan Ubaleht, Taisto-Kalevi Raudalainen

In this paper we present the speech corpus for the Siberian Ingrian Finnish language. The speech corpus includes audio data, annotations, software tools for data-processing, two databases and a web application. We have published part of the audio data and annotations. The software tool for parsing annotation files and feeding a relational database is developed and published under a free license. A web application is developed and available. At this moment, about 300 words and 200 phrases can be displayed using this web application.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Lahjoita puhetta -- a large-scale corpus of spoken Finnish with some benchmarks

2022-03-24 · Anssi Moisio, Dejan Porjazovski, Aku Rouhe, Yaroslav Getman 외

The Donate Speech campaign has so far succeeded in gathering approximately 3600 hours of ordinary, colloquial Finnish speech into the Lahjoita puhetta (Donate Speech) corpus. The corpus includes over twenty thousand spea…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Finnish Parliament ASR corpus - Analysis, benchmarks and statistics

2022-03-28 · Anja Virkkunen, Aku Rouhe, Nhan Phan, Mikko Kurimo

Public sources like parliament meeting recordings and transcripts provide ever-growing material for the training and evaluation of automatic speech recognition (ASR) systems. In this paper, we publish and analyse the Fin…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

UralicNLP: An NLP Library for Uralic Languages

2019-05-09 · Journal of Open Source Software 2019 5 · Mika Hämäläinen

UralicNLP is a natural language processing library for small Uralic languages. It can produce morphological analysis, generate morphological forms, lemmatize words and give lexical information about words in Uralic langu…

Morphological Analysis

A Broad-coverage Corpus for Finnish Named Entity Recognition

2020-05-01 · LREC 2020 5 · Jouni Luoma, Miika Oinonen, Maria Pyyk{\"o}nen, Veronika Laippala 외

We present a new manually annotated corpus for broad-coverage named entity recognition for Finnish. Building on the original Universal Dependencies Finnish corpus of 754 documents (200,000 tokens) representing ten differ…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER

FinChat: Corpus and evaluation setup for Finnish chat conversations on everyday topics

2020-08-19 · Katri Leino, Juho Leinonen, Mittul Singh, Sami Virpioja 외

Creating open-domain chatbots requires large amounts of conversational data and related benchmark tasks to evaluate them. Standardized evaluation tasks are crucial for creating automatic evaluation metrics for model deve…

ChatbotRetrieval