paper-with-me

홈 › Papers

FinChat: Corpus and evaluation setup for Finnish chat conversations on everyday topics

2020-08-19 · Katri Leino, Juho Leinonen, Mittul Singh, Sami Virpioja, Mikko Kurimo

Creating open-domain chatbots requires large amounts of conversational data and related benchmark tasks to evaluate them. Standardized evaluation tasks are crucial for creating automatic evaluation metrics for model development; otherwise, comparing the models would require resource-expensive human evaluation. While chatbot challenges have recently managed to provide a plethora of such resources for English, resources in other languages are not yet available. In this work, we provide a starting point for Finnish open-domain chatbot research. We describe our collection efforts to create the Finnish chat conversation corpus FinChat, which is made available publicly. FinChat includes unscripted conversations on seven topics from people of different ages. Using this corpus, we also construct a retrieval-based evaluation task for Finnish chatbot development. We observe that off-the-shelf chatbot models trained on conversational corpora do not perform better than chance at choosing the right answer based on automatic metrics, while humans can do the same task almost perfectly. Similarly, in a human evaluation, responses to questions from the evaluation set generated by the chatbots are predominantly marked as incoherent. Thus, FinChat provides a challenging evaluation set, meant to encourage chatbot development in Finnish.

📄 PDF Abstract BibTeX arXiv:2008.08315

Code (1)

aalto-speech/FinChat 공식 구현 pytorch

Tasks

ChatbotRetrieval

Similar Papers 제목 키워드 기반

A Broad-coverage Corpus for Finnish Named Entity Recognition

2020-05-01 · LREC 2020 5 · Jouni Luoma, Miika Oinonen, Maria Pyyk{\"o}nen, Veronika Laippala 외

We present a new manually annotated corpus for broad-coverage named entity recognition for Finnish. Building on the original Universal Dependencies Finnish corpus of 754 documents (200,000 tokens) representing ten differ…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER

Quantitative Evaluation of Alternative Translations in a Corpus of Highly Dissimilar Finnish Paraphrases

2021-05-06 · MoTra (NoDaLiDa) 2021 5 · Li-Hsin Chang, Sampo Pyysalo, Jenna Kanerva, Filip Ginter

In this paper, we present a quantitative evaluation of differences between alternative translations in a large recently released Finnish paraphrase corpus focusing in particular on non-trivial variation in translation. W…

Translation

The Corpus of Finnish Sign Language

2020-05-01 · LREC 2020 5 · Juhana Salonen, Antti Kronqvist, Tommi Jantunen

This paper presents the Corpus of Finnish Sign Language (Corpus FinSL), a structured and annotated collection of Finnish Sign Language (FinSL) videos published in May 2019 in FIN-CLARIN{'}s Language Bank of Finland. The …

Dialect Text Normalization to Normative Standard Finnish

2019-11-01 · WS 2019 11 · Niko Partanen, Mika H{\"a}m{\"a}l{\"a}inen, Khalid Alnajjar

We compare different LSTMs and transformer models in terms of their effectiveness in normalizing dialectal Finnish into the normative standard Finnish. As dialect is the common way of communication for people online in F…

Text Normalization

Out-of-Domain Evaluation of Finnish Dependency Parsing

2022-04-22 · LREC 2022 6 · Jenna Kanerva, Filip Ginter

The prevailing practice in the academia is to evaluate the model performance on in-domain evaluation data typically set aside from the training corpus. However, in many real world applications the data on which the model…

Dependency Parsing