paper-with-me

Papers

TallVocabL2Fi: A Tall Dataset of 15 Finnish L2 Learners’ Vocabulary

2022-06-01 · LREC 2022 6 · Frankie Robertson, Li-Hsin Chang, Sini Söyrinki

Previous work concerning measurement of second language learners has tended to focus on the knowledge of small numbers of words, often geared towards measuring vocabulary size. This paper presents a “tall” dataset containing information about a few learners’ knowledge of many words, suitable for evaluating Vocabulary Inventory Prediction (VIP) techniques, including those based on Computerised Adaptive Testing (CAT). In comparison to previous comparable datasets, the learners are from varied backgrounds, so as to reduce the risk of overfitting when used for machine learning based VIP. The dataset contains both a self-rating test and a translation test, used to derive a measure of reliability for learner responses. The dataset creation process is documented, and the relationship between variables concerning the participants, such as their completion time, their language ability level, and the triangulated reliability of their self-assessment responses, are analysed. The word list is constructed by taking into account the extensive derivation morphology of Finnish, and infrequent words are included in order to account for explanatory variables beyond word frequency.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Character-Word Compositional Neural Language Model for Finnish

2016-12-10 · Matti Lankinen, Hannes Heikinheimo, Pyry Takala, Tapani Raiko 외

Inspired by recent research, we explore ways to model the highly morphological Finnish language at the level of characters while maintaining the performance of word-level models. We propose a new Character-to-Word-to-Cha…

Language ModelingLanguage Modelling

Automatic Speech Recognition with Very Large Conversational Finnish and Estonian Vocabularies

2017-07-13 · Seppo Enarvi, Peter Smit, Sami Virpioja, Mikko Kurimo

Today, the vocabulary size for language models in large vocabulary speech recognition is typically several hundreds of thousands of words. While this is already sufficient in some applications, the out-of-vocabulary word…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Personalized Text Retrieval for Learners of Chinese as a Foreign Language

2018-08-01 · COLING 2018 8 · Chak Yan Yeung, John Lee

This paper describes a personalized text retrieval algorithm that helps language learners select the most suitable reading material in terms of vocabulary complexity. The user first rates their knowledge of a small set o…

Active LearningComplex Word IdentificationRetrievalText Retrieval

An Intelligent Testing Strategy for Vocabulary Assessment of Chinese Second Language Learners

2019-08-01 · WS 2019 8 · Wei Zhou, Renfen Hu, Feipeng Sun, Ronghuai Huang

Vocabulary is one of the most important parts of language competence. Testing of vocabulary knowledge is central to research on reading and language. However, it usually costs a large amount of time and human labor to bu…

Building an English Vocabulary Knowledge Dataset of Japanese English-as-a-Second-Language Learners Using Crowdsourcing

2018-05-01 · LREC 2018 5 · Yo Ehara