Complex Word Identification Based on Frequency in a Learner Corpus
We introduce the TMU systems for the Complex Word Identification (CWI) Shared Task 2018. TMU systems use random forest classifiers and regressors whose features are the number of characters, the number of words, and the frequency of target words in various corpora. Our simple systems performed best on 5 tracks out of 12 tracks. Our ablation analysis revealed the usefulness of a learner corpus for CWI task.
Code (0)
등록된 구현이 없습니다.
Tasks
Complex Word IdentificationLexical SimplificationReading ComprehensionText SimplificationSimilar Papers 제목 키워드 기반
CLexIS2: A New Corpus for Complex Word Identification Research in Computing Studies
Reading is a complex process not only because of the words or sections that are difficult for the reader to understand. Complex word identification (CWI) is the task of detecting in the content of documents the words tha…
Complex Word IdentificationLexical SimplificationSeCoDa: Sense Complexity Dataset
The Sense Complexity Dataset (SeCoDa) provides a corpus that is annotated jointly for complexity and word senses. It thus provides a valuable resource for both word sense disambiguation and the task of complex word ident…
Complex Word IdentificationWord Sense DisambiguationMulti-task Learning for Chinese Word Usage Errors Detection
Chinese word usage errors often occur in non-native Chinese learners' writing. It is very helpful for non-native Chinese learners to detect them automatically when learning writing. In this paper, we propose a novel appr…
Multi-Task LearningPOSPOS TaggingPredictionComplex Word Identification in Vietnamese: Towards Vietnamese Text Simplification
Text Simplification has been an extensively researched problem in English, but has not been investigated in Vietnamese. We focus on the Vietnamese-specific Complex Word Identification task, often the first step in Lexica…
Complex Word IdentificationLexical SimplificationText SimplificationVietnamese DatasetsAn evaluation of the role of statistical measures and frequency for MWE identification
We report on an experiment to evaluate the role of statistical association measures and frequency for the identification of MWE. We base our evaluation on a lexicon of 14.000 MWE comprising different types of word combin…