CWIG3G2 - Complex Word Identification Task across Three Text Genres and Two User Groups
Complex word identification (CWI) is an important task in text accessibility. However, due to the scarcity of CWI datasets, previous studies have only addressed this problem on Wikipedia sentences and have solely taken into account the needs of non-native English speakers. We collect a new CWI dataset (CWIG3G2) covering three text genres News, WikiNews, and Wikipedia) annotated by both native and non-native English speakers. Unlike previous datasets, we cover single words, as well as complex phrases, and present them for judgment in a paragraph context. We present the first study on cross-genre and cross-group CWI, showing measurable influences in native language and genre types.
Code (0)
등록된 구현이 없습니다.
Tasks
Complex Word IdentificationLexical SimplificationReading ComprehensionSimilar Papers 제목 키워드 기반
Cross-lingual complex word identification with multitask learning
We approach the 2018 Shared Task on Complex Word Identification by leveraging a cross-lingual multitask learning approach. Our method is highly language agnostic, as evidenced by the ability of our system to generalize a…
Complex Word IdentificationLexical SimplificationManchester Metropolitan at SemEval-2021 Task 1: Convolutional Networks for Complex Word Identification
We present two convolutional neural networks for predicting the complexity of words and phrases in context on a continuous scale. Both models utilize word and character embeddings alongside lexical features as inputs. Ou…
Complex Word IdentificationregressionComplex Word Identification in Vietnamese: Towards Vietnamese Text Simplification
Text Simplification has been an extensively researched problem in English, but has not been investigated in Vietnamese. We focus on the Vietnamese-specific Complex Word Identification task, often the first step in Lexica…
Complex Word IdentificationLexical SimplificationText SimplificationVietnamese DatasetsComplex Word Identification Based on Frequency in a Learner Corpus
We introduce the TMU systems for the Complex Word Identification (CWI) Shared Task 2018. TMU systems use random forest classifiers and regressors whose features are the number of characters, the number of words, and the …
Complex Word IdentificationLexical SimplificationReading ComprehensionText SimplificationCLaC at SemEval-2016 Task 11: Exploring linguistic and psycho-linguistic Features for Complex Word Identification
This paper describes the system deployed by the CLaC-EDLK team to the "SemEval 2016, Complex Word Identification task". The goal of the task is to identify if a given word in a given context is "simple" or "complex". Our…
Complex Word Identification