paper-with-me

Papers

The Write & Improve Corpus 2024: Error-annotated and CEFR-labelled essays by learners of English

2024-10-23 · Apollo - University of Cambridge Repository 2024 10 · Diane Nicholls, Andrew Caines, Paula Buttery

We present a new annotated corpus of written learner English, derived from essays submitted to the learning platform Write & Improve (W&I). Users of W&I are presented with automated scoring and feedback on grammatical errors, and are encouraged to act on their error feedback, submitting multiple versions of their essays for any given prompt. We build the corpus on this interplay between users and prompts, collecting sets of essays submitted by users for a selected list of 50 popular prompts. The prompts include 20 aimed at beginner learners of English, 20 aimed at intermediate learners, and 10 at advanced learners. This distribution reflects the greater use of W&I by beginner and intermediate learners of English. We ensured that the prompts were not likely to elicit personal information and covered a broad range of tasks and topics. This list of prompts enabled us to identify 5050 essay sets written by 766 users, forming the basis for the Write & Improve Corpus, which is being made available for non-commercial use on the ELiT website. We describe the steps we took to ensure the corpus contains appropriate texts, does not include personal information, and will come with annotations relating to Common European Framework of Reference (CEFR) level and grammatical errors. All essays were submitted between 2020 and 2022 by registered users of W&I who have supplied their first language (L1) in an optional questionnaire. In total, there are more than 23K essays containing more than 3.5 million word tokens. The final versions of each essay set amount to 762K word tokens. There are 22 different L1s in the corpus, with the most common being Spanish, Portuguese, Japanese, Arabic and Vietnamese. We present some descriptive statistics for the corpus, and consider some research use cases.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

DescriptiveGrammatical Error Correction

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

ZAEBUC: An Annotated Arabic-English Bilingual Writer Corpus

2022-06-01 · LREC 2022 6 · Nizar Habash, David Palfreyman

We present ZAEBUC, an annotated Arabic-English bilingual writer corpus comprising short essays by first-year university students at Zayed University in the United Arab Emirates. We describe and discuss the various guidel…

LemmatizationPart-Of-Speech TaggingPOSPOS Tagging

The MERLIN corpus: Learner language and the CEFR

2014-05-01 · LREC 2014 5 · Adriane Boyd, Jirka Hana, Lionel Nicolas, Detmar Meurers 외

The MERLIN corpus is a written learner corpus for Czech, German,and Italian that has been designed to illustrate the Common European Framework of Reference for Languages (CEFR) with authentic learner data. The corpus con…

Language AcquisitionLanguage IdentificationNative Language Identification

CEFR-Based Sentence Difficulty Annotation and Assessment

2022-10-21 · Yuki Arase, Satoru Uchida, Tomoyuki Kajiwara

Controllable text simplification is a crucial assistive technique for language learning and teaching. One of the primary factors hindering its advancement is the lack of a corpus annotated with sentence difficulty levels…

SentenceText Simplification

CEFR-Annotated WordNet: LLM-Based Proficiency-Guided Semantic Database for Language Learning

2025-10-21 · Masato Kikuchi, Masatsugu Ono, Toshioki Soga, Tetsu Tanabe 외 arxiv

Although WordNet is a valuable resource because of its structured semantic networks and extensive vocabulary, its fine-grained sense distinctions can be challenging for second-language learners. To address this issue, we…

Semantic Similarity

SVALex: a CEFR-graded Lexical Resource for Swedish Foreign and Second Language Learners

2016-05-01 · LREC 2016 5 · Thomas Fran{\c{c}}ois, Elena Volodina, Ildik{\'o} Pil{\'a}n, Ana{\"\i}s Tack

The paper introduces SVALex, a lexical resource primarily aimed at learners and teachers of Swedish as a foreign and second language that describes the distribution of 15,681 words and expressions across the Common Europ…