paper-with-me

Papers

Ace-CEFR -- A Dataset for Automated Evaluation of the Linguistic Difficulty of Conversational Texts for LLM Applications

2025-06-16 · David Kogan, Max Schumacher, Sam Nguyen, Masanori Suzuki, Melissa Smith, Chloe Sophia Bellows, Jared Bernstein

There is an unmet need to evaluate the language difficulty of short, conversational passages of text, particularly for training and filtering Large Language Models (LLMs). We introduce Ace-CEFR, a dataset of English conversational text passages expert-annotated with their corresponding level of text difficulty. We experiment with several models on Ace-CEFR, including Transformer-based models and LLMs. We show that models trained on Ace-CEFR can measure text difficulty more accurately than human experts and have latency appropriate to production environments. Finally, we release the Ace-CEFR dataset to the public for research and development.

📄 PDF Abstract BibTeX arXiv:2506.14046

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

UniversalCEFR: Enabling Open Multilingual Research on Language Proficiency Assessment

2025-06-02 · Joseph Marvin Imperial, Abdullah Barayan, Regina Stodden, Rodrigo Wilkens 외

We introduce UniversalCEFR, a large-scale multilingual multidimensional dataset of texts annotated according to the CEFR (Common European Framework of Reference) scale in 13 languages. To enable open research in both aut…

EvalYaks: Instruction Tuning Datasets and LoRA Fine-tuned Models for Automated Scoring of CEFR B2 Speaking Assessment Transcripts

2024-08-22 · Nicy Scaria, Silvester John Joseph Kennedy, Thomas Latinovich, Deepak Subramani

Relying on human experts to evaluate CEFR speaking assessments in an e-learning environment creates scalability challenges, as it limits how quickly and widely assessments can be conducted. We aim to automate the evaluat…

Human and Automated CEFR-based Grading of Short Answers

2017-09-01 · WS 2017 9 · Ana{\"\i}s Tack, Thomas Fran{\c{c}}ois, Sophie Roekhaut, C{\'e}drick Fairon

This paper is concerned with the task of automatically assessing the written proficiency level of non-native (L2) learners of English. Drawing on previous research on automated L2 writing assessment following the Common …

ComplexityMT: Benchmarking the Interaction Between Text Complexity and Machine Translation

2026-06-03 · Joseph Marvin Imperial, Junhong Liang, Belal Shoer, Abdullah Barayan 외 arxiv

When a text is translated, does the translation retain the complexity of the original? We introduce ComplexityMT, a new challenge for assessing how text complexity and machine translation interact with and influence each…

Machine Translation

Can LLMs Control Readability? A Multi-Dimensional Evaluation Framework for CEFR-Controlled Arabic Generation

2026-06-20 · Nour Rabih, Chatrine Qwaider, Ted Briscoe arxiv

While Large Language Models (LLMs) can generate fluent Arabic text, their ability to reliably control readability levels remains unclear. We propose a multi-dimensional evaluation framework for Common European Framework …

Text Generation