paper-with-me

홈 › Papers

Inter-Rater Agreement Study on Readability Assessment in Bengali

2014-07-08 · Shanta Phani, Shibamouli Lahiri, Arindam Biswas

An inter-rater agreement study is performed for readability assessment in Bengali. A 1-7 rating scale was used to indicate different levels of readability. We obtained moderate to fair agreement among seven independent annotators on 30 text passages written by four eminent Bengali authors. As a by product of our study, we obtained a readability-annotated ground truth dataset in Bengali. .

📄 PDF Abstract BibTeX arXiv:1407.1976

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Large and Balanced Corpus for Fine-grained Arabic Readability Assessment

2025-02-19 · Khalid N. Elmadani, Nizar Habash, Hanada Taha-Thomure

This paper introduces the Balanced Arabic Readability Evaluation Corpus BAREC, a large-scale, fine-grained dataset for Arabic readability assessment. BAREC consists of 68,182 sentences spanning 1+ million words, carefull…

Diversity

Event-Aligned Analysis of Multi-Rater Pain Assessments Using Continuous Wearable Physiology

2026-06-11 · Saba A. Farahani, Elahe Khatibi, Thomas D. Hughes, Ariana M. Nelson 외 arxiv

Pain is assessed differently by patients, nurses, and clinicians, yet most computational approaches assume a single ground-truth label - effectively ignoring who is doing the rating. We introduce a rater-aware, event-ali…

Predicting Disagreement with Human Raters in LLM-as-a-Judge Difficulty Assessment without Using Generation-Time Probability Signals

2026-05-12 · Yo Ehara arxiv

Automatic generation of educational materials using large language models (LLMs) is becoming increasingly common, but assigning difficulty levels to such materials still requires substantial human effort. LLM-as-a-Judge …

Automatic Construction of Large Readability Corpora

2016-12-01 · WS 2016 12 · Jorge Alberto Wagner Filho, Rodrigo Wilkens, Aline Villavicencio

This work presents a framework for the automatic construction of large Web corpora classified by readability level. We compare different Machine Learning classifiers for the task of readability assessment focusing on Por…

Text ClassificationText Simplification

Quantifying the Statistical Effect of Rubric Modifications on Human-Autorater Agreement

2026-05-07 · Jessica Huynh, Alfredo Gomez, Athiya Deviyani, Renee Shelby 외 arxiv

Autoraters, also referred to as LLM-as-judges, are increasingly used for evaluation and automated content moderation. However, there is limited statistical analysis of how modifications in a rubric presented to both huma…