Reading Comprehension
7개 벤치마크 · 논문 1,821편 · 이 태스크의 논문 보기 →
Benchmarks
ReClor
RACE
MuSeRC
AdversarialQA
CrowdSource QA
RadQA
ReCAM
Most implemented
Language Models are Few-Shot Learners
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Listen, Attend and Spell
XLNet: Generalized Autoregressive Pretraining for Language Understanding
Bidirectional Attention Flow for Machine Comprehension
Language Models are Unsupervised Multitask Learners
Papers
Language Proficiency Assessment from Eye Movements in Naturalistic Passage Reading
Standard language proficiency tests rely on linguistic tasks such as vocabulary, grammar and reading comprehension quizzes. An alternative, cognitively motivated approach, introduced in Berzak et al. (2018), proposed ins…
Reading ComprehensionArkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer
We present Arkios, a 1.04B-parameter dense transformer pretrained from scratch on 150B tokens of bilingual English-Nepali text, using a custom single-file C/CUDA training stack and a Devanagari-aware byte-level BPE token…
Reading ComprehensionLEXIC: Lightweight Eye-tracking eXtension via Injected Complexity
On the recent EyeBench benchmark, predicting reading comprehension from eye movements exposes a stark gap: text-aware models using pretrained language models reach 56--63% AUROC, while gaze-only models operate at chance.…
Reading ComprehensionLLMs Struggle to Measure What Distinguishes Students of Different Proficiency Levels: A Study of Item Discrimination in Reading Comprehension Assessment
Item discrimination is a fundamental psychometric property of educational assessment, which measures whether an item meaningfully distinguishes students with higher proficiency from students with lower proficiency. While…
Reading ComprehensionModeling semantic association in self-paced reading with language model embeddings
Semantic association between a word and its context has been identified as an important component of reading comprehension, even when word predictability is accounted for. Recent research has highlighted the potential of…
Reading ComprehensionMemoryDocDataSet: A Benchmark for Joint Conversational Memory and Long Document Reasoning
AI systems increasingly need to combine two demanding capabilities: navigating multi-session conversation history and performing deep reading comprehension within long documents. Yet no existing benchmark evaluates both …
Reading Comprehension