paper-with-me

Papers

ORB: An Open Reading Benchmark for Comprehensive Evaluation of Machine Reading Comprehension

2019-12-29 · Dheeru Dua, Ananth Gottumukkala, Alon Talmor, Sameer Singh, Matt Gardner

Reading comprehension is one of the crucial tasks for furthering research in natural language understanding. A lot of diverse reading comprehension datasets have recently been introduced to study various phenomena in natural language, ranging from simple paraphrase matching and entity typing to entity tracking and understanding the implications of the context. Given the availability of many such datasets, comprehensive and reliable evaluation is tedious and time-consuming for researchers working on this problem. We present an evaluation server, ORB, that reports performance on seven diverse reading comprehension datasets, encouraging and facilitating testing a single model's capability in understanding a wide variety of reading phenomena. The evaluation server places no restrictions on how models are trained, so it is a suitable test bed for exploring training paradigms and representation learning for general reading facility. As more suitable datasets are released, they will be added to the evaluation server. We also collect and include synthetic augmentations for these datasets, testing how well models can handle out-of-domain questions.

📄 PDF Abstract BibTeX arXiv:1912.12598

Code (0)

등록된 구현이 없습니다.

Tasks

Entity TypingMachine Reading ComprehensionNatural Language UnderstandingReading ComprehensionRepresentation Learning

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

MRCEval: A Comprehensive, Challenging and Accessible Machine Reading Comprehension Benchmark

2025-03-10 · Shengkun Ma, Hao Peng, Lei Hou, Juanzi Li

Machine Reading Comprehension (MRC) is an essential task in evaluating natural language understanding. Existing MRC datasets primarily assess specific aspects of reading comprehension (RC), lacking a comprehensive MRC be…

Machine Reading ComprehensionNatural Language UnderstandingReading Comprehension

A Survey on Machine Reading Comprehension: Tasks, Evaluation Metrics and Benchmark Datasets

2020-06-21 · Changchang Zeng, Shaobo Li, Qin Li, Jie Hu 외

Machine Reading Comprehension (MRC) is a challenging Natural Language Processing(NLP) research field with wide real-world applications. The great progress of this field in recent years is mainly due to the emergence of l…

Machine Reading ComprehensionReading Comprehension

A Comprehensive Survey on Multi-hop Machine Reading Comprehension Datasets and Metrics

2022-12-08 · Azade Mohammadi, Reza Ramezani, Ahmad Baraani

Multi-hop Machine reading comprehension is a challenging task with aim of answering a question based on disjoint pieces of information across the different passages. The evaluation metrics and datasets are a vital part o…

Machine Reading ComprehensionReading Comprehension

A Survey on Explainability in Machine Reading Comprehension

2020-10-01 · Mokanarangan Thayaparan, Marco Valentino, André Freitas

This paper presents a systematic review of benchmarks and approaches for explainability in Machine Reading Comprehension (MRC). We present how the representation and inference challenges evolved and the steps which were …

Machine Reading ComprehensionReading ComprehensionSurvey

FunBench: Benchmarking Fundus Reading Skills of MLLMs

2025-03-02 · Qijie Wei, Kaiheng Qian, Xirong Li

Multimodal Large Language Models (MLLMs) have shown significant potential in medical image analysis. However, their capabilities in interpreting fundus images, a critical skill for ophthalmology, remain under-evaluated. …

AnatomyBenchmarkingLanguage ModelingLanguage Modelling+5