paper-with-me

홈 › Papers

R4C: A Benchmark for Evaluating RC Systems to Get the Right Answer for the Right Reason

2019-10-10 · ACL 2020 6 · Naoya Inoue, Pontus Stenetorp, Kentaro Inui

Recent studies have revealed that reading comprehension (RC) systems learn to exploit annotation artifacts and other biases in current datasets. This prevents the community from reliably measuring the progress of RC systems. To address this issue, we introduce R4C, a new task for evaluating RC systems' internal reasoning. R4C requires giving not only answers but also derivations: explanations that justify predicted answers. We present a reliable, crowdsourced framework for scalably annotating RC datasets with derivations. We create and publicly release the R4C dataset, the first, quality-assured dataset consisting of 4.6k questions, each of which is annotated with 3 reference derivations (i.e. 13.8k derivations). Experiments show that our automatic evaluation metrics using multiple reference derivations are reliable, and that R4C assesses different skills from an existing benchmark.

📄 PDF Abstract BibTeX arXiv:1910.04601

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-Hop Reading ComprehensionReading Comprehension

Similar Papers 제목 키워드 기반

Do LLMs Know to Respect Copyright Notice?

2024-11-02 · Jialiang Xu, Shenglan Li, Zhaozhuo Xu, Denghui Zhang

Prior study shows that LLMs sometimes generate content that violates copyright. In this paper, we study another important yet underexplored problem, i.e., will LLMs respect copyright information in user input, and behave…

Articles

Building and Evaluating Open-Domain Dialogue Corpora with Clarifying Questions

2021-09-13 · EMNLP 2021 11 · Mohammad Aliannejadi, Julia Kiseleva, Aleksandr Chuklin, Jeffrey Dalton 외

Enabling open-domain dialogue systems to ask clarifying questions when appropriate is an important direction for improving the quality of the system response. Namely, for cases when a user request is not specific enough …

Evaluating Investment Logic in Large Language Models: A Real-World Benchmark Towards Personalzied Financial Agents

2026-08-06 · Yuanhong Jiang, Jingjie Zou, Zhenghong Lin, Xusheng Yu 외 arxiv

Investment competence is inherently personalized: the same market evidence can justify different actions for investors with different goals, horizons, portfolios, and risk boundaries. Yet financial LLMs are evaluated eit…

Question Answering

CLR-Bench: Evaluating Large Language Models in College-level Reasoning

2024-10-23 · Junnan Dong, Zijin Hong, Yuanchen Bei, Feiran Huang 외

Large language models (LLMs) have demonstrated their remarkable performance across various language understanding tasks. While emerging benchmarks have been proposed to evaluate LLMs in various domains such as mathematic…

COLUMBUS: Evaluating COgnitive Lateral Understanding through Multiple-choice reBUSes

2024-09-06 · Koen Kraaijveld, Yifan Jiang, Kaixin Ma, Filip Ilievski

While visual question-answering (VQA) benchmarks have catalyzed the development of reasoning techniques, they have focused on vertical thinking. Effective problem-solving also necessitates lateral thinking, which remains…

Multiple-choiceQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)