paper-with-me

Papers

A Systematic Classification of Knowledge, Reasoning, and Context within the ARC Dataset

2018-06-01 · WS 2018 7 · Michael Boratko, Harshit Padigela, Divyendra Mikkilineni, Pritish Yuvraj, Rajarshi Das, Andrew McCallum, Maria Chang, Achille Fokoue-Nkoutche, Pavan Kapanipathi, Nicholas Mattei, Ryan Musa, Kartik Talamadupula, Michael Witbrock

The recent work of Clark et al. introduces the AI2 Reasoning Challenge (ARC) and the associated ARC dataset that partitions open domain, complex science questions into an Easy Set and a Challenge Set. That paper includes an analysis of 100 questions with respect to the types of knowledge and reasoning required to answer them; however, it does not include clear definitions of these types, nor does it offer information about the quality of the labels. We propose a comprehensive set of definitions of knowledge and reasoning types necessary for answering the questions in the ARC dataset. Using ten annotators and a sophisticated annotation interface, we analyze the distribution of labels across the Challenge Set and statistics related to them. Additionally, we demonstrate that although naive information retrieval methods return sentences that are irrelevant to answering the query, sufficient supporting text is often present in the (ARC) corpus. Evaluating with human-selected relevant sentences improves the performance of a neural machine comprehension model by 42 points.

📄 PDF Abstract BibTeX arXiv:1806.00358

Code (0)

등록된 구현이 없습니다.

Tasks

AI2 Reasoning ChallengeARCGeneral ClassificationInformation RetrievalReading ComprehensionRetrieval

Similar Papers 제목 키워드 기반

A MIND for Reasoning: Meta-learning for In-context Deduction

2025-05-20 · Leonardo Bertolazzi, Manuel Vargas Guzmán, Raffaella Bernardi, Maciej Malicki 외

Large language models (LLMs) are increasingly evaluated on formal tasks, where strong reasoning abilities define the state of the art. However, their ability to generalize to out-of-distribution problems remains limited.…

Meta-Learning

Ontologies in Digital Twins: A Systematic Literature Review

2023-08-29 · Erkan Karabulut, Salvatore F. Pileggi, Paul Groth, Victoria Degeler

Digital Twins (DT) facilitate monitoring and reasoning processes in cyber-physical systems. They have progressively gained popularity over the past years because of intense research activity and industrial advancements. …

ArticlesKnowledge GraphsSystematic Literature Review

Exploring Multi-Temperature Strategies for Token- and Rollout-Level Control in RLVR

2025-10-10 · Haomin Zhuang, Yujun Zhou, Taicheng Guo, Yue Huang 외 arxiv

Reinforcement Learning has demonstrated substantial improvements in the reasoning abilities of Large Language Models (LLMs), exhibiting significant applicability across various domains. Recent research has identified tha…

Reinforcement Learning

MediEval: A Unified Medical Benchmark for Patient-Contextual and Knowledge-Grounded Reasoning in LLMs

2025-12-23 · Zhan Qu, Michael Färber arxiv

Large Language Models (LLMs) are increasingly applied to medicine, yet their adoption is limited by concerns over reliability and safety. Existing evaluations either test factual medical knowledge in isolation or assess …

Large Language Models are Limited in Out-of-Context Knowledge Reasoning

2024-06-11 · Peng Hu, Changjiang Gao, Ruiqi Gao, Jiajun Chen 외

Large Language Models (LLMs) possess extensive knowledge and strong capabilities in performing in-context reasoning. However, previous work challenges their out-of-context reasoning ability, i.e., the ability to infer in…

AttributeLogical ReasoningRetrievalTransfer Learning