paper-with-me

홈 › Papers

From 'F' to 'A' on the N.Y. Regents Science Exams: An Overview of the Aristo Project

2019-09-04 · Peter Clark, Oren Etzioni, Daniel Khashabi, Tushar Khot, Bhavana Dalvi Mishra, Kyle Richardson, Ashish Sabharwal, Carissa Schoenick, Oyvind Tafjord, Niket Tandon, Sumithra Bhakthavatsalam, Dirk Groeneveld, Michal Guerquin, Michael Schmitz

AI has achieved remarkable mastery over games such as Chess, Go, and Poker, and even Jeopardy, but the rich variety of standardized exams has remained a landmark challenge. Even in 2016, the best AI system achieved merely 59.3% on an 8th Grade science exam challenge. This paper reports unprecedented success on the Grade 8 New York Regents Science Exam, where for the first time a system scores more than 90% on the exam's non-diagram, multiple choice (NDMC) questions. In addition, our Aristo system, building upon the success of recent language models, exceeded 83% on the corresponding Grade 12 Science Exam NDMC questions. The results, on unseen test questions, are robust across different test years and different variations of this kind of test. They demonstrate that modern NLP methods can result in mastery on this task. While not a full solution to general question-answering (the questions are multiple choice, and the domain is restricted to 8th Grade science), it represents a significant milestone for the field.

📄 PDF Abstract BibTeX arXiv:1909.01958

Code (0)

등록된 구현이 없습니다.

Tasks

Multiple-choiceQuestion Answering

Similar Papers 제목 키워드 기반

The Limitations of Standardized Science Tests as Benchmarks for Artificial Intelligence Research: Position Paper

2014-11-06 · Ernest Davis

In this position paper, I argue that standardized tests for elementary science such as SAT or Regents tests are not very good benchmarks for measuring the progress of artificial intelligence systems in understanding basi…

Position

Assessing the Comprehensibility of Automatic Translations (ArisToCAT)

2020-11-01 · EAMT 2020 11 · Lieve Macken, Margot Fonteyne, Arda Tezcan, Joke Daems

The ArisToCAT project aims to assess the comprehensibility of ‘raw’ (unedited) MT output for readers who can only rely on the MT output. In this project description, we summarize the main results of the project and prese…

ReGentS: Real-World Safety-Critical Driving Scenario Generation Made Stable

2024-09-12 · Yuan Yin, Pegah Khayatan, Éloi Zablocki, Alexandre Boulch 외

Machine learning based autonomous driving systems often face challenges with safety-critical scenarios that are rare in real-world data, hindering their large-scale deployment. While increasing real-world training data c…

Autonomous Driving

Humans Keep It One Hundred: an Overview of AI Journey

2020-05-01 · LREC 2020 5 · Tatiana Shavrina, Anton Emelyanov, Alena Fenogenova, Vadim Fomin 외

Artificial General Intelligence (AGI) is showing growing performance in numerous applications - beating human performance in Chess and Go, using knowledge bases and text sources to answer questions (SQuAD) and even pass …

Text Generation

EXAMS: A Multi-Subject High School Examinations Dataset for Cross-Lingual and Multilingual Question Answering

2020-11-05 · EMNLP 2020 11 · Momchil Hardalov, Todor Mihaylov, Dimitrina Zlatkova, Yoan Dinkov 외

We propose EXAMS -- a new benchmark dataset for cross-lingual and multilingual question answering for high school examinations. We collected more than 24,000 high-quality high school exam questions in 16 languages, cover…

Question AnsweringTransfer Learning