paper-with-me

Papers

EXAMS: A Multi-Subject High School Examinations Dataset for Cross-Lingual and Multilingual Question Answering

2020-11-05 · EMNLP 2020 11 · Momchil Hardalov, Todor Mihaylov, Dimitrina Zlatkova, Yoan Dinkov, Ivan Koychev, Preslav Nakov

We propose EXAMS -- a new benchmark dataset for cross-lingual and multilingual question answering for high school examinations. We collected more than 24,000 high-quality high school exam questions in 16 languages, covering 8 language families and 24 school subjects from Natural Sciences and Social Sciences, among others. EXAMS offers a fine-grained evaluation framework across multiple languages and subjects, which allows precise analysis and comparison of various models. We perform various experiments with existing top-performing multilingual pre-trained models and we show that EXAMS offers multiple challenges that require multilingual knowledge and reasoning in multiple domains. We hope that EXAMS will enable researchers to explore challenging reasoning and knowledge transfer methods and pre-trained models for school question answering in various languages which was not possible before. The data, code, pre-trained models, and evaluation are available at https://github.com/mhardalov/exams-qa.

📄 PDF Abstract BibTeX arXiv:2011.03080

Code (2)

mhardalov/exams-qa 공식 구현 pytorch
checkstep/mole-stance pytorch

Tasks

Question AnsweringTransfer Learning

Similar Papers 제목 키워드 기반

LHMKE: A Large-scale Holistic Multi-subject Knowledge Evaluation Benchmark for Chinese Large Language Models

2024-03-19 · Chuang Liu, Renren Jin, Yuqi Ren, Deyi Xiong

Chinese Large Language Models (LLMs) have recently demonstrated impressive capabilities across various NLP benchmarks and real-world applications. However, the existing benchmarks for comprehensively evaluating these LLM…

Multiple-choice

IJCNLP-2017 Task 5: Multi-choice Question Answering in Examinations

2017-12-01 · IJCNLP 2017 12 · Shangmin Guo, Kang Liu, Shizhu He, Cao Liu 외

The IJCNLP-2017 Multi-choice Question Answering(MCQA) task aims at exploring the performance of current Question Answering(QA) techniques via the realworld complex questions collected from Chinese Senior High School Entr…

Question Answering

Examining Monitoring System: Detecting Abnormal Behavior In Online Examinations

2024-02-19 · Dinh An Ngo, Thanh Dat Nguyen, Thi Le Chi Dang, Huy Hoan Le 외

Cheating in online exams has become a prevalent issue over the past decade, especially during the COVID-19 pandemic. To address this issue of academic dishonesty, our "Exam Monitoring System: Detecting Abnormal Behavior …

Decision Making

Alvorada-Bench: Can Language Models Solve Brazilian University Entrance Exams?

2025-08-19 · Henrique Godoy arxiv

Language models are increasingly used in Brazil, but most evaluation remains English-centric. This paper presents Alvorada-Bench, a 4,515-question, text-only benchmark drawn from five Brazilian university entrance examin…

Reasoning Models Ace the CFA Exams

2025-12-09 · Jaisal Patel, Yunzhe Chen, Kaiwen He, Keyi Wang 외 arxiv

Previous research has reported that large language models (LLMs) demonstrate poor performance on the Chartered Financial Analyst (CFA) exams. However, recent reasoning models have achieved strong results on graduate-leve…