paper-with-me

홈 › Papers

BLUEX: A benchmark based on Brazilian Leading Universities Entrance eXams

2023-07-11 · Thales Sales Almeida, Thiago Laitz, Giovana K. Bonás, Rodrigo Nogueira

One common trend in recent studies of language models (LMs) is the use of standardized tests for evaluation. However, despite being the fifth most spoken language worldwide, few such evaluations have been conducted in Portuguese. This is mainly due to the lack of high-quality datasets available to the community for carrying out evaluations in Portuguese. To address this gap, we introduce the Brazilian Leading Universities Entrance eXams (BLUEX), a dataset of entrance exams from the two leading universities in Brazil: UNICAMP and USP. The dataset includes annotated metadata for evaluating the performance of NLP models on a variety of subjects. Furthermore, BLUEX includes a collection of recently administered exams that are unlikely to be included in the training data of many popular LMs as of 2023. The dataset is also annotated to indicate the position of images in each question, providing a valuable resource for advancing the state-of-the-art in multimodal language understanding and reasoning. We describe the creation and characteristics of BLUEX and establish a benchmark through experiments with state-of-the-art LMs, demonstrating its potential for advancing the state-of-the-art in natural language understanding and reasoning in Portuguese. The data and relevant code can be found at https://github.com/Portuguese-Benchmark-Datasets/BLUEX

📄 PDF Abstract BibTeX arXiv:2307.05410

Code (1)

portuguese-benchmark-datasets/bluex 공식 구현

Tasks

Natural Language Understanding

Similar Papers 제목 키워드 기반

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams

2026-06-21 · João Guilherme Alves Santos, Giovana Kerche Bonás, Thiago Laitz, Thales Sales Almeida 외 arxiv

Although Large Language Models (LLMs) excel in many tasks, their assessment in Portuguese has received less attention, particularly for open-ended, discursive tasks that demand deeper reasoning and generation capabilitie…

Mathematical Reasoning

Evaluating GPT-4's Vision Capabilities on Brazilian University Admission Exams

2023-11-23 · Ramon Pires, Thales Sales Almeida, Hugo Abonizio, Rodrigo Nogueira

Recent advancements in language models have showcased human-comparable performance in academic entrance exams. However, existing studies often overlook questions that require the integration of visual comprehension, thus…

Evaluating GPT-3.5 and GPT-4 Models on Brazilian University Admission Exams

2023-03-29 · Desnes Nunes, Ricardo Primi, Ramon Pires, Roberto Lotufo 외

The present study aims to explore the capabilities of Language Models (LMs) in tackling high-stakes multiple-choice tests, represented here by the Exame Nacional do Ensino M\'edio (ENEM), a multidisciplinary entrance exa…

Multiple-choice

Alvorada-Bench: Can Language Models Solve Brazilian University Entrance Exams?

2025-08-19 · Henrique Godoy arxiv

Language models are increasingly used in Brazil, but most evaluation remains English-centric. This paper presents Alvorada-Bench, a 4,515-question, text-only benchmark drawn from five Brazilian university entrance examin…

BLUEX Revisited: Enhancing Benchmark Coverage with Automatic Captioning

2025-08-29 · João Guilherme Alves Santos, Giovana Kerche Bonás, Thales Sales Almeida arxiv

With the growing capabilities of Large Language Models (LLMs), there is an increasing need for robust evaluation methods, especially in multilingual and non-English contexts. We present an updated version of the BLUEX da…