paper-with-me

Papers

The Invalsi Benchmarks: measuring Linguistic and Mathematical understanding of Large Language Models in Italian

2024-03-27 · Giovanni Puccetti, Maria Cassese, Andrea Esuli

While Italian is a high-resource language, there are few Italian-native benchmarks to evaluate generative Large Language Models (LLMs) in this language. This work presents three new benchmarks: Invalsi MATE to evaluate models performance on mathematical understanding in Italian, Invalsi ITA to evaluate language understanding in Italian and Olimpiadi MATE for more complex mathematical understanding. The first two benchmarks are based on the Invalsi tests, which are administered to students of age between 6 and 18 within the Italian school system and have been validated by several experts in teaching and pedagogy, the third one comes from the Italian high school math Olympics. We evaluate 10 powerful language models on these benchmarks and find that they are bound by 71% accuracy on Invasli MATE, achieved by Llama 3.1 70b instruct and by 88% on Invalsi ITA. For both Invalsi MATE and Invalsi ITA we compare LLMs with the average performance of Italian students to show that Llama 3.1 is the only one to outperform them on Invalsi MATE while most models do so on Invalsi ITA, we then show that Olimpiadi MATE is more challenging than Invalsi MATE and the highest accuracy, achieved by Llama 3.1 405b instruct is 45%. We will make data and evaluation code openly available upon acceptance of the paper.

📄 PDF Abstract BibTeX arXiv:2403.18697

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModellingMath

Methods 이 논문이 사용한 방법론

MATE MATE is a Transformer architecture designed to model the structure of web tables. It uses sparse attention in a way that…
LLaMA LLaMA is a collection of foundation language models ranging from 7B to 65B parameters. It is based on the transformer architecture with various improvements that were…

Similar Papers 제목 키워드 기반

Disce aut Deficere: Evaluating LLMs Proficiency on the INVALSI Italian Benchmark

2024-06-25 · Fabio Mercorio, Mario Mezzanzanica, Daniele Potertì, Antonio Serino 외

Recent advancements in Large Language Models (LLMs) have significantly enhanced their ability to generate and manipulate human language, highlighting their potential across various applications. Evaluating LLMs in langua…

The Qiyas Benchmark: Measuring ChatGPT Mathematical and Language Understanding in Arabic

2024-06-28 · Shahad Al-Khalifa, Hend Al-Khalifa

Despite the growing importance of Arabic as a global language, there is a notable lack of language models pre-trained exclusively on Arabic data. This shortage has led to limited benchmarks available for assessing langua…

Language ModelingLanguage ModellingMathematical Reasoning

Language Model Metrics and Procrustes Analysis for Improved Vector Transformation of NLP Embeddings

2021-06-04 · ICON 2020 12 · Thomas Conley, Jugal Kalita

Artificial Neural networks are mathematical models at their core. This truismpresents some fundamental difficulty when networks are tasked with Natural Language Processing. A key problem lies in measuring the similarity …

Language ModelingLanguage Modelling

TokEval: A Tokenizer Evaluation Suite

2026-08-18 · Clara Meister arxiv

Language model tokenizers are typically selected with minimal evaluation, despite the fact that their design choices directly impact model capabilities. This can be partly attributed to a limited understanding of which t…

Mathematical ReasoningCode Generation

How and where does CLIP process negation?

2024-07-15 · Vincent Quantmeyer, Pablo Mosteiro, Albert Gatt

Various benchmarks have been proposed to test linguistic understanding in pre-trained vision \& language (VL) models. Here we build on the existence task from the VALSE benchmark (Parcalabescu et al, 2022) which we use t…

Language ModellingNegation