paper-with-me

Papers

SuperGLEBer: German Language Understanding Evaluation Benchmark

2024-06-20 · NAACL 2024 6 · Jan Pfister, Andreas Hotho

We assemble a broad Natural Language Understanding benchmark suite for the German language and consequently evaluate a wide array of existing German-capable models in order to create a better understanding of the current state of German LLMs. Our benchmark consists of 29 different tasks ranging over different types such as document classification, sequence tagging, sentence similarity, and question answering, on which we evaluate 10 different German-pretrained models, thereby charting the landscape of German LLMs. In our comprehensive evaluation we find that encoder models are a good choice for most tasks, but also that the largest encoder model does not necessarily perform best for all tasks. We make our benchmark suite and a leaderboard publically available at https://supergleber.professor-x.de and encourage the community to contribute new tasks and evaluate more models on it (https://github.com/LSX-UniWue/SuperGLEBer).

📄 PDF Abstract BibTeX

Code (1)

LSX-UniWue/SuperGLEBer pytorch

Tasks

Document ClassificationNatural Language UnderstandingQuestion AnsweringSentenceSentence Similarity

Similar Papers 제목 키워드 기반

LLäMmlein: Compact and Competitive German-Only Language Models from Scratch

2024-11-17 · Jan Pfister, Julia Wunderle, Andreas Hotho

We create two German-only decoder models, LL\"aMmlein 120M and 1B, transparently from scratch and publish them, along with the training data, for the German NLP research community to use. The model training involved seve…

Decoder

AI-assisted German Employment Contract Review: A Benchmark Dataset

2025-01-27 · Oliver Wardas, Florian Matthes

Employment contracts are used to agree upon the working conditions between employers and employees all over the world. Understanding and reviewing contracts for void or unfair clauses requires extensive knowledge of the …

Fairness

Encoder vs Decoder: Comparative Analysis of Encoder and Decoder Language Models on Multilingual NLU Tasks

2024-06-19 · Dan Saattrup Nielsen, Kenneth Enevoldsen, Peter Schneider-Kamp

This paper explores the performance of encoder and decoder language models on multilingual Natural Language Understanding (NLU) tasks, with a broad focus on Germanic languages. Building upon the ScandEval benchmark, init…

DecoderLanguage ModelingLanguage ModellingModel Selection+1

VLM@school -- Evaluation of AI image understanding on German middle school knowledge

2025-06-13 · René Peinl, Vincent Tischler

This paper introduces a novel benchmark dataset designed to evaluate the capabilities of Vision Language Models (VLMs) on tasks that combine visual reasoning with subject-specific background knowledge in the German langu…

Visual Reasoning

M4U: Evaluating Multilingual Understanding and Reasoning for Large Multimodal Models

2024-05-24 · Hongyu Wang, Jiayu Xu, Senwei Xie, Ruiping Wang 외

Multilingual multimodal reasoning is a core component in achieving human-level intelligence. However, most existing benchmarks for multilingual multimodal reasoning struggle to differentiate between models of varying per…

Multimodal Reasoning