paper-with-me

홈 › Papers

DUMB: A Benchmark for Smart Evaluation of Dutch Models

2023-05-22 · Wietse de Vries, Martijn Wieling, Malvina Nissim

We introduce the Dutch Model Benchmark: DUMB. The benchmark includes a diverse set of datasets for low-, medium- and high-resource tasks. The total set of nine tasks includes four tasks that were previously not available in Dutch. Instead of relying on a mean score across tasks, we propose Relative Error Reduction (RER), which compares the DUMB performance of language models to a strong baseline which can be referred to in the future even when assessing different sets of language models. Through a comparison of 14 pre-trained language models (mono- and multi-lingual, of varying sizes), we assess the internal consistency of the benchmark tasks, as well as the factors that likely enable high performance. Our results indicate that current Dutch monolingual models under-perform and suggest training larger Dutch models with other architectures and pre-training objectives. At present, the highest performance is achieved by DeBERTaV3 (large), XLM-R (large) and mDeBERTaV3 (base). In addition to highlighting best strategies for training larger Dutch models, DUMB will foster further research on Dutch. A public leaderboard is available at https://dumbench.nl.

📄 PDF Abstract BibTeX arXiv:2305.13026

Code (2)

wietsedv/dumb 공식 구현
rijgersberg/geitje pytorch

Tasks

XLM-R

Methods 이 논문이 사용한 방법론

XLM-R XLM-R

Similar Papers 제목 키워드 기반

Four Ways to Scale Up: Smart, Dumb, Forced, and Fumbled

2021-01-13 · Bent Flyvbjerg

Scale-up is the process of growing a venture in size. The paper identifies modularity and speed as keys to successful scale-up. On that basis four types of scale-up are identified: Smart, dumb, forced, and fumbled. Smart…

A Dutch Financial Large Language Model

2024-10-03 · Sander Noels, Jorne De Blaere, Tijl De Bie

This paper presents FinGEITje, the first Dutch financial Large Language Model (LLM) specifically designed and optimized for various financial tasks. Together with the model, we release a specialized Dutch financial instr…

Language ModelingLanguage ModellingLarge Language Modelmodel

RDumb++: Drift-Aware Continual Test-Time Adaptation

2026-01-22 · Himanshu Mishra arxiv

Continual Test-Time Adaptation (CTTA) seeks to update a pretrained model during deployment using only the incoming, unlabeled data stream. Although prior approaches such as Tent, EATA etc. provide meaningful improvements…

Test-time Adaptation

MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch

2025-09-15 · Nikolay Banar, Ehsan Lotfi, Jens Van Nooten, Cristina Arhiliuc 외 arxiv

Recently, embedding resources, including models, benchmarks, and datasets, have been widely released to support a variety of languages. However, the Dutch language remains underrepresented, typically comprising only a sm…

High‐purity Solid SiO2 Nanodumbbells

2021-01-06 · Silicon

Abstract:Optically levitated nanodumbbells in vacuum are excellent candidates for thermodynamics, macroscopic quantum mechanics, precision measurements and quantum sensing. Silica (SiO2) material, with extremely low abso…

Vocal Bursts Intensity Prediction