paper-with-me

홈 › Papers

Charting the European LLM Benchmarking Landscape: A New Taxonomy and a Set of Best Practices

2025-10-28 · Špela Vintar, Taja Kuzman Pungeršek, Mojca Brglez, Nikola Ljubešić arxiv

While new benchmarks for large language models (LLMs) are being developed continuously to catch up with the growing capabilities of new models and AI in general, using and evaluating LLMs in non-English languages remains a little-charted landscape. We give a concise overview of recent developments in LLM benchmarking, and then propose a new taxonomy for the categorization of benchmarks that is tailored to multilingual or non-English use scenarios. We further propose a set of best practices and quality standards that could lead to a more coordinated development of benchmarks for European languages. Among other recommendations, we advocate for a higher language and culture sensitivity of evaluation methods.

📄 PDF Abstract BibTeX arXiv:2510.24450

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

European Energy Vision 2060: Charting Diverse Pathways for Europe's Energy Transition

2025-01-22 · Mostafa Barani, Konstantin Löffler, Pedro Crespo del Granado, Nikita Moskalenko 외

Europe is warming at the fastest rate of all continents, experiencing a temperature increase of about 1{\deg}C higher than the corresponding global increase. Aiming to be the first climate-neutral continent by 2050 under…

Beyond the Hook: Predicting Billboard Hot 100 Chart Inclusion with Machine Learning from Streaming, Audio Signals, and Perceptual Features

2025-09-29 · Christos Mountzouris arxiv

The advent of digital streaming platforms have recently revolutionized the landscape of music industry, with the ensuing digitalization providing structured data collections that open new research avenues for investigati…

Charting and Navigating Hugging Face's Model Atlas

2025-03-13 · Eliahu Horwitz, Nitzan Kurer, Jonathan Kahana, Liel Amar 외

As there are now millions of publicly available neural networks, searching and analyzing large model repositories becomes increasingly important. Navigating so many models requires an atlas, but as most models are poorly…

model

ErrorMap and ErrorAtlas: Charting the Failure Landscape of Large Language Models

2026-01-22 · Shir Ashury-Tahan, Yifan Mai, Elron Bandel, Michal Shmueli-Scheuer 외 arxiv

Large Language Models (LLM) benchmarks tell us when models fail, but not why they fail. A wrong answer on a reasoning dataset may stem from formatting issues, calculation errors, or dataset noise rather than weak reasoni…

AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies

2024-06-25 · Yi Zeng, Kevin Klyman, Andy Zhou, Yu Yang 외

We present a comprehensive AI risk taxonomy derived from eight government policies from the European Union, United States, and China and 16 company policies worldwide, making a significant step towards establishing a uni…