paper-with-me

홈 › Papers

From Global Benchmarks to Local Evaluations: Benchmarking LLMs for the German Public Sector

2026-08-18 · Camilla Dalerci, Thilo Michael, Robin Schaefer, Daniel Weinland arxiv

Public institutions face a persistent challenge in selecting LLMs suited to their specific context. Existing benchmarks, however, are of limited use as they primarily reflect English-language and US-centric settings, and often only evaluate task performance. In this paper, we present first results of MÖVE, a holistic evaluation framework for the German public sector, examining three rarely considered governance dimensions: energy consumption, provider transparency, and knowledge of German-party positions. Our results reveal significant trade-offs, with no single model excelling across all dimensions: estimated energy consumption varies more than 60-fold and is not explained by model size alone, information disclosure varies systematically across providers, and European models do not exhibit stronger knowledge of German party positions. Model selection for public institutions thus cannot rely on performance rankings alone. Instead, evaluations should also reflect the governance requirements of the deployment context.

📄 PDF Abstract BibTeX arXiv:2608.17827

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints

2026-01-26 · Yinger Zhang, Shutong Jiang, Renhao Li, Jianhong Tu 외 arxiv

While agent evaluation has shifted toward long-horizon tasks, most benchmarks still emphasize local, step-level reasoning rather than the global constrained optimization (e.g., time and financial budgets) that demands ge…

LongIns: A Challenging Long-context Instruction-based Exam for LLMs

2024-06-25 · Shawn Gavin, Tuney Zheng, Jiaheng Liu, Quehry Que 외

The long-context capabilities of large language models (LLMs) have been a hot topic in recent years. To evaluate the performance of LLMs in different scenarios, various assessment benchmarks have emerged. However, as mos…

16k4k

QualBench: Benchmarking Chinese LLMs with Localized Professional Qualifications for Vertical Domain Evaluation

2025-05-08 · Mengze Hong, Wailing Ng, Di Jiang, Chen Jason Zhang

The rapid advancement of Chinese large language models (LLMs) underscores the need for domain-specific evaluations to ensure reliable applications. However, existing benchmarks often lack coverage in vertical domains and…

BenchmarkingFederated LearningRAG

EnviroLLM: Resource Tracking and Optimization for Local AI

2025-12-12 · Troy Allen arxiv

Large language models (LLMs) are increasingly deployed locally for privacy and accessibility, yet users lack tools to measure their resource usage, environmental impact, and efficiency metrics. This paper presents Enviro…

LAG-MMLU: Benchmarking Frontier LLM Understanding in Latvian and Giriama

2025-03-14 · Naome A. Etori, Kevin Lu, Randu Karisa, Arturs Kanepajs

As large language models (LLMs) rapidly advance, evaluating their performance is critical. LLMs are trained on multilingual data, but their reasoning abilities are mainly evaluated using English datasets. Hence, robust e…

BenchmarkingMMLU