paper-with-me

홈 › Papers

Swiss-Bench SBP-002: A Frontier Model Comparison on Swiss Legal and Regulatory Tasks

2026-03-24 · Fatih Uenal arxiv

While recent work has benchmarked large language models on Swiss legal translation (Niklaus et al., 2025) and academic legal reasoning from university exams (Fan et al., 2025), no existing benchmark evaluates frontier model performance on applied Swiss regulatory compliance tasks. I introduce Swiss-Bench SBP-002, a trilingual benchmark of 395 expert-crafted items spanning three Swiss regulatory domains (FINMA, Legal-CH, EFK), seven task types, and three languages (German, French, Italian), and evaluate ten frontier models from March 2026 using a structured three-dimension scoring framework assessed via a blind three-judge LLM panel (GPT-4o, Claude Sonnet 4, Qwen3-235B) with majority-vote aggregation and weighted kappa = 0.605, with reference answers validated by an independent human legal expert on a 100-item subset (73% rated Correct, 0% Incorrect, perfect Legal Accuracy). Results reveal three descriptive performance clusters: Tier A (35-38% correct), Tier B (26-29%), and Tier C (13-21%). The benchmark proves difficult: even the top-ranked model (Qwen 3.5 Plus) achieves only 38.2% correct, with 47.3% incorrect and 14.4% partially correct. Task type difficulty varies widely: legal translation and case analysis yield 69-72% correct rates, while regulatory Q&A, hallucination detection, and gap analysis remain below 9%. Within this roster (seven open-weight, three closed-source), an open-weight model leads the ranking, and several open-weight models match or outperform their closed-source counterparts. These findings provide an initial empirical reference point for assessing frontier model capability on Swiss regulatory tasks under zero-retrieval conditions.

📄 PDF Abstract BibTeX arXiv:2603.23646

Code (0)

등록된 구현이 없습니다.

Tasks

Legal Reasoning

Similar Papers 제목 키워드 기반

Swiss-Bench 003: Evaluating LLM Reliability and Adversarial Security for Swiss Regulatory Contexts

2026-04-07 · Fatih Uenal arxiv

The deployment of large language models (LLMs) in Swiss financial and regulatory contexts demands empirical evidence of both production reliability and adversarial security, dimensions not jointly operationalized in exis…

Customized Neural Machine Translation Systems for the Swiss Legal Domain

2020-10-01 · AMTA 2020 10 · Rubén Martínez-Domínguez, Matīss Rikters, Artūrs Vasiļevskis, Mārcis Pinnis 외
Machine TranslationTranslation

Automated Boilerplate: Prevalence and Quality of Contract Generators in the Context of Swiss Privacy Policies

2025-10-07 · Luka Nenadic, David Rodriguez arxiv

It has become increasingly challenging for firms to comply with a plethora of novel digital regulations. This is especially true for smaller businesses that often lack both the resources and know-how to draft complex leg…

Standard German Subtitling of Swiss German TV content: the PASSAGE Project

2022-06-01 · LREC 2022 6 · Jonathan David Mutal, Pierrette Bouillon, Johanna Gerlach, Veronika Haberkorn

In Switzerland, two thirds of the population speak Swiss German, a primarily spoken language with no standardised written form. It is widely used on Swiss TV, for example in news reports, interviews or talk shows, and su…

speech-recognitionSpeech RecognitionTranslation

Text-to-Speech Pipeline for Swiss German -- A comparison

2023-05-31 · Tobias Bollinger, Jan Deriu, Manfred Vogel

In this work, we studied the synthesis of Swiss German speech using different Text-to-Speech (TTS) models. We evaluated the TTS models on three corpora, and we found, that VITS models performed best, hence, using them fo…

Speech Synthesistext-to-speechText to Speech