paper-with-me

Papers

SwissBERT: The Multilingual Language Model for Switzerland

2023-03-23 · Jannis Vamvas, Johannes Graën, Rico Sennrich

We present SwissBERT, a masked language model created specifically for processing Switzerland-related text. SwissBERT is a pre-trained model that we adapted to news articles written in the national languages of Switzerland -- German, French, Italian, and Romansh. We evaluate SwissBERT on natural language understanding tasks related to Switzerland and find that it tends to outperform previous models on these tasks, especially when processing contemporary news and/or Romansh Grischun. Since SwissBERT uses language adapters, it may be extended to Swiss German dialects in future work. The model and our open-source code are publicly released at https://github.com/ZurichNLP/swissbert.

📄 PDF Abstract BibTeX arXiv:2303.13310

Code (1)

zurichnlp/swissbert 공식 구현

Tasks

ArticlesLanguage ModelingLanguage ModellingmodelNatural Language Understanding

Similar Papers 제목 키워드 기반

Fine-tuning the SwissBERT Encoder Model for Embedding Sentences and Documents

2024-05-13 · Juri Grosjean, Jannis Vamvas

Encoder models trained for the embedding of sentences or short documents have proven useful for tasks such as semantic search and topic modeling. In this paper, we present a version of the SwissBERT encoder model that we…

ArticlesContrastive LearningRetrievaltext-classification+1

SwissADT: An Audio Description Translation System for Swiss Languages

2024-11-22 · Lukas Fischer, Yingqiang Gao, Alexa Lintner, Sarah Ebling

Audio description (AD) is a crucial accessibility service provided to blind persons and persons with visual impairment, designed to convey visual information in acoustic form. Despite recent advancements in multilingual …

Machine TranslationTranslation

Unlocking Legal Knowledge: A Multilingual Dataset for Judicial Summarization in Switzerland

2024-10-17 · Luca Rolshoven, Vishvaksenan Rasiah, Srinanda Brügger Bose, Matthias Stürmer 외

Legal research is a time-consuming task that most lawyers face on a daily basis. A large part of legal research entails looking up relevant caselaw and bringing it in relation to the case at hand. Lawyers heavily rely on…

Human-in-the-Loop Hate Speech Classification in a Multilingual Context

2022-12-05 · Ana Kotarcic, Dominik Hangartner, Fabrizio Gilardi, Selina Kurer 외

The shift of public debate to the digital sphere has been accompanied by a rise in online hate speech. While many promising approaches for hate speech classification have been proposed, studies often focus only on a sing…

Classification

FairLex: A Multilingual Benchmark for Evaluating Fairness in Legal Text Processing

2022-03-14 · ACL 2022 5 · Ilias Chalkidis, Tommaso Pasini, Sheng Zhang, Letizia Tomada 외

We present a benchmark suite of four datasets for evaluating the fairness of pre-trained language models and the techniques used to fine-tune them for downstream tasks. Our benchmarks cover four jurisdictions (European C…

Fairness