paper-with-me

Papers

Measuring Harmful Representations in Scandinavian Language Models

2022-11-21 · Samia Touileb, Debora Nozza

Scandinavian countries are perceived as role-models when it comes to gender equality. With the advent of pre-trained language models and their widespread usage, we investigate to what extent gender-based harmful and toxic content exist in selected Scandinavian language models. We examine nine models, covering Danish, Swedish, and Norwegian, by manually creating template-based sentences and probing the models for completion. We evaluate the completions using two methods for measuring harmful and toxic completions and provide a thorough analysis of the results. We show that Scandinavian pre-trained language models contain harmful and gender-based stereotypes with similar values across all languages. This finding goes against the general expectations related to gender equality in Scandinavian countries and shows the possible problematic outcomes of using such models in real-world settings.

📄 PDF Abstract BibTeX arXiv:2211.11678

Code (1)

samiatouileb/scandinavianhonest 공식 구현

Similar Papers 제목 키워드 기반

ScandEval: A Benchmark for Scandinavian Natural Language Processing

2023-04-03 · Dan Saattrup Nielsen

This paper introduces a Scandinavian benchmarking platform, ScandEval, which can benchmark any pretrained model on four different tasks in the Scandinavian languages. The datasets used in two of the tasks, linguistic acc…

BenchmarkingCross-Lingual TransferLinguistic AcceptabilityQuestion Answering

SWEb: A Large Web Dataset for the Scandinavian Languages

2024-10-06 · Tobias Norlund, Tim Isbister, Amaru Cuba Gyllensten, Paul Dos Santos 외

This paper presents the hitherto largest pretraining dataset for the Scandinavian languages: the Scandinavian WEb (SWEb), comprising over one trillion tokens. The paper details the collection and processing pipeline, and…

The Scandinavian Embedding Benchmarks: Comprehensive Assessment of Multilingual and Monolingual Text Embedding

2024-06-04 · Kenneth Enevoldsen, Márton Kardos, Niklas Muennighoff, Kristoffer Laigaard Nielbo

The evaluation of English text embeddings has transitioned from evaluating a handful of datasets to broad coverage across many tasks through benchmarks such as MTEB. However, this is not the case for multilingual text em…

Multi-label Scandinavian Language Identification (SLIDE)

2025-02-10 · Mariia Fedorova, Jonas Sebulon Frydenberg, Victoria Handford, Victoria Ovedie Chruickshank Langø 외

Identifying closely related languages at sentence level is difficult, in particular because it is often impossible to assign a sentence to a single language. In this paper, we focus on multi-label sentence-level Scandina…

Language IdentificationSentence

Building Sentiment Lexicons for Mainland Scandinavian Languages Using Machine Translation and Sentence Embeddings

2022-06-01 · LREC 2022 6 · Peng Liu, Cristina Marco, Jon Atle Gulla

This paper presents a simple but effective method to build sentiment lexicons for the three Mainland Scandinavian languages: Danish, Norwegian and Swedish. This method benefits from the English Sentiwordnet and a thesaur…

Machine TranslationSentenceSentence EmbeddingSentence-Embedding+3