Why Not Simply Translate? A First Swedish Evaluation Benchmark for Semantic Similarity
This paper presents the first Swedish evaluation benchmark for textual semantic similarity. The benchmark is compiled by simply running the English STS-B dataset through the Google machine translation API. This paper discusses potential problems with using such a simple approach to compile a Swedish evaluation benchmark, including translation errors, vocabulary variation, and productive compounding. Despite some obvious problems with the resulting dataset, we use the benchmark to compare the majority of the currently existing Swedish text representations, demonstrating that native models outperform multilingual ones, and that simple bag of words performs remarkably well.
Code (1)
Tasks
Machine TranslationSemantic SimilaritySemantic Textual SimilaritySTSSTS-BTranslationSimilar Papers 제목 키워드 기반
Bipol: Multi-axes Evaluation of Bias with Explainability in Benchmark Datasets
We investigate five English NLP benchmark datasets (on the superGLUE leaderboard) and two Swedish datasets for bias, along multiple axes. The datasets are the following: Boolean Question (Boolq), CommitmentBank (CB), Win…
Bias DetectionDiagnosticNatural Language InferenceRTEPreferences for Idiomatic Language are Acquired Slowly -- and Forgotten Quickly: A Case Study on Swedish
In this study, we investigate how language models develop preferences for \textit{idiomatic} as compared to \textit{linguistically acceptable} Swedish, both during pretraining and when adapting a model from English to Sw…
Linguistic AcceptabilityA Parallel WordNet for English, Swedish and Bulgarian
We present the parallel creation of a WordNet resource for Swedish and Bulgarian which is tightly aligned with the Princeton WordNet. The alignment is not only on the synset level, but also on word level, by matching wor…
ArticlesMachine TranslationText GenerationTranslationThe FISKM\"O Project: Resources and Tools for Finnish-Swedish Machine Translation and Cross-Linguistic Research
This paper presents FISKM{\"O}, a project that focuses on the development of resources and tools for cross-linguistic research and machine translation between Finnish and Swedish. The goal of the project is the compilati…
Machine TranslationTranslationLessons Learned from GPT-SW3: Building the First Large-Scale Generative Language Model for Swedish
We present GTP-SW3, a 3.5 billion parameter autoregressive language model, trained on a newly created 100 GB Swedish corpus. This paper provides insights with regards to data collection and training, while highlights the…
Language ModelingLanguage ModellingText Generation