20min-XD: A Comparable Corpus of Swiss News Articles
We present 20min-XD (20 Minuten cross-lingual document-level), a French-German, document-level comparable corpus of news articles, sourced from the Swiss online news outlet 20 Minuten/20 minutes. Our dataset comprises around 15,000 article pairs spanning 2015 to 2024, automatically aligned based on semantic similarity. We detail the data collection process and alignment methodology. Furthermore, we provide a qualitative and quantitative analysis of the corpus. The resulting dataset exhibits a broad spectrum of cross-lingual similarity, ranging from near-translations to loosely related articles, making it valuable for various NLP applications and broad linguistically motivated studies. We publicly release the dataset in document- and sentence-aligned versions and code for the described experiments.
Code (1)
Tasks
ArticlesSemantic SimilaritySemantic Textual SimilaritySentenceSimilar Papers 제목 키워드 기반
Fine-tuning the SwissBERT Encoder Model for Embedding Sentences and Documents
Encoder models trained for the embedding of sentences or short documents have proven useful for tasks such as semantic search and topic modeling. In this paper, we present a version of the SwissBERT encoder model that we…
ArticlesContrastive LearningRetrievaltext-classification+1SwissBERT: The Multilingual Language Model for Switzerland
We present SwissBERT, a masked language model created specifically for processing Switzerland-related text. SwissBERT is a pre-trained model that we adapted to news articles written in the national languages of Switzerla…
ArticlesLanguage ModelingLanguage Modellingmodel+1Lest We Forget: A Dataset of Coronavirus-Related News Headlines in Swiss Media
We release our COVID-19 news dataset, containing more than 10,000 links to news articles related to the Coronavirus pandemic published in the Swiss media since early January 2020. This collection can prove beneficial in …
ArticlesUpgrading the Newsroom: An Automated Image Selection System for News Articles
We propose an automated image selection system to assist photo editors in selecting suitable images for news articles. The system fuses multiple textual sources extracted from news articles and accepts multilingual input…
ArticlesImage RetrievalRetrievalWeakly-supervised Learning+1Swiss-AL: A Multilingual Swiss Web Corpus for Applied Linguistics
The Swiss Web Corpus for Applied Linguistics (Swiss-AL) is a multilingual (German, French, Italian) collection of texts from selected web sources. Unlike most other web corpora it is not intended for NLP purposes, but ra…