paper-with-me

홈 › Papers

20min-XD: A Comparable Corpus of Swiss News Articles

2025-04-30 · Michelle Wastl, Jannis Vamvas, Selena Calleri, Rico Sennrich

We present 20min-XD (20 Minuten cross-lingual document-level), a French-German, document-level comparable corpus of news articles, sourced from the Swiss online news outlet 20 Minuten/20 minutes. Our dataset comprises around 15,000 article pairs spanning 2015 to 2024, automatically aligned based on semantic similarity. We detail the data collection process and alignment methodology. Furthermore, we provide a qualitative and quantitative analysis of the corpus. The resulting dataset exhibits a broad spectrum of cross-lingual similarity, ranging from near-translations to loosely related articles, making it valuable for various NLP applications and broad linguistically motivated studies. We publicly release the dataset in document- and sentence-aligned versions and code for the described experiments.

📄 PDF Abstract BibTeX arXiv:2504.21677

Code (1)

ZurichNLP/20min-XD 공식 구현

Tasks

ArticlesSemantic SimilaritySemantic Textual SimilaritySentence

Similar Papers 제목 키워드 기반

Fine-tuning the SwissBERT Encoder Model for Embedding Sentences and Documents

2024-05-13 · Juri Grosjean, Jannis Vamvas

Encoder models trained for the embedding of sentences or short documents have proven useful for tasks such as semantic search and topic modeling. In this paper, we present a version of the SwissBERT encoder model that we…

ArticlesContrastive LearningRetrievaltext-classification+1

SwissBERT: The Multilingual Language Model for Switzerland

2023-03-23 · Jannis Vamvas, Johannes Graën, Rico Sennrich

We present SwissBERT, a masked language model created specifically for processing Switzerland-related text. SwissBERT is a pre-trained model that we adapted to news articles written in the national languages of Switzerla…

ArticlesLanguage ModelingLanguage Modellingmodel+1

Lest We Forget: A Dataset of Coronavirus-Related News Headlines in Swiss Media

2020-06-25 · Alireza Ghasemi, Amina Chebira

We release our COVID-19 news dataset, containing more than 10,000 links to news articles related to the Coronavirus pandemic published in the Swiss media since early January 2020. This collection can prove beneficial in …

Articles

Upgrading the Newsroom: An Automated Image Selection System for News Articles

2020-04-23 · Fangyu Liu, Rémi Lebret, Didier Orel, Philippe Sordet 외

We propose an automated image selection system to assist photo editors in selecting suitable images for news articles. The system fuses multiple textual sources extracted from news articles and accepts multilingual input…

ArticlesImage RetrievalRetrievalWeakly-supervised Learning+1

Swiss-AL: A Multilingual Swiss Web Corpus for Applied Linguistics

2020-05-01 · LREC 2020 5 · Julia Krasselt, Philipp Dressen, Matthias Fluor, Cerstin Mahlow 외

The Swiss Web Corpus for Applied Linguistics (Swiss-AL) is a multilingual (German, French, Italian) collection of texts from selected web sources. Unlike most other web corpora it is not intended for NLP purposes, but ra…