paper-with-me

Papers

Swiss-AL: A Multilingual Swiss Web Corpus for Applied Linguistics

2020-05-01 · LREC 2020 5 · Julia Krasselt, Philipp Dressen, Matthias Fluor, Cerstin Mahlow, Klaus Rothenh{\"a}usler, Maren Runte

The Swiss Web Corpus for Applied Linguistics (Swiss-AL) is a multilingual (German, French, Italian) collection of texts from selected web sources. Unlike most other web corpora it is not intended for NLP purposes, but rather designed to support data-based and data-driven research on societal and political discourses in Switzerland. It currently contains 8 million texts (approx. 1.55 billion tokens), including news and specialist publications, governmental opinions, and parliamentary records, web sites of political parties, companies, and universities, statements from industry associations and NGOs, etc. A flexible processing pipeline using state-of-the-art components allows researchers in applied linguistics to create tailor-made subcorpora for studying discourse in a wide range of domains. So far, Swiss-AL has been used successfully in research on Swiss public discourses on energy and on antibiotic resistance.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SwissAdmin: A multilingual tagged parallel corpus of press releases

2014-05-01 · LREC 2014 5 · Yves Scherrer, Luka Nerima, Lorenza Russo, Maria Ivanova 외

SwissAdmin is a new multilingual corpus of press releases from the Swiss Federal Administration, available in German, French, Italian and English. We provide SwissAdmin in three versions: (i) plain texts of approximately…

Language IdentificationSentence

Modular Adaptation of Multilingual Encoders to Written Swiss German Dialect

2024-01-25 · Jannis Vamvas, Noëmi Aepli, Rico Sennrich

Creating neural text encoders for written Swiss German is challenging due to a dearth of training data combined with dialectal variation. In this paper, we build on several existing multilingual encoders and adapt them t…

Automatic Creation of Text Corpora for Low-Resource Languages from the Internet: The Case of Swiss German

2019-11-30 · LREC 2020 5 · Lucy Linder, Michael Jungo, Jean Hennebert, Claudiu Musat 외

This paper presents SwissCrawl, the largest Swiss German text corpus to date. Composed of more than half a million sentences, it was generated using a customized web scraping tool that could be applied to other low-resou…

Language ModelingLanguage Modelling

SwissGPC v1.0 -- The Swiss German Podcasts Corpus

2025-09-24 · Samuel Stucki, Mark Cieliebak, Jan Deriu arxiv

We present SwissGPC v1.0, the first mid-to-large-scale corpus of spontaneous Swiss German speech, developed to support research in ASR, TTS, dialect identification, and related fields. The dataset consists of links to ta…

SwissBERT: The Multilingual Language Model for Switzerland

2023-03-23 · Jannis Vamvas, Johannes Graën, Rico Sennrich

We present SwissBERT, a masked language model created specifically for processing Switzerland-related text. SwissBERT is a pre-trained model that we adapted to news articles written in the national languages of Switzerla…

ArticlesLanguage ModelingLanguage Modellingmodel+1