paper-with-me

홈 › Papers

The Nordic Pile: A 1.2TB Nordic Dataset for Language Modeling

2023-03-30 · Joey Öhman, Severine Verlinden, Ariel Ekgren, Amaru Cuba Gyllensten, Tim Isbister, Evangelia Gogoulou, Fredrik Carlsson, Magnus Sahlgren

Pre-training Large Language Models (LLMs) require massive amounts of text data, and the performance of the LLMs typically correlates with the scale and quality of the datasets. This means that it may be challenging to build LLMs for smaller languages such as Nordic ones, where the availability of text corpora is limited. In order to facilitate the development of the LLMS in the Nordic languages, we curate a high-quality dataset consisting of 1.2TB of text, in all of the major North Germanic languages (Danish, Icelandic, Norwegian, and Swedish), as well as some high-quality English data. This paper details our considerations and processes for collecting, cleaning, and filtering the dataset.

📄 PDF Abstract BibTeX arXiv:2303.17183

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

Towards Multilingual LLM Evaluation for Baltic and Nordic languages: A study on Lithuanian History

2025-01-15 · Yevhen Kostiuk, Oxana Vitman, Łukasz Gagała, Artur Kiulian

In this work, we evaluated Lithuanian and general history knowledge of multilingual Large Language Models (LLMs) on a multiple-choice question-answering task. The models were tested on a dataset of Lithuanian national an…

Multiple-choiceQuestion Answering

Discriminating Between Similar Nordic Languages

2020-12-11 · EACL (VarDial) 2021 4 · René Haas, Leon Derczynski

Automatic language identification is a challenging problem. Discriminating between closely related languages is especially difficult. This paper presents a machine learning approach for automatic language identification …

BIG-bench Machine LearningLanguage Identification

Creation of an Open Shared Language Resource Repository in the Nordic and Baltic Countries

2012-05-01 · LREC 2012 5 · Andrejs Vasi{\c{l}}jevs, Markus Forsberg, Tatiana Gornostay, Dorte Haltrup Hansen 외

The META-NORD project has contributed to an open infrastructure for language resources (data and tools) under the META-NET umbrella. This paper presents the key objectives of META-NORD and reports on the results achieved…

Probabilistic Behavioral Aggregation: A Case Study on the Nordic Power Grid

2024-12-16 · Anna Büttner, Frank Hellmann

This study applies the Probabilistic Behavioral Tuning (ProBeTune) framework to transient power grid simulations to address challenges posed by increasing grid complexity. ProBeTune offers a probabilistic approach to mod…

Material Philology Meets Digital Onomastic Lexicography: The NordiCon Database of Medieval Nordic Personal Names in Continental Sources

2020-05-01 · LREC 2020 5 · Michelle Waldisp{\"u}hl, Dana Dannells, Lars Borin

We present NordiCon, a database containing medieval Nordic personal names attested in Continental sources. The database combines formally interpreted and richly interlinked onomastic data with digitized versions of the m…

Lemmatization