Operationalizing a National Digital Library: The Case for a Norwegian Transformer Model
In this work, we show the process of building a large-scale training set from digital and digitized collections at a national library. The resulting Bidirectional Encoder Representations from Transformers (BERT)-based language model for Norwegian outperforms multilingual BERT (mBERT) models in several token and sequence classification tasks for both Norwegian Bokm{\aa}l and Norwegian Nynorsk. Our model also improves the mBERT performance for other languages present in the corpus such as English, Swedish, and Danish. For languages not included in the corpus, the weights degrade moderately while keeping strong multilingual properties. Therefore, we show that building high-quality models within a memory institution using somewhat noisy optical character recognition (OCR) content is feasible, and we hope to pave the way for other memory institutions to follow.
Code (2)
Tasks
Language ModelingLanguage ModellingOptical Character RecognitionOptical Character Recognition (OCR)Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
The Norwegian Dependency Treebank
The Norwegian Dependency Treebank is a new syntactic treebank for Norwegian Bokm{\"a}l and Nynorsk with manual syntactic and morphological annotation, developed at the National Library of Norway in collaboration with the…
Dependency ParsingMachine TranslationPart-Of-Speech TaggingSemantic Parsing+1An open source part-of-speech tagger for Norwegian: Building on existing language resources
This paper presents an open source part-of-speech tagger for the Norwegian language. It describes how an existing language processing library (FreeLing) was used to build a new part-of-speech tagger for this language. Th…
Dependency ParsingMachine TranslationMorphological AnalysisMorphological Tagging+1Visual Navigation of Digital Libraries: Retrieval and Classification of Images in the National Library of Norway's Digitised Book Collection
Digital tools for text analysis have long been essential for the searchability and accessibility of digitised library collections. Recent computer vision advances have introduced similar capabilities for visual materials…
Classificationimage-classificationImage ClassificationImage Retrieval+2Perspective Chapter: MOOCs in India: Evolution, Innovation, Impact, and Roadmap
With the largest population of the world and one of the highest enrolments in higher education, India needs efficient and effective means to educate its learners. India started focusing on open and digital education in 1…
Constructing a Norwegian Academic Wordlist
We present the development of a Norwegian Academic Wordlist (AKA list) for the Norwegian Bokm{\"a}l variety. To identify specific academic vocabulary we developed a 100-million-word academic corpus based on the Universit…