Constructing a Norwegian Academic Wordlist
We present the development of a Norwegian Academic Wordlist (AKA list) for the Norwegian Bokm{\"a}l variety. To identify specific academic vocabulary we developed a 100-million-word academic corpus based on the University of Oslo archive of digital publications. Other corpora were used for testing and developing general word lists. We tried two different methods, those of Carlund et al. (2012) and Gardner {\&} Davies (2013), and compared them. The resulting list is presented on a web site, where the words can be inspected in different ways, and freely downloaded.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Unstable Grounds for Beautiful Trees? Testing the Robustness of Concept Translations in the Compilation of Multilingual Wordlists
Multilingual wordlists play a crucial role in comparative linguistics. While many studies have been carried out to test the power of computational methods for language subgrouping or divergence time estimation, few studi…
The Norwegian Colossal Corpus: A Text Corpus for Training Large Norwegian Language Models
Norwegian has been one of many languages lacking sufficient available text to train quality language models. In an attempt to bridge this gap, we introduce the Norwegian Colossal Corpus (NCC), which comprises 49GB of cle…
A Health Focused Text Classification Tool (HFTCT)
Due to the high number of users on social media and the massive amounts of queries requested every second to share a new video, picture, or message, social platforms struggle to manage this humungous amount of data that …
Classificationtext-classificationText ClassificationInline Detection of Domain Generation Algorithms with Context-Sensitive Word Embeddings
Domain generation algorithms (DGAs) are frequently employed by malware to generate domains used for connecting to command-and-control (C2) servers. Recent work in DGA detection leveraged deep learning architectures like …
Word Embeddings