paper-with-me

홈 › Papers

High Quality Word Lists as a Resource for Multiple Purposes

2014-05-01 · LREC 2014 5 · Uwe Quasthoff, Dirk Goldhahn, Thomas Eckart, Erla Hallsteinsd{\'o}ttir, Sabine Fiedler

Since 2011 the comprehensive, electronically available sources of the Leipzig Corpora Collection have been used consistently for the compilation of high quality word lists. The underlying corpora include newspaper texts, Wikipedia articles and other randomly collected Web texts. For many of the languages featured in this collection, it is the first comprehensive compilation to use a large-scale empirical base. The word lists have been used to compile dictionaries with comparable frequency data in the Frequency Dictionaries series. This includes frequency data of up to 1,000,000 word forms presented in alphabetical order. This article provides an introductory description of the data and the methodological approach used. In addition, language-specific statistical information is provided with regard to letters, word structure and structural changes. Such high quality word lists also provide the opportunity to explore comparative linguistic topics and such monolingual issues as studies of word formation and frequency-based examinations of lexical areas for use in dictionaries or language teaching. The results presented here can provide initial suggestions for subsequent work in several areas of research.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesVocal Bursts Intensity Prediction

Similar Papers 제목 키워드 기반

Novel Language Resources for Hindi: An Aesthetics Text Corpus and a Comprehensive Stop Lemma List

2020-02-01 · Gayatri Venugopal-Wairagade, Jatinderkumar R. Saini, Dhanya Pramod

This paper is an effort to complement the contributions made by researchers working toward the inclusion of non-English languages in natural language processing studies. Two novel Hindi language resources have been creat…

LEMMA

Attention-based Vocabulary Selection for NMT Decoding

2017-06-12 · Baskaran Sankaran, Markus Freitag, Yaser Al-Onaizan

Neural Machine Translation (NMT) models usually use large target vocabulary sizes to capture most of the words in the target language. The vocabulary size is a big factor when decoding new sentences as the final softmax …

Machine TranslationNMTSentenceTranslation

Keyword Augmentation via Generative Methods

2021-08-01 · ACL (ECNLP) 2021 8 · Haoran Shi, Zhibiao Rao, Yongning Wu, Zuohua Zhang 외

Keyword augmentation is a fundamental problem for sponsored search modeling and business. Machine generated keywords can be recommended to advertisers for better campaign discoverability as well as used as features for s…

valid

Building a Linguistic Resource : A Word Frequency List for Sinhala

2021-12-01 · ICON 2021 12 · Aloka Fernando, Gihan Dias

A word frequency list is a list of unique words in a language along with their frequency count. It is generally sorted by frequency. Such a list is essential for many NLP tasks, including building language models, POS ta…

POS

Information-Theoretic Characterization of Vowel Harmony: A Cross-Linguistic Study on Word Lists

2023-08-09 · Julius Steuer, Badr Abdullah, Johann-Mattis List, Dietrich Klakow

We present a cross-linguistic study that aims to quantify vowel harmony using data-driven computational modeling. Concretely, we define an information-theoretic measure of harmonicity based on the predictability of vowel…

LEMMA