paper-with-me

홈 › Papers

Generating abbreviations using Google Books library

2014-10-04 · Valery D. Solovyev, Vladimir V. Bochkarev

The article describes the original method of creating a dictionary of abbreviations based on the Google Books Ngram Corpus. The dictionary of abbreviations is designed for Russian, yet as its methodology is universal it can be applied to any language. The dictionary can be used to define the function of the period during text segmentation in various applied systems of text processing. The article describes difficulties encountered in the process of its construction as well as the ways to overcome them. A model of evaluating a probability of first and second type errors (extraction accuracy and fullness) is constructed. Certain statistical data for the use of abbreviations are provided.

📄 PDF Abstract BibTeX arXiv:1410.1080

Code (0)

등록된 구현이 없습니다.

Tasks

Text Segmentation

Similar Papers 제목 키워드 기반

Characterizing the Google Books corpus: Strong limits to inferences of socio-cultural and linguistic evolution

2015-01-05 · Eitan Adam Pechenick, Christopher M. Danforth, Peter Sheridan Dodds

It is tempting to treat frequency trends from the Google Books data sets as indicators of the "true" popularity of various words and phrases. Doing so allows us to draw quantitatively strong conclusions about the evoluti…

Articles

MaintNet: A Collaborative Open-Source Library for Predictive Maintenance Language Resources

2020-05-25 · COLING 2020 8 · Farhad Akhbardeh, Travis Desell, Marcos Zampieri

Maintenance record logbooks are an emerging text type in NLP. They typically consist of free text documents with many domain specific technical terms, abbreviations, as well as non-standard spelling and grammar, which po…

Clustering

Dealing with Abbreviations in the Slovenian Biographical Lexicon

2022-11-04 · Angel Daza, Antske Fokkens, Tomaž Erjavec

Abbreviations present a significant challenge for NLP systems because they cause tokenization and out-of-vocabulary errors. They can also make the text less readable, especially in reference printed books, where they are…

Institutional Books 1.0: A 242B token dataset from Harvard Library's collections, refined for accuracy and usability

2025-06-10 · Matteo Cargnelutti, Catherine Brobston, John Hess, Jack Cushman 외

Large language models (LLMs) use data to learn about the world in order to produce meaningful correlations and predictions. As such, the nature, scale, quality, and diversity of the datasets used to train these models, o…

Optical Character Recognition (OCR)

NLP Tools for Predictive Maintenance Records in MaintNet

2020-12-01 · Asian Chapter of the Association for Computational Linguistics 2020 · Farhad Akhbardeh, Travis Desell, Marcos Zampieri

Processing maintenance logbook records is an important step in the development of predictive maintenance systems. Logbooks often include free text fields with domain specific terms, abbreviations, and non-standard spelli…

ClusteringPOSPOS Tagging