paper-with-me

홈 › Papers

Neologisms on Facebook

2018-04-13 · Nikita Muravyev, Alexander Panchenko, Sergei Obiedkov

In this paper, we present a study of neologisms and loan words frequently occurring in Facebook user posts. We have analyzed a dataset of several million publically available posts written during 2006-2013 by Russian-speaking Facebook users. From these, we have built a vocabulary of most frequent lemmatized words missing from the OpenCorpora dictionary the assumption being that many such words have entered common use only recently. This assumption is certainly not true for all the words extracted in this way; for that reason, we manually filtered the automatically obtained list in order to exclude non-Russian or incorrectly lemmatized words, as well as words recorded by other dictionaries or those occurring in texts from the Russian National Corpus. The result is a list of 168 words that can potentially be considered neologisms. We present an attempt at an etymological classification of these neologisms (unsurprisingly, most of them have recently been borrowed from English, but there are also quite a few new words composed of previously borrowed stems) and identify various derivational patterns. We also classify words into several large thematic areas, "internet", "marketing", and "multimedia" being among those with the largest number of words. We believe that, together with the word base collected in the process, they can serve as a starting point in further studies of neologisms and lexical processes that lead to their acceptance into the mainstream language.

📄 PDF Abstract BibTeX arXiv:1804.05831

Code (0)

등록된 구현이 없습니다.

Tasks

Marketing

Similar Papers 제목 키워드 기반

NEO-BENCH: Evaluating Robustness of Large Language Models with Neologisms

2024-02-19 · Jonathan Zheng, Alan Ritter, Wei Xu

The performance of Large Language Models (LLMs) degrades from the temporal drift between data used for model training and newer text seen during inference. One understudied avenue of language change causing data drift is…

Machine TranslationNatural Language UnderstandingSentence

A primer on getting neologisms from foreign languages to under-resourced languages

2023-03-07 · Luis Camacho

Mainly due to lack of support, most under-resourced languages have a reduced lexicon in most realms and domains of increasing importance, then their speakers need to significantly augment it. Although neologisms should a…

Common Sense Reasoning

Unsupervised Neologism Normalization Using Embedding Space Mapping

2019-11-01 · WS 2019 11 · Nasser Zalmout, Kapil Thadani, Aasish Pappu

This paper presents an approach for detecting and normalizing neologisms in social media content. Neologisms refer to recent expressions that are specific to certain entities or events and are being increasingly used by …

Natural Language UnderstandingText Normalization

Skill Neologisms: Towards Skill-based Continual Learning

2026-05-06 · Antonin Berthon, Nicolas Astorga, Mihaela van der Schaar arxiv

Modern LLMs show mastery over an ever-growing range of skills, as well as the ability to compose them flexibly. However, extending model capabilities to new skills in a scalable manner is an open problem: fine-tuning and…

Continual Learning

CNeo-Bench: Diagnosing Large Language Models on Chinese Neologisms

2026-08-28 · Kaiyan Zhao, Zhongtao Miao, Zheyong Xie, Shaosheng Cao 외 arxiv

Chinese neologisms exploit diverse and unique linguistic mechanisms, such as phonetic substitution (e.g., 886 for ``bye-bye'') and visual character decomposition that are rare in other languages. We introduce CNeo-Bench,…