paper-with-me

홈 › Papers

Unsupervised Neologism Normalization Using Embedding Space Mapping

2019-11-01 · WS 2019 11 · Nasser Zalmout, Kapil Thadani, Aasish Pappu

This paper presents an approach for detecting and normalizing neologisms in social media content. Neologisms refer to recent expressions that are specific to certain entities or events and are being increasingly used by the public, but have not yet been accepted in mainstream language. Automated methods for handling neologisms are important for natural language understanding and normalization, especially for informal genres with user generated content. We present an unsupervised approach for detecting neologisms and then normalizing them to canonical words without relying on parallel training data. Our approach builds on the text normalization literature and introduces adaptations to fit the specificities of this task, including phonetic and etymological considerations. We evaluate the proposed techniques on a dataset of Reddit comments, with detected neologisms and corresponding normalizations.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Natural Language UnderstandingText Normalization

Similar Papers 제목 키워드 기반

NeoN: A Tool for Automated Detection, Linguistic and LLM-Driven Analysis of Neologisms in Polish

2025-05-21 · Aleksandra Tomaszewska, Dariusz Czerski, Bartosz Żuk, Maciej Ogrodniczuk

NeoN, a tool for detecting and analyzing Polish neologisms. Unlike traditional dictionary-based methods requiring extensive manual review, NeoN combines reference corpora, Polish-specific linguistic filters, an LLM-drive…

Lemmatization

Cross-Lingual BERT Contextual Embedding Space Mapping with Isotropic and Isometric Conditions

2021-07-19 · Haoran Xu, Philipp Koehn

Typically, a linearly orthogonal transformation mapping is learned by aligning static type-level embeddings to build a shared semantic space. In view of the analysis that contextual embeddings contain richer semantic fea…

Unsupervised Joint Training of Bilingual Word Embeddings

2019-07-01 · ACL 2019 7 · Benjamin Marie, Atsushi Fujita

State-of-the-art methods for unsupervised bilingual word embeddings (BWE) train a mapping function that maps pre-trained monolingual word embeddings into a bilingual space. Despite its remarkable results, unsupervised ma…

Machine TranslationTranslationUnsupervised Machine TranslationWord Embeddings

FeelsGoodMan: Inferring Semantics of Twitch Neologisms

2021-08-18 · Pavel Dolin, Luc d'Hauthuille, Andrea Vattani

Twitch chats pose a unique problem in natural language understanding due to a large presence of neologisms, specifically emotes. There are a total of 8.06 million emotes, over 400k of which were used in the week studied.…

Natural Language UnderstandingSentiment AnalysisWord Embeddings

FeelsGoodMan: Inferring Semantics of Twitch Neologisms

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Twitch chat messages pose a unique problem in natural language understanding due to a large presence of neologisms, specifically emotes. There are a total of 8.06 million emotes, over 400k of which were observed during t…

Natural Language UnderstandingSentiment AnalysisWord Embeddings