paper-with-me

Papers

The Impact of Text Normalization on Multiword Expressions Discovery in Persian

2021-09-01 · RANLP 2021 9 · Katarzyna Marszałek-Kowalewska

This paper evaluates normalization procedures of Persian text for a downstream NLP task - multiword expressions (MWEs) discovery. We discuss the challenges the Persian language poses for NLP and evaluate open-source tools that try to address these difficulties. The best-performing tool is later used in the main task - MWEs discovery. In order to discover MWEs, we use association measures and a subpart of the MirasText corpus. The results show that an F-score is 26% higher in the case of normalized input data.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Text Normalization

Similar Papers 제목 키워드 기반

Evaluating Diversity of Multiword Expressions in Annotated Text

2022-10-01 · COLING 2022 10 · Adam Lion-Bouton, Yagmur Ozturk, Agata Savary, Jean-Yves Antoine

Diversity can be decomposed into three distinct concepts, namely: variety, balance and disparity. This paper borrows from the extensive formalization and measures of diversity developed in ecology in order to evaluate th…

DiversityLemmatization

Discovery of Multiword Expressions with Loanwords and Their Equivalents in the Persian Language

2021-09-01 · RANLP 2021 9 · Katarzyna Marszałek-Kowalewska

This paper presents an attempt at multiword expressions (MWEs) discovery in the Persian language. It focuses on extracting MWEs containing lemmas of a particular group: loanwords in Persian and their equivalents proposed…

Evaluating the Impact of Verbal Multiword Expressions on Machine Translation

2025-08-24 · Linfeng Liu, Saptarshi Ghosh, Tianyu Jiang arxiv

Verbal multiword expressions (VMWEs) remain difficult for machine translation because their meanings are often not recoverable from their component words. In this study, we analyze the impact of three VMWE categories -- …

Machine Translation

Identification of Multiword Expressions in the brWaC

2014-05-01 · LREC 2014 5 · Rodrigo Boos, Kassius Prestes, Aline Villavicencio

Although corpus size is a well known factor that affects the performance of many NLP tasks, for many languages large freely available corpora are still scarce. In this paper we describe one effort to build a very large c…

Information RetrievalMachine TranslationSpelling Correction

How Well Do Embedding Models Capture Non-compositionality? A View from Multiword Expressions

2019-06-01 · WS 2019 6 · N, Navnita akumar, Timothy Baldwin, Bahar Salehi

In this paper, we apply various embedding methods on multiword expressions to study how well they capture the nuances of non-compositional data. Our results from a pool of word-, character-, and document-level embbedings…