paper-with-me

Papers

MultiLexNorm: A Shared Task on Multilingual Lexical Normalization

2021-11-01 · EMNLP (WNUT) 2021 11 · Rob van der Goot, Alan Ramponi, Arkaitz Zubiaga, Barbara Plank, Benjamin Muller, Iñaki San Vicente Roncal, Nikola Ljubešić, Özlem Çetinoğlu, Rahmad Mahendra, Talha Çolakoğlu, Timothy Baldwin, Tommaso Caselli, Wladimir Sidorenko

Lexical normalization is the task of transforming an utterance into its standardized form. This task is beneficial for downstream analysis, as it provides a way to harmonize (often spontaneous) linguistic variation. Such variation is typical for social media on which information is shared in a multitude of ways, including diverse languages and code-switching. Since the seminal work of Han and Baldwin (2011) a decade ago, lexical normalization has attracted attention in English and multiple other languages. However, there exists a lack of a common benchmark for comparison of systems across languages with a homogeneous data and evaluation setup. The MultiLexNorm shared task sets out to fill this gap. We provide the largest publicly available multilingual lexical normalization benchmark including 13 language variants. We propose a homogenized evaluation setup with both intrinsic and extrinsic evaluation. As extrinsic evaluation, we use dependency parsing and part-of-speech tagging with adapted evaluation metrics (a-LAS, a-UAS, and a-POS) to account for alignment discrepancies. The shared task hosted at W-NUT 2021 attracted 9 participants and 18 submissions. The results show that neural normalization systems outperform the previous state-of-the-art system by a large margin. Downstream parsing and part-of-speech tagging performance is positively affected but to varying degrees, with improvements of up to 1.72 a-LAS, 0.85 a-UAS, and 1.54 a-POS for the winning system.

📄 PDF Abstract BibTeX

Code (1)

https://bitbucket.org/robvanderg/multilexnorm 공식 구현

Tasks

Dependency ParsingLexical NormalizationPart-Of-Speech TaggingPOS

Similar Papers 제목 키워드 기반

ÚFAL at MultiLexNorm 2021: Improving Multilingual Lexical Normalization by Fine-tuning ByT5

2021-10-28 · WNUT (ACL) 2021 11 · David Samuel, Milan Straka

We present the winning entry to the Multilingual Lexical Normalization (MultiLexNorm) shared task at W-NUT 2021 (van der Goot et al., 2021a), which evaluates lexical-normalization systems on 12 social media datasets in 1…

Dependency ParsingLanguage ModelingLanguage ModellingLexical Normalization

Sesame Street to Mount Sinai: BERT-constrained character-level Moses models for multilingual lexical normalization

2021-11-01 · WNUT (ACL) 2021 11 · Yves Scherrer, Nikola Ljubešić

This paper describes the HEL-LJU submissions to the MultiLexNorm shared task on multilingual lexical normalization. Our system is based on a BERT token classification preprocessing step, where for each token the type of …

Lexical Normalizationtoken-classificationToken Classification

MultiLexNorm++: A Unified Benchmark and a Generative Model for Lexical Normalization for Asian Languages

2026-01-23 · Weerayut Buaphet, Thanh-Nhi Nguyen, Risa Kondo, Tomoyuki Kajiwara 외 arxiv

Social media data has been of interest to Natural Language Processing (NLP) practitioners for over a decade, because of its richness in information, but also challenges for automatic processing. Since language use is mor…

Lexical Normalization

Multilingual Sequence Labeling Approach to solve Lexical Normalization

2021-11-01 · WNUT (ACL) 2021 11 · Divesh Kubal, Apurva Nagvenkar

The task of converting a nonstandard text to a standard and readable text is known as lexical normalization. Almost all the Natural Language Processing (NLP) applications require the text data in normalized form to build…

Language ModellingLexical NormalizationWord Alignment

Findings of the TSAR-2022 Shared Task on Multilingual Lexical Simplification

2023-02-06 · Horacio Saggion, Sanja Štajner, Daniel Ferrés, Kim Cheng SHEANG 외

We report findings of the TSAR-2022 shared task on multilingual lexical simplification, organized as part of the Workshop on Text Simplification, Accessibility, and Readability TSAR-2022 held in conjunction with EMNLP 20…

Lexical SimplificationText Simplification