paper-with-me

Papers

Evaluating Shortest Edit Script Methods for Contextual Lemmatization

2024-03-25 · Olia Toporkov, Rodrigo Agerri

Modern contextual lemmatizers often rely on automatically induced Shortest Edit Scripts (SES), namely, the number of edit operations to transform a word form into its lemma. In fact, different methods of computing SES have been proposed as an integral component in the architecture of several state-of-the-art contextual lemmatizers currently available. However, previous work has not investigated the direct impact of SES in the final lemmatization performance. In this paper we address this issue by focusing on lemmatization as a token classification task where the only input that the model receives is the word-label pairs in context, where the labels correspond to previously induced SES. Thus, by modifying in our lemmatization system only the SES labels that the model needs to learn, we may then objectively conclude which SES representation produces the best lemmatization results. We experiment with seven languages of different morphological complexity, namely, English, Spanish, Basque, Russian, Czech, Turkish and Polish, using multilingual and language-specific pre-trained masked language encoder-only models as a backbone to build our lemmatizers. Comprehensive experimental results, both in- and out-of-domain, indicate that computing the casing and edit operations separately is beneficial overall, but much more clearly for languages with high-inflected morphology. Notably, multilingual pre-trained language models consistently outperform their language-specific counterparts in every evaluation setting.

📄 PDF Abstract BibTeX arXiv:2403.16968

Code (1)

hitz-zentroa/ses-lemma 공식 구현

Tasks

LEMMALemmatizationtoken-classificationToken Classification

Similar Papers 제목 키워드 기반

Novel algorithm to generate shortest edit script using Levenshtein distance algorithm

2022-08-16 · Github 2022 8 · P. Prakash Maria Liju

String similarity, longest common subsequence and shortest edit scripts are the triplets of problem that related to each other. There are different algorithms exist to generate edit script by solving longest common subse…

Edit script generationFile difference

Robust Data-driven Prescriptiveness Optimization

2023-06-09 · Mehran Poursoltani, Erick Delage, Angelos Georghiou

The abundance of data has led to the emergence of a variety of optimization techniques that attempt to leverage available side information to provide more anticipative decisions. The wide range of methods and contexts of…

CA-Edit: Causality-Aware Condition Adapter for High-Fidelity Local Facial Attribute Editing

2024-12-18 · Xiaole Xian, Xilin He, Zenghao Niu, Junliang Zhang 외

For efficient and high-fidelity local facial attribute editing, most existing editing methods either require additional fine-tuning for different editing effects or tend to affect beyond the editing regions. Alternativel…

Attribute

SubER: A Metric for Automatic Evaluation of Subtitle Quality

2022-05-11 · Patrick Wilken, Panayota Georgakopoulou, Evgeny Matusov

This paper addresses the problem of evaluating the quality of automatically generated subtitles, which includes not only the quality of the machine-transcribed or translated speech, but also the quality of line segmentat…

SegmentationTranslation

SubER - A Metric for Automatic Evaluation of Subtitle Quality

2022-05-01 · IWSLT (ACL) 2022 5 · Patrick Wilken, Panayota Georgakopoulou, Evgeny Matusov

This paper addresses the problem of evaluating the quality of automatically generated subtitles, which includes not only the quality of the machine-transcribed or translated speech, but also the quality of line segmentat…

SegmentationTranslation