paper-with-me

홈 › Papers

LexiClean: An annotation tool for rapid multi-task lexical normalisation

2021-11-01 · EMNLP (ACL) 2021 11 · Tyler Bikaun, Tim French, Melinda Hodkiewicz, Michael Stewart, Wei Liu

NLP systems are often challenged by difficulties arising from noisy, non-standard, and domain specific corpora. The task of lexical normalisation aims to standardise such corpora, but currently lacks suitable tools to acquire high-quality annotated data to support deep learning based approaches. In this paper, we present LexiClean, the first open-source web-based annotation tool for multi-task lexical normalisation. LexiClean’s main contribution is support for simultaneous in situ token-level modification and annotation that can be rapidly applied corpus wide. We demonstrate the usefulness of our tool through a case study on two sets of noisy corpora derived from the specialised-domain of industrial mining. We show that LexiClean allows for the rapid and efficient development of high-quality parallel corpora. A demo of our system is available at: https://youtu.be/P7_ooKrQPDU.

📄 PDF Abstract BibTeX

Code (1)

nlp-tlp/lexiclean 공식 구현

Similar Papers 제목 키워드 기반

QuickGraph: A Rapid Annotation Tool for Knowledge Graph Extraction from Technical Text

2022-05-01 · ACL 2022 5 · Tyler Bikaun, Michael Stewart, Wei Liu

Acquiring high-quality annotated corpora for complex multi-task information extraction (MT-IE) is an arduous and costly process for human-annotators. Adoption of unsupervised techniques for automated annotation have thus…

Clustering

Controlled Propagation of Concept Annotations in Textual Corpora

2016-05-01 · LREC 2016 5 · Cyril Grouin

In this paper, we presented the annotation propagation tool we designed to be used in conjunction with the BRAT rapid annotation tool. We designed two experiments to annotate a corpus of 60 files, first not using our too…

TeamTat: a collaborative text annotation tool

2020-04-24 · Rezarta Islamaj, Dongseop Kwon, Sun Kim, Zhiyong Lu

Manually annotated data is key to developing text-mining and information-extraction algorithms. However, human annotation requires considerable time, effort and expertise. Given the rapid growth of biomedical literature,…

Managementtext annotation

From Bounding Boxes to Visual Reasoning: An On-Policy Data Annotation Tool for Vision-Language Models

2026-06-17 · Like Zhang, Runliang Niu, Shiqi Wang, Xiyu Hu 외 arxiv

Vision-language models (VLMs) are rapidly advancing toward sophisticated grounded structured visual reasoning. Training models for such advanced capabilities demands a new genre of data that seamlessly unifies spatial co…

Visual Reasoning

AIANO: Enhancing Information Retrieval with AI-Augmented Annotation

2026-02-04 · Sameh Khattab, Marie Bauer, Lukas Heine, Till Rostalski 외 arxiv

The rise of Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) has rapidly increased the need for high-quality, curated information retrieval datasets. These datasets, however, are currently created wi…

Information Retrieval