paper-with-me

Papers

Semi-Automatic Data Annotation, POS Tagging and Mildly Context-Sensitive Disambiguation: the eXtended Revised AraMorph (XRAM)

2016-03-06 · Giuliano Lancioni, Valeria Pettinari, Laura Garofalo, Marta Campanelli, Ivana Pepe, Simona Olivieri, Ilaria Cicola

An extended, revised form of Tim Buckwalter's Arabic lexical and morphological resource AraMorph, eXtended Revised AraMorph (henceforth XRAM), is presented which addresses a number of weaknesses and inconsistencies of the original model by allowing a wider coverage of real-world Classical and contemporary (both formal and informal) Arabic texts. Building upon previous research, XRAM enhancements include (i) flag-selectable usage markers, (ii) probabilistic mildly context-sensitive POS tagging, filtering, disambiguation and ranking of alternative morphological analyses, (iii) semi-automatic increment of lexical coverage through extraction of lexical and morphological information from existing lexical resources. Testing of XRAM through a front-end Python module showed a remarkable success level.

📄 PDF Abstract BibTeX arXiv:1603.01833

Code (0)

등록된 구현이 없습니다.

Tasks

POSPOS Tagging

Similar Papers 제목 키워드 기반

Improving corpus annotation productivity: a method and experiment with interactive tagging

2012-05-01 · LREC 2012 5 · Atro Voutilainen

Corpus linguistic and language technological research needs empirical corpus data with nearly correct annotation and high volume to enable advances in language modelling and theorising. Recent work on improving corpus an…

Language Modelling

A Modality Lexicon and its use in Automatic Tagging

2014-10-17 · Kathryn Baker, Michael Bloodgood, Bonnie J. Dorr, Nathaniel W. Filardo 외

This paper describes our resource-building results for an eight-week JHU Human Language Technology Center of Excellence Summer Camp for Applied Language Exploration (SCALE-2009) on Semantically-Informed Machine Translati…

Machine TranslationTranslation

Correcting Errors in a New Gold Standard for Tagging Icelandic Text

2014-05-01 · LREC 2014 5 · Sigr{\'u}n Helgad{\'o}ttir, Hrafn Loftsson, Eir{\'\i}kur R{\"o}gnvaldsson

In this paper, we describe the correction of PoS tags in a new Icelandic corpus, MIM-GOLD, consisting of about 1 million tokens sampled from the Tagged Icelandic Corpus, M{\'I}M, released in 2013. The goal is to use the …

Part-Of-Speech TaggingPOS

Visualization Framework for Colonoscopy Videos

2018-10-21 · Saad Nadeem, Arie Kaufman

We present a visualization framework for annotating and comparing colonoscopy videos, where these annotations can then be used for semi-automatic report generation at the end of the procedure. Currently, there are approx…

Parser agreement and disagreement in L2 Korean UD: Implications for human-in-the-loop annotation

2026-05-07 · Hakyung Sung, Gyu-Ho Shin arxiv

We propose a simplified human-in-the-loop workflow for second language (L2) Korean morphosyntactic annotation by leveraging agreement between two domain-adapted parsers. We first evaluate whether parser agreement can ser…