paper-with-me

홈 › Papers

Tagger for Polish Computer Mediated Communication Texts

2019-09-01 · RANLP 2019 9 · Wiktor Walentynowicz, Maciej Piasecki, Marcin Oleksy

In this paper we present a morpho-syntactic tagger dedicated to Computer-mediated Communication texts in Polish. Its construction is based on an expanded RNN-based neural network adapted to the work on noisy texts. Among several techniques, the tagger utilises fastText embedding vectors, sequential character embedding vectors, and Brown clustering for the coarse-grained representation of sentence structures. In addition a set of manually written rules was proposed for post-processing. The system was trained to disambiguate descriptions of words in relation to Parts of Speech tags together with the full morphological information in terms of values for the different grammatical categories. We present also evaluation of several model variants on the gold standard annotated CMC data, comparison to the state-of-the-art taggers for Polish and error analysis. The proposed tagger shows significantly better results in this domain and demonstrates the viability of adaptation.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ClusteringSentence

Methods 이 논문이 사용한 방법론

fastText fastText embeddings exploit subword information to construct word embeddings. Representations are learnt of character $n$-grams, and words represented as the sum of the…

Similar Papers 제목 키워드 기반

ProtagonistTagger -- a Tool for Entity Linkage of Persons in Texts from Various Languages and Domains

2022-03-13 · Weronika Lajewska, Anna Wroblewska

Named entities recognition (NER) and disambiguation (NED) can add semantic context to the recognized named entities in texts. Named entity linkage in texts, regardless of a domain, provides links between the entities men…

NER

PoliTa: A multitagger for Polish

2014-05-01 · LREC 2014 5 · {\L}ukasz Kobyli{\'n}ski

Part-of-Speech (POS) tagging is a crucial task in Natural Language Processing (NLP). POS tags may be assigned to tokens in text manually, by trained linguists, or using algorithmic approaches. Particularly, in the case o…

Part-Of-Speech TaggingPOSPOS TaggingSentiment Analysis+1

Polish -English Statistical Machine Translation of Medical Texts

2015-09-29 · Krzysztof Wołk, Krzysztof Marasek

This new research explores the effects of various training methods on a Polish to English Statistical Machine Translation system for medical texts. Various elements of the EMEA parallel text corpora from the OPUS project…

Machine TranslationPOSPOS TaggingTranslation

Detecting spelling variants in non-standard texts

2017-04-01 · EACL 2017 4 · Fabian Barteld

Spelling variation in non-standard language, e.g. computer-mediated communication and historical texts, is usually treated as a deviation from a standard spelling, e.g. 2mr as an non-standard spelling for tomorrow. Conse…

ADMEDTAGGER: an annotation framework for distillation of expert knowledge for the Polish medical language

2025-12-27 · Franciszek Górski, Andrzej Czyżewski arxiv

In this work, we present an annotation framework that demonstrates how a multilingual LLM pretrained on a large corpus can be used as a teacher model to distill the expert knowledge needed for tagging medical texts in Po…