paper-with-me

홈 › Papers

PoliTa: A multitagger for Polish

2014-05-01 · LREC 2014 5 · {\L}ukasz Kobyli{\'n}ski

Part-of-Speech (POS) tagging is a crucial task in Natural Language Processing (NLP). POS tags may be assigned to tokens in text manually, by trained linguists, or using algorithmic approaches. Particularly, in the case of annotated text corpora, the quantity of textual data makes it unfeasible to rely on manual tagging and automated methods are used extensively. The quality of such methods is of critical importance, as even 1{\%} tagger error rate results in introducing millions of errors in a corpus consisting of a billion tokens. In case of Polish several POS taggers have been proposed to date, but even the best of the taggers achieves an accuracy of ca. 93{\%}, as measured on the one million subcorpus of the National Corpus of Polish (NCP). As the task of tagging is an example of classification, in this article we introduce a new POS tagger for Polish, which is based on the idea of combining several classifiers to produce higher quality tagging results than using any of the taggers individually.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Part-Of-Speech TaggingPOSPOS TaggingSentiment AnalysisWord Sense Disambiguation

Similar Papers 제목 키워드 기반

Something Old, Something New: Grammar-based CCG Parsing with Transformer Models

2021-09-21 · Stephen Clark

This report describes the parsing problem for Combinatory Categorial Grammar (CCG), showing how a combination of Transformer-based neural models and a symbolic CCG grammar can lead to substantial gains over existing appr…

CCG SupertaggingSentence

Big City Bias: Evaluating the Impact of Metropolitan Size on Computational Job Market Abilities of Language Models

2024-03-12 · Charlie Campanella, Rob van der Goot

Large language models (LLMs) have emerged as a useful technology for job matching, for both candidates and employers. Job matching is often based on a particular geographic location, such as a city or region. However, LL…

PL-MTEB: Polish Massive Text Embedding Benchmark

2024-05-16 · Rafał Poświata, Sławomir Dadas, Michał Perełkiewicz

In this paper, we introduce the Polish Massive Text Embedding Benchmark (PL-MTEB), a comprehensive benchmark for text embeddings in Polish. The PL-MTEB consists of 28 diverse NLP tasks from 5 task types. We adapted the t…

Coupling Coordinated Development among Digital Economy, Regional Innovation and Talent Employment A case study of Hangzhou Metropolitan Circle, China

2023-10-29 · Luyi Qiu

Coordination development across various subsystems, particularly economic, social, cultural, and human resources subsystems, is a key aspect of urban sustainability that has a direct impact on the quality of urbanization…

Apollo: A Sequencing-Technology-Independent, Scalable, and Accurate Assembly Polishing Algorithm

2019-02-12 · Can Firtina, Jeremie S. Kim, Mohammed Alser, Damla Senol Cali 외

Long reads produced by third-generation sequencing technologies are used to construct an assembly (i.e., the subject's genome), which is further used in downstream genome analysis. Unfortunately, long reads have high seq…