paper-with-me

홈 › Papers

TaLAPi --- A Thai Linguistically Annotated Corpus for Language Processing

2014-05-01 · LREC 2014 5 · AiTi Aw, Sharifah Mahani Aljunied, Nattadaporn Lertcheva, Sasiwimon Kalunsima

This paper discusses a Thai corpus, TaLAPi, fully annotated with word segmentation (WS), part-of-speech (POS) and named entity (NE) information with the aim to provide a high-quality and sufficiently large corpus for real-life implementation of Thai language processing tools. The corpus contains 2,720 articles (1,043,471words) from the entertainment and lifestyle (NE{\&}L) domain and 5,489 articles (3,181,487 words) in the news (NEWS) domain, with a total of 35 POS tags and 10 named entity categories. In particular, we present an approach to segment and tag foreign and loan words expressed in transliterated or original form in Thai text corpora. We see this as an area for study as adapted and un-adapted foreign language sequences have not been well addressed in the literature and this poses a challenge to the annotation process due to the increasing use and adoption of foreign words in the Thai language nowadays. To reduce the ambiguities in POS tagging and to provide rich information for facilitating Thai syntactic analysis, we adapted the POS tags used in ORCHID and propose a framework to tag Thai text and also addresses the tagging of loan and foreign words based on the proposed segmentation strategy. TaLAPi also includes a detailed guideline for tagging the 10 named entity categories

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesPOSPOS TaggingTAG

Similar Papers 제목 키워드 기반

The Annotation Guideline of LST20 Corpus

2020-08-12 · Prachya Boonkwan, Vorapon Luantangsrisuk, Sitthaa Phaholphinyo, Kanyanat Kriengket 외

This report presents the annotation guideline for LST20, a large-scale corpus with multiple layers of linguistic annotation for Thai language processing. Our guideline consists of five layers of linguistic annotation: wo…

POSPOS TaggingSentence

A Word Labeling Approach to Thai Sentence Boundary Detection and POS Tagging

2016-12-01 · COLING 2016 12 · Nina Zhou, AiTi Aw, Nattadaporn Lertcheva, Xuancong Wang

Previous studies on Thai Sentence Boundary Detection (SBD) mostly assumed sentence ends at a space disambiguation problem, which classified space either as an indicator for Sentence Boundary (SB) or non-Sentence Boundary…

Boundary DetectionMachine TranslationPart-Of-Speech TaggingPOS+2

A Comparative Study of Pretrained Language Models on Thai Social Text Categorization

2019-12-03 · Thanapapas Horsuwan, Kasidis Kanwatchara, Peerapon Vateekul, Boonserm Kijsirikul

The ever-growing volume of data of user-generated content on social media provides a nearly unlimited corpus of unlabeled data even in languages where resources are scarce. In this paper, we demonstrate that state-of-the…

General ClassificationLanguage ModelingLanguage ModellingText Categorization

The SETimes.HR Linguistically Annotated Corpus of Croatian

2014-05-01 · LREC 2014 5 · {\v{Z}}eljko Agi{\'c}, Nikola Ljube{\v{s}}i{\'c}

We present SETimes.HR ― the first linguistically annotated corpus of Croatian that is freely available for all purposes. The corpus is built on top of the SETimes parallel corpus of nine Southeast European languages an…

AllBoundary DetectionDependency ParsingLemmatization+4

Mangosteen: An Open Thai Corpus for Language Model Pretraining

2025-07-19 · Wannaphong Phatthiyaphaibun, Can Udomcharoenchaikit, Pakpoom Singkorapoom, Kunat Pipatanakul 외 arxiv

Pre-training data shapes a language model's quality, but raw web text is noisy and demands careful cleaning. Existing large-scale corpora rely on English-centric or language-agnostic pipelines whose heuristics do not cap…