paper-with-me

홈 › Papers

A Comparative Study of Pretrained Language Models on Thai Social Text Categorization

2019-12-03 · Thanapapas Horsuwan, Kasidis Kanwatchara, Peerapon Vateekul, Boonserm Kijsirikul

The ever-growing volume of data of user-generated content on social media provides a nearly unlimited corpus of unlabeled data even in languages where resources are scarce. In this paper, we demonstrate that state-of-the-art results on two Thai social text categorization tasks can be realized by pretraining a language model on a large noisy Thai social media corpus of over 1.26 billion tokens and later fine-tuned on the downstream classification tasks. Due to the linguistically noisy and domain-specific nature of the content, our unique data preprocessing steps designed for Thai social media were utilized to ease the training comprehension of the model. We compared four modern language models: ULMFiT, ELMo with biLSTM, OpenAI GPT, and BERT. We systematically compared the models across different dimensions including speed of pretraining and fine-tuning, perplexity, downstream classification benchmarks, and performance in limited pretraining data.

📄 PDF Abstract BibTeX arXiv:1912.01580

Code (0)

등록된 구현이 없습니다.

Tasks

General ClassificationLanguage ModelingLanguage ModellingText Categorization

Methods 이 논문이 사용한 방법론

Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…
Sigmoid Activation 설명 없음
Tanh Activation 설명 없음
Activation Regularization Activation Regularization (AR), or $L\_{2}$ activation regularization, is regularization performed on activations as opposed to weights. It is usually used in conjunction with…
Weight Decay 설명 없음
Residual Connection 설명 없음

Similar Papers 제목 키워드 기반

PhayaThaiBERT: Enhancing a Pretrained Thai Language Model with Unassimilated Loanwords

2023-11-21 · Panyut Sriwirote, Jalinee Thapiang, Vasan Timtong, Attapol T. Rutherford

While WangchanBERTa has become the de facto standard in transformer-based Thai language modeling, it still has shortcomings in regard to the understanding of foreign words, most notably English words, which are often bor…

Language ModelingLanguage Modelling

Quantifying fair income distribution in Thailand

2024-04-15 · Thitithep Sitthiyot, Kanyarat Holasut

Given a vast concern about high income inequality in Thailand as opposed to empirical findings around the world showing people's preference for fair income inequality over unfair income equality, it is therefore importan…

Fairness

Parsing Thai Social Data: A New Challenge for Thai NLP

2020-03-06 · Sattaya Singkul, Borirat Khampingyot, Nattasit Maharattamalai, Supawat Taerungruang 외

Dependency parsing (DP) is a task that analyzes text for syntactic structure and relationship between words. DP is widely used to improve natural language processing (NLP) applications in many languages such as English. …

Dependency Parsing

ChatHealthAI: Aligning Electronic Health Record Representations with Large Language Models for Grounded Clinical Reasoning

2026-06-01 · Bo-Hong Wang, Baicheng Peng, Ruilin Wang, Jun Bai 외 arxiv

Large language models (LLMs) exhibit strong natural-language reasoning abilities for clinical decision support, but struggle to effectively model structured longitudinal electronic health records (EHRs). In contrast, EHR…

Multimodal Reasoning

Thai Universal Dependency Treebank

2024-05-13 · Panyur Sriwirote, Wei Qi Leong, Charin Polpanumas, Santhawat Thanyawong 외

Automatic dependency parsing of Thai sentences has been underexplored, as evidenced by the lack of large Thai dependency treebanks with complete dependency structures and the lack of a published systematic evaluation of …

Dependency Parsing