paper-with-me

홈 › Papers

ViHealthBERT: Pre-trained Language Models for Vietnamese in Health Text Mining

2022-06-01 · LREC 2022 6 · Minh, Nguyen and Tran, Vu Hoang and Hoang, Vu and Ta, Huy Duc and Bui, Trung Huu and Truong, Steven Quoc Hung

Pre-trained language models have become crucial to achieving competitive results across many Natural Language Processing (NLP) problems. For monolingual pre-trained models in low-resource languages, the quantity has been significantly increased. However, most of them relate to the general domain, and there are limited strong baseline language models for domain-specific. We introduce ViHealthBERT, the first domain-specific pre-trained language model for Vietnamese healthcare. The performance of our model shows strong results while outperforming the general domain language models in all health-related datasets. Moreover, we also present Vietnamese datasets for the healthcare domain for two tasks are Acronym Disambiguation (AD) and Frequently Asked Questions (FAQ) Summarization. We release our ViHealthBERT to facilitate future research and downstream application for Vietnamese NLP in domain-specific.

📄 PDF Abstract BibTeX

Code (1)

demdecuong/vihealthbert pytorch

Tasks

Language ModelingLanguage ModellingMedical Named Entity RecognitionNamed Entity Recognition In VietnameseText SummarizationVietnamese DatasetsWord Sense Disambiguation

Similar Papers 제목 키워드 기반

ViHealthBERT: Pre-trained Language Models for Vietnamese in Health Text Mining

2022-06-01 · LREC 2022 6 · Nguyen Minh, Vu Hoang Tran, Vu Hoang, Huy Duc Ta 외

Pre-trained language models have become crucial to achieving competitive results across many Natural Language Processing (NLP) problems. For monolingual pre-trained models in low-resource languages, the quantity has been…

Language ModelingLanguage ModellingVietnamese Datasets

ViSoBERT: A Pre-Trained Language Model for Vietnamese Social Media Text Processing

2023-10-17 · Quoc-Nam Nguyen, Thang Chau Phan, Duc-Vu Nguyen, Kiet Van Nguyen

English and Chinese, known as resource-rich languages, have witnessed the strong development of transformer-based language models for natural language processing tasks. Although Vietnam has approximately 100M people spea…

Language ModelingLanguage ModellingVietnamese Hate Speech DetectionVietnamese Language Models+2

ViT5: Pretrained Text-to-Text Transformer for Vietnamese Language Generation

2022-05-13 · NAACL (ACL) 2022 7 · Long Phan, Hieu Tran, Hieu Nguyen, Trieu H. Trinh

We present ViT5, a pretrained Transformer-based encoder-decoder model for the Vietnamese language. With T5-style self-supervised pretraining, ViT5 is trained on a large corpus of high-quality and diverse Vietnamese texts…

Abstractive Text SummarizationDecodernamed-entity-recognitionNamed Entity Recognition+4

BamiBERT: A New BERT-based Language Model for Vietnamese

2026-07-02 · Dat Quoc Nguyen, Thinh Pham, Chi Tran, Linh The Nguyen arxiv

In this paper, we introduce BamiBERT, a new BERT-based pre-trained language model for Vietnamese that addresses key limitations of PhoBERT -- the current de facto Vietnamese text encoder. Trained from scratch on a 129GB …

Domain Generalization

KTVIC: A Vietnamese Image Captioning Dataset on the Life Domain

2024-01-16 · Anh-Cuong Pham, Van-Quang Nguyen, Thi-Hong Vuong, Quang-Thuy Ha

Image captioning is a crucial task with applications in a wide range of domains, including healthcare and education. Despite extensive research on English image captioning datasets, the availability of such datasets for …

Image CaptioningVietnamese Image Captioning