paper-with-me

홈 › Papers

ViMQ: A Vietnamese Medical Question Dataset for Healthcare Dialogue System Development

2023-04-27 · Ta Duc Huy, Nguyen Anh Tu, Tran Hoang Vu, Nguyen Phuc Minh, Nguyen Phan, Trung H. Bui, Steven Q. H. Truong

Existing medical text datasets usually take the form of ques- tion and answer pairs that support the task of natural language gener- ation, but lacking the composite annotations of the medical terms. In this study, we publish a Vietnamese dataset of medical questions from patients with sentence-level and entity-level annotations for the Intent Classification and Named Entity Recognition tasks. The tag sets for two tasks are in medical domain and can facilitate the development of task- oriented healthcare chatbots with better comprehension of queries from patients. We train baseline models for the two tasks and propose a simple self-supervised training strategy with span-noise modelling that substan- tially improves the performance. Dataset and code will be published at https://github.com/tadeephuy/ViMQ

📄 PDF Abstract BibTeX arXiv:2304.14405

Code (1)

tadeephuy/vimq 공식 구현 pytorch

Tasks

intent-classificationIntent Classificationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)SentenceTAG

Similar Papers 제목 키워드 기반

VIMQA: A Vietnamese Dataset for Advanced Reasoning and Explainable Multi-hop Question Answering

2022-06-01 · LREC 2022 6 · Khang Le, Hien Nguyen, Tung Le Thanh, Minh Nguyen

Vietnamese is the native language of over 98 million people in the world. However, existing Vietnamese Question Answering (QA) datasets do not explore the model’s ability to perform advanced reasoning and provide evidenc…

Multi-hop Question AnsweringQuestion AnsweringSentence

SPBERTQA: A Two-Stage Question Answering System Based on Sentence Transformers for Medical Texts

2022-06-20 · Nhung Thi-Hong Nguyen, Phuong Phan-Dieu Ha, Luan Thanh Nguyen, Kiet Van Nguyen 외

Question answering (QA) systems have gained explosive attention in recent years. However, QA tasks in Vietnamese do not have many datasets. Significantly, there is mostly no dataset in the medical domain. Therefore, we b…

Question AnsweringSentence

Multilingual LLM Prompting Strategies for Medical English-Vietnamese Machine Translation

2025-09-19 · Nhu Vo, Nu-Uyen-Phuong Le, Dung D. Le, Massimo Piccardi 외 arxiv

Medical English-Vietnamese machine translation (En-Vi MT) is essential for healthcare access and communication in Vietnam, yet Vietnamese remains a low-resource and under-studied language. We systematically evaluate prom…

Machine Translation

ViHERMES: A Graph-Grounded Multihop Question Answering Benchmark and System for Vietnamese Healthcare Regulations

2026-02-07 · Long S. T. Nguyen, Quan M. Bui, Tin T. Ngo, Quynh T. N. Vo 외 arxiv

Question Answering (QA) over regulatory documents is inherently challenging due to the need for multihop reasoning across legally interdependent texts, a requirement that is particularly pronounced in the healthcare doma…

Question Answering

ViHealthBERT: Pre-trained Language Models for Vietnamese in Health Text Mining

2022-06-01 · LREC 2022 6 · Minh, Nguyen and Tran, Vu Hoang and Hoang, Vu and Ta 외

Pre-trained language models have become crucial to achieving competitive results across many Natural Language Processing (NLP) problems. For monolingual pre-trained models in low-resource languages, the quantity has been…

Language ModelingLanguage ModellingMedical Named Entity RecognitionNamed Entity Recognition In Vietnamese+3