paper-with-me

Papers

ViNLI: A Vietnamese Corpus for Studies on Open-Domain Natural Language Inference

2022-10-01 · COLING 2022 10 · Tin Van Huynh, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen

Over a decade, the research field of computational linguistics has witnessed the growth of corpora and models for natural language inference (NLI) for rich-resource languages such as English and Chinese. A large-scale and high-quality corpus is necessary for studies on NLI for Vietnamese, which can be considered a low-resource language. In this paper, we introduce ViNLI (Vietnamese Natural Language Inference), an open-domain and high-quality corpus for evaluating Vietnamese NLI models, which is created and evaluated with a strict process of quality control. ViNLI comprises over 30,000 human-annotated premise-hypothesis sentence pairs extracted from more than 800 online news articles on 13 distinct topics. In this paper, we introduce the guidelines for corpus creation which take the specific characteristics of the Vietnamese language in expressing entailment and contradiction into account. To evaluate the challenging level of our corpus, we conduct experiments with state-of-the-art deep neural networks and pre-trained models on our dataset. The best system performance is still far from human performance (a 14.20% gap in accuracy). The ViNLI corpus is a challenging corpus to accelerate progress in Vietnamese computational linguistics. Our corpus is available publicly for research purposes.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesNatural Language InferenceSentenceVietnamese Natural Language Inference

Similar Papers 제목 키워드 기반

Transformer-Based Contextualized Language Models Joint with Neural Networks for Natural Language Inference in Vietnamese

2024-11-20 · Dat Van-Thanh Nguyen, Tin Van Huynh, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen

Natural Language Inference (NLI) is a task within Natural Language Processing (NLP) that holds value for various AI applications. However, there have been limited studies on Natural Language Inference in Vietnamese that …

Natural Language InferenceXLM-R

KC4MT: A High-Quality Corpus for Multilingual Machine Translation

2022-06-01 · LREC 2022 6 · Vinh Van Nguyen, Ha Nguyen, Huong Thanh Le, Thai Phuong Nguyen 외

The multilingual parallel corpus is an important resource for many applications of natural language processing (NLP). For machine translation, the size and quality of the training corpus mainly affects the quality of the…

Machine TranslationSentenceTranslationVocal Bursts Intensity Prediction

ViWikiFC: Fact-Checking for Vietnamese Wikipedia-Based Textual Knowledge Source

2024-05-13 · Hung Tuan Le, Long Truong To, Manh Trong Nguyen, Kiet Van Nguyen

Fact-checking is essential due to the explosion of misinformation in the media ecosystem. Although false information exists in every language and country, most research to solve the problem mainly concentrated on huge co…

ArticlesFact CheckingFact VerificationLanguage Modelling+3

ViCLSR: A Supervised Contrastive Learning Framework with Natural Language Inference for Natural Language Understanding Tasks

2026-03-22 · Tin Van Huynh, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen arxiv

High-quality text representations are crucial for natural language understanding (NLU), but low-resource languages like Vietnamese face challenges due to limited annotated data. While pre-trained models like PhoBERT and …

Natural Language UnderstandingNatural Language InferenceRepresentation LearningContrastive Learning

MTet: Multi-domain Translation for English and Vietnamese

2022-10-11 · Chinh Ngo, Trieu H. Trinh, Long Phan, Hieu Tran 외

We introduce MTet, the largest publicly available parallel corpus for English-Vietnamese translation. MTet consists of 4.2M high-quality training sentence pairs and a multi-domain test set refined by the Vietnamese resea…

Machine TranslationSentenceTranslation