paper-with-me

홈 › Papers

Vietnamese Legal Information Retrieval in Question-Answering System

2024-09-05 · Thiem Nguyen Ba, Vinh Doan The, Tung Pham Quang, Toan Tran Van

In the modern era of rapidly increasing data volumes, accurately retrieving and recommending relevant documents has become crucial in enhancing the reliability of Question Answering (QA) systems. Recently, Retrieval Augmented Generation (RAG) has gained significant recognition for enhancing the capabilities of large language models (LLMs) by mitigating hallucination issues in QA systems, which is particularly beneficial in the legal domain. Various methods, such as semantic search using dense vector embeddings or a combination of multiple techniques to improve results before feeding them to LLMs, have been proposed. However, these methods often fall short when applied to the Vietnamese language due to several challenges, namely inefficient Vietnamese data processing leading to excessive token length or overly simplistic ensemble techniques that lead to instability and limited improvement. Moreover, a critical issue often overlooked is the ordering of final relevant documents which are used as reference to ensure the accuracy of the answers provided by LLMs. In this report, we introduce our three main modifications taken to address these challenges. First, we explore various practical approaches to data processing to overcome the limitations of the embedding model. Additionally, we enhance Reciprocal Rank Fusion by normalizing order to combine results from keyword and vector searches effectively. We also meticulously re-rank the source pieces of information used by LLMs with Active Retrieval to improve user experience when refining the information generated. In our opinion, this technique can also be considered as a new re-ranking method that might be used in place of the traditional cross encoder. Finally, we integrate these techniques into a comprehensive QA system, significantly improving its performance and reliability

📄 PDF Abstract BibTeX arXiv:2409.13699

Code (0)

등록된 구현이 없습니다.

Tasks

HallucinationInformation RetrievalQuestion AnsweringRAGRe-RankingRetrievalRetrieval-augmented Generation

Similar Papers 제목 키워드 기반

VLSP 2025 MLQA-TSR Challenge: Vietnamese Multimodal Legal Question Answering on Traffic Sign Regulation

2025-10-23 · Son T. Luu, Trung Vo, Hiep Nguyen, Khanh Quoc Tran 외 arxiv

This paper presents the VLSP 2025 MLQA-TSR - the multimodal legal question answering on traffic sign regulation shared task at VLSP 2025. VLSP 2025 MLQA-TSR comprises two subtasks: multimodal legal retrieval and multimod…

Question Answering

Improving Vietnamese Legal Document Retrieval using Synthetic Data

2024-12-01 · Son Pham Tien, Hieu Nguyen Doan, An Nguyen Dai, Sang Dinh Viet

In the field of legal information retrieval, effective embedding-based models are essential for accurate question-answering systems. However, the scarcity of large annotated datasets poses a significant challenge, partic…

Information RetrievalQuestion AnsweringRetrieval

Answering Legal Questions by Learning Neural Attentive Text Representation

2020-12-01 · COLING 2020 8 · Phi Manh Kien, Ha-Thanh Nguyen, Ngo Xuan Bach, Vu Tran 외

Text representation plays a vital role in retrieval-based question answering, especially in the legal domain where documents are usually long and complicated. The better the question and the legal documents are represent…

ArticlesQuestion AnsweringRetrievalvalid

Improving Vietnamese Legal Question--Answering System based on Automatic Data Enrichment

2023-06-08 · Thi-Hai-Yen Vuong, Ha-Thanh Nguyen, Quang-Huy Nguyen, Le-Minh Nguyen 외

Question answering (QA) in law is a challenging problem because legal documents are much more complicated than normal texts in terms of terminology, structure, and temporal and logical relationships. It is even more diff…

Question AnsweringRetrieval

VLQA: The First Comprehensive, Large, and High-Quality Vietnamese Dataset for Legal Question Answering

2025-07-26 · Tan-Minh Nguyen, Hoang-Trung Nguyen, Trong-Khoi Dao, Xuan-Hieu Phan 외 arxiv

The advent of large language models (LLMs) has led to significant achievements in various domains, including legal text processing. Leveraging LLMs for legal tasks is a natural evolution and an increasingly compelling ch…

Information RetrievalQuestion Answering