paper-with-me

홈 › Papers

Re-Ex: Revising after Explanation Reduces the Factual Errors in LLM Responses

2024-02-27 · Juyeon Kim, Jeongeun Lee, Yoonho Chang, Chanyeol Choi, JunSeong Kim, Jy-yong Sohn

Mitigating hallucination issues is a key challenge that must be overcome to reliably deploy large language models (LLMs) in real-world scenarios. Recently, various methods have been proposed to detect and revise factual errors in LLM-generated texts, in order to reduce hallucination. In this paper, we propose Re-Ex, a method for post-editing LLM-generated responses. Re-Ex introduces a novel reasoning step dubbed as the factual error explanation step. Re-Ex revises the initial response of LLMs using 3-steps : first, external tools are used to retrieve the evidences of the factual errors in the initial LLM response; next, LLM is instructed to explain the problematic parts of the response based on the gathered evidence; finally, LLM revises the initial response using the explanations provided in the previous step. In addition to the explanation step, Re-Ex also incorporates new prompting techniques to reduce the token count and inference time required for the response revision process. Compared with existing methods including FacTool, CoVE, and RARR, Re-Ex provides better detection and revision performance with less inference time and fewer tokens in multiple benchmarks.

📄 PDF Abstract BibTeX arXiv:2402.17097

Code (1)

juyeonnn/reex 공식 구현

Tasks

Hallucination

Similar Papers 제목 키워드 기반

Improving Factual Error Correction by Learning to Inject Factual Errors

2023-12-12 · Xingwei He, Qianru Zhang, A-Long Jin, Jun Ma 외

Factual error correction (FEC) aims to revise factual errors in false claims with minimal editing, making them faithful to the provided evidence. This task is crucial for alleviating the hallucination problem encountered…

Hallucination

On Improving Summarization Factual Consistency from Natural Language Feedback

2022-12-20 · Yixin Liu, Budhaditya Deb, Milagro Teruel, Aaron Halfaker 외

Despite the recent progress in language generation models, their outputs may not always meet user expectations. In this work, we study whether informational feedback in natural language can be leveraged to improve genera…

Text GenerationZero-Shot Learning

SummExecEdit: A Factual Consistency Benchmark in Summarization with Executable Edits

2024-12-17 · Onkar Thorat, Philippe Laban, Chien-Sheng Wu

Detecting factual inconsistencies in summarization is critical, yet existing benchmarks lack the necessary challenge and interpretability for robust evaluation. In this paper, we introduce SummExecEdit, a novel benchmark…

GenAudit: Fixing Factual Errors in Language Model Outputs with Evidence

2024-02-19 · Kundan Krishna, Sanjana Ramprasad, Prakhar Gupta, Byron C. Wallace 외

LLMs can generate factually incorrect statements even when provided access to reference documents. Such errors can be dangerous in high-stakes applications (e.g., document-grounded QA for healthcare or finance). We prese…

Fact CheckingLanguage ModelingLanguage Modelling

BoxingGym: Benchmarking Progress in Automated Experimental Design and Model Discovery

2025-01-02 · Kanishk Gandhi, Michael Y. Li, Lyle Goodyear, Louise Li 외

Understanding the world and explaining it with scientific theories is a central aspiration of artificial intelligence research. Proposing theories, designing experiments to test them, and then revising them based on data…

BenchmarkingExperimental DesignModel Discoveryscientific discovery