paper-with-me

홈 › Papers

CHARP: Conversation History AwaReness Probing for Knowledge-grounded Dialogue Systems

2024-05-24 · Abbas Ghaddar, David Alfonso-Hermelo, Philippe Langlais, Mehdi Rezagholizadeh, Boxing Chen, Prasanna Parthasarathi

In this work, we dive deep into one of the popular knowledge-grounded dialogue benchmarks that focus on faithfulness, FaithDial. We show that a significant portion of the FaithDial data contains annotation artifacts, which may bias models towards completely ignoring the conversation history. We therefore introduce CHARP, a diagnostic test set, designed for an improved evaluation of hallucinations in conversational model. CHARP not only measures hallucination but also the compliance of the models to the conversation task. Our extensive analysis reveals that models primarily exhibit poor performance on CHARP due to their inability to effectively attend to and reason over the conversation history. Furthermore, the evaluation methods of FaithDial fail to capture these shortcomings, neglecting the conversational history. Our findings indicate that there is substantial room for contribution in both dataset creation and hallucination evaluation for knowledge-grounded dialogue, and that CHARP can serve as a tool for monitoring the progress in this particular research area. CHARP is publicly available at https://huggingface.co/datasets/huawei-noah/CHARP

📄 PDF Abstract BibTeX arXiv:2405.15110

Code (0)

등록된 구현이 없습니다.

Tasks

DiagnosticHallucinationHallucination Evaluation

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

CharPoet: A Chinese Classical Poetry Generation System Based on Token-free LLM

2024-01-07 · Chengyue Yu, Lei Zang, Jiaotuan Wang, Chenyi Zhuang 외

Automatic Chinese classical poetry generation has attracted much research interest, but achieving effective control over format and content simultaneously remains challenging. Traditional systems usually accept keywords …

Language Modelling

Enhancing LLM-Based Human-Robot Interaction with Nuances for Diversity Awareness

2024-06-25 · Lucrezia Grassi, Carmine Tommaso Recchiuto, Antonio Sgorbissa

This paper presents a system for diversity-aware autonomous conversation leveraging the capabilities of large language models (LLMs). The system adapts to diverse populations and individuals, considering factors like bac…

Diversity

Machine Learning-based Correlation of Charpy Impact Properties Between Sub-sized and Standard-sized Specimens for Nuclear Structural Materials

2026-07-11 · Yugandhar Kasala Sreenivasulu, Isshu Lee, John W. Merickel, Fei Xu 외 arxiv

Reliable correlations of Charpy impact test results between sub-sized and full-sized specimens are essential for structural integrity assessments, particularly in nuclear applications, where spatial constraints and limit…

Human-like informative conversations: Better acknowledgements using conditional mutual information

2021-04-16 · NAACL 2021 4 · Ashwin Paranjape, Christopher D. Manning

This work aims to build a dialogue agent that can weave new factual content into conversations as naturally as humans. We draw insights from linguistic principles of conversational analysis and annotate human-human conve…

Specificity

Learning to Trust Your Feelings: Leveraging Self-awareness in LLMs for Hallucination Mitigation

2024-01-27 · Yuxin Liang, Zhuoyang Song, Hao Wang, Jiaxing Zhang

We evaluate the ability of Large Language Models (LLMs) to discern and express their internal knowledge state, a key factor in countering factual hallucination and ensuring reliable application of LLMs. We observe a robu…

HallucinationKnowledge Probingreinforcement-learningReinforcement Learning