What if you said that differently?: How Explanation Formats Affect Human Feedback Efficacy and User Perception
Eliciting feedback from end users of NLP models can be beneficial for improving models. However, how should we present model responses to users so they are most amenable to be corrected from user feedback? Further, what properties do users value to understand and trust responses? We answer these questions by analyzing the effect of rationales (or explanations) generated by QA models to support their answers. We specifically consider decomposed QA models that first extract an intermediate rationale based on a context and a question and then use solely this rationale to answer the question. A rationale outlines the approach followed by the model to answer the question. Our work considers various formats of these rationales that vary according to well-defined properties of interest. We sample rationales from language models using few-shot prompting for two datasets, and then perform two user studies. First, we present users with incorrect answers and corresponding rationales in various formats and ask them to provide natural language feedback to revise the rationale. We then measure the effectiveness of this feedback in patching these rationales through in-context learning. The second study evaluates how well different rationale formats enable users to understand and trust model answers, when they are correct. We find that rationale formats significantly affect how easy it is (1) for users to give feedback for rationales, and (2) for models to subsequently execute this feedback. In addition, formats with attributions to the context and in-depth reasoning significantly enhance user-reported understanding and trust of model outputs.
Code (2)
Tasks
In-Context LearningQuestion AnsweringReading ComprehensionSimilar Papers 제목 키워드 기반
Retail Pricing Format and Rigidity of Regular Prices
We study the price rigidity of regular and sale prices, and how it is affected by pricing formats (pricing strategies). We use data from three large Canadian stores with different pricing formats (Every-Day-Low-Price, Hi…
It's not what you said, it's how you said it: discriminative perception of speech as a multichannel communication system
People convey information extremely effectively through spoken interaction using multiple channels of information transmission: the lexical channel of what is said, and the non-lexical channel of how it is said. We propo…
Supporting Calibrated Reliance in Human-AI Collaboration: Different Strategies for Different Tasks
As AI systems increasingly support human decision making, a central challenge is determining what information helps people recognize when to rely on AI predictions and when to question or override them. Across three cont…
Logical ReasoningVisual ReasoningMeasuring the Impact of Explanation Bias: A Study of Natural Language Justifications for Recommender Systems
Despite the potential impact of explanations on decision making, there is a lack of research on quantifying their effect on users' choices. This paper presents an experimental protocol for measuring the degree to which p…
Decision MakingRecommendation SystemsNot All Explanations Simulate Equally: Comparing Verbalized Feature Attributions and Self-Generated Rationales
Natural-language explanations are often treated as a unified interface for understanding model behavior, but different explanation sources may support simulation in different ways. This paper compares two families of exp…
Question Answering