paper-with-me

홈 › Papers

RATE: Causal Explainability of Reward Models with Imperfect Counterfactuals

2024-10-15 · David Reber, Sean Richardson, Todd Nief, Cristina Garbacea, Victor Veitch

Reward models are widely used as proxies for human preferences when aligning or evaluating LLMs. However, reward models are black boxes, and it is often unclear what, exactly, they are actually rewarding. In this paper we develop Rewrite-based Attribute Treatment Estimator (RATE) as an effective method for measuring the sensitivity of a reward model to high-level attributes of responses, such as sentiment, helpfulness, or complexity. Importantly, RATE measures the causal effect of an attribute on the reward. RATE uses LLMs to rewrite responses to produce imperfect counterfactuals examples that can be used to measure causal effects. A key challenge is that these rewrites are imperfect in a manner that can induce substantial bias in the estimated sensitivity of the reward model to the attribute. The core idea of RATE is to adjust for this imperfect-rewrite effect by rewriting twice. We establish the validity of the RATE procedure and show empirically that it is an effective estimator.

📄 PDF Abstract BibTeX arXiv:2410.11348

Code (1)

toddnief/rate 공식 구현 pytorch

Tasks

AttributeLanguage ModelingLanguage ModellingSensitivity

Methods 이 논문이 사용한 방법론

Counterfactuals 설명 없음

Similar Papers 제목 키워드 기반

LIBERTy: A Causal Framework for Benchmarking Concept-Based Explanations of LLMs with Structural Counterfactuals

2026-01-15 · Gilat Toker, Nitay Calderon, Ohad Amosy, Roi Reichart arxiv

Concept-based explanations quantify how high-level concepts (e.g., gender or experience) influence model behavior, which is crucial for decision-makers in high-stakes domains. Recent work evaluates the faithfulness of su…

Text Generation

CausalKG: Causal Knowledge Graph Explainability using interventional and counterfactual reasoning

2022-01-06 · Utkarshani Jaimini, Amit Sheth

Humans use causality and hypothetical retrospection in their daily decision-making, planning, and understanding of life events. The human mind, while retrospecting a given situation, think about questions such as "What w…

counterfactualCounterfactual ReasoningDecision MakingKnowledge Graphs

Application of Causal Inference to Analytical Customer Relationship Management in Banking and Insurance

2022-08-19 · Satyam Kumar, Vadlamani Ravi

Of late, in order to have better acceptability among various domain, researchers have argued that machine intelligence algorithms must be able to provide explanations that humans can understand causally. This aspect, als…

Causal InferenceFraud DetectionManagement

Towards Characterizing Domain Counterfactuals For Invertible Latent Causal Models

2023-06-20 · Zeyu Zhou, Ruqi Bai, Sean Kulinski, Murat Kocaoglu 외

Answering counterfactual queries has important applications such as explainability, robustness, and fairness but is challenging when the causal variables are unobserved and the observations are non-linear mixtures of the…

Causal DiscoverycounterfactualFairness

Causal Proxy Models for Concept-Based Model Explanations

2022-09-28 · Zhengxuan Wu, Karel D'Oosterlinck, Atticus Geiger, Amir Zur 외

Explainability methods for NLP systems encounter a version of the fundamental problem of causal inference: for a given ground-truth input text, we never truly observe the counterfactual texts necessary for isolating the …

Causal Inferencecounterfactualmodel