paper-with-me

Papers

Leveraging Contextual Counterfactuals Toward Belief Calibration

2023-07-13 · Qiuyi, Zhang, Michael S. Lee, Sherol Chen

Beliefs and values are increasingly being incorporated into our AI systems through alignment processes, such as carefully curating data collection principles or regularizing the loss function used for training. However, the meta-alignment problem is that these human beliefs are diverse and not aligned across populations; furthermore, the implicit strength of each belief may not be well calibrated even among humans, especially when trying to generalize across contexts. Specifically, in high regret situations, we observe that contextual counterfactuals and recourse costs are particularly important in updating a decision maker's beliefs and the strengths to which such beliefs are held. Therefore, we argue that including counterfactuals is key to an accurate calibration of beliefs during alignment. To do this, we first segment belief diversity into two categories: subjectivity (across individuals within a population) and epistemic uncertainty (within an individual across different contexts). By leveraging our notion of epistemic uncertainty, we introduce `the belief calibration cycle' framework to more holistically calibrate this diversity of beliefs with context-driven counterfactual reasoning by using a multi-objective optimization. We empirically apply our framework for finding a Pareto frontier of clustered optimal belief strengths that generalize across different contexts, demonstrating its efficacy on a toy dataset for credit decisions.

📄 PDF Abstract BibTeX arXiv:2307.06513

Code (0)

등록된 구현이 없습니다.

Tasks

counterfactualCounterfactual ReasoningDiversity

Methods 이 논문이 사용한 방법론

Counterfactuals 설명 없음

Similar Papers 제목 키워드 기반

Reasoning about Counterfactuals to Improve Human Inverse Reinforcement Learning

2022-03-03 · Michael S. Lee, Henny Admoni, Reid Simmons

To collaborate well with robots, we must be able to understand their decision making. Humans naturally infer other agents' beliefs and desires by reasoning about their observable behavior in a way that resembles inverse …

counterfactualCounterfactual ReasoningDecision Makingreinforcement-learning+2

Retrieval-Augmented Linguistic Calibration

2026-05-19 · Yi-Fan Yeh, Linwei Tao, Minjing Dong, Tao Huang 외 arxiv

Linguistic cues such as "I believe" and "probably" offer an intuitive interface for communicating confidence, yet a generalisable, principled calibration framework for linguistic confidence expressions remains underexplo…

Quantifying the Privacy of Counterfactuals by Leveraging Membership Inference Attacks Against Synthetic Data

2026-06-04 · Maryam Babaei, Yingke Wang, Hadrien Lautraite, Heber H. Arcolezi 외 arxiv

Counterfactuals are typically used in high-stakes decision areas to explain a machine learning model by showing how changes to the user profiles result in the desired outcome. However, explaining the model's decisions th…

The Case for Negative Data: From Crash Reports to Counterfactuals for Reasonable Driving

2025-09-23 · Jay Patrikar, Apoorva Sharma, Sushant Veer, Boyi Li 외 arxiv

Learning-based autonomous driving systems are trained mostly on incident-free data, offering little guidance near safety-performance boundaries. Real crash reports contain precisely the contrastive evidence needed, but t…

Autonomous Driving

Causality Without Causal Models

2025-11-26 · Joseph Y. Halpern, Rafael Pass arxiv

Perhaps the most prominent current definition of (actual) causality is due to Halpern and Pearl. It is defined using causal models (also known as structural equations models). We abstract the definition, extracting its k…