paper-with-me

홈 › Papers

A Geometric View of Counterfactual Behavior: Interaction of Boundary Proximity and Local Support

2026-06-02 · Ioanna Gemou, Matteo Gamba, Randall Balestriero, Ritambhara Singh arxiv

Counterfactual explanations seek small, semantically meaningful changes to an input that alter a model's prediction, and are widely used to interpret and audit machine learning systems. In modern vision, language, and multimodal systems, pretrained encoders map inputs to representation spaces, and downstream classifier heads impose decision boundaries within those spaces. As a result, the feasibility and distance of nearby counterfactuals depend on boundary placement relative to the data. Yet models with similar predictive performance can differ substantially in whether such changes are achievable and how far representations must move. This work examines this variation using a standardized local search probe across several pretrained encoders and linear classifier heads. Results show that despite similar predictive performance, models differ substantially in their counterfactual behavior. Under fixed representations, varying only the classifier head alters counterfactual outcomes while leaving predictive performance largely unchanged. This variation is explained by the interaction of decision-boundary proximity and local data support, which jointly determine whether prediction changes are both feasible and lie in regions supported by the data, and can also improve counterfactual search within fixed models. Together, these findings identify counterfactual behavior as a distinct dimension beyond predictive performance and show that it can be altered without changing accuracy, with implications for model selection, robustness, and the reliability of counterfactual methods.

📄 PDF Abstract BibTeX arXiv:2606.04209

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards Personalized Treatment Plan: Geometrical Model-Agnostic Approach to Counterfactual Explanations

2025-10-27 · Daniel Sin, Milad Toutounchian arxiv

In our article, we describe a method for generating counterfactual explanations in high-dimensional spaces using four steps that involve fitting our dataset to a model, finding the decision boundary, determining constrai…

CAMO: An Agentic Framework for Automated Causal Discovery from Micro Behaviors to Macro Emergence in LLM Agent Simulations

2026-04-16 · Xiangning Yu, Yuwei Guo, Yuqi Hou, Xiao Xue 외 arxiv

LLM-empowered agent simulations are increasingly used to study social emergence, yet the micro-to-macro causal mechanisms behind macro outcomes often remain unclear. This is challenging because emergence arises from inte…

Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations

2025-10-24 · Faisal Hamman, Pasan Dissanayake, Yanjun Fu, Sanghamitra Dutta arxiv

Knowledge distillation is a promising approach to transfer capabilities from complex teacher models to smaller, resource-efficient student models that can be deployed easily, particularly in task-aware scenarios. However…

Knowledge Distillation

Classifier Reconstruction Through Counterfactual-Aware Wasserstein Prototypes

2025-12-11 · Xuan Zhao, Zhuo Cao, Arya Bangun, Hanno Scharr 외 arxiv

Counterfactual explanations provide actionable insights by identifying minimal input changes required to achieve a desired model prediction. Beyond their interpretability benefits, counterfactuals can also be leveraged f…

Counterfactual Interventions Reveal the Causal Effect of Relative Clause Representations on Agreement Prediction

2021-05-14 · CoNLL (EMNLP) 2021 11 · Shauli Ravfogel, Grusha Prasad, Tal Linzen, Yoav Goldberg

When language models process syntactically complex sentences, do they use their representations of syntax in a manner that is consistent with the grammar of the language? We propose AlterRep, an intervention-based method…

counterfactualSentence