paper-with-me

Papers

Removing Spurious Correlation from Neural Network Interpretations

2024-12-03 · Milad Fotouhi, Mohammad Taha Bahadori, Oluwaseyi Feyisetan, Payman Arabshahi, David Heckerman

The existing algorithms for identification of neurons responsible for undesired and harmful behaviors do not consider the effects of confounders such as topic of the conversation. In this work, we show that confounders can create spurious correlations and propose a new causal mediation approach that controls the impact of the topic. In experiments with two large language models, we study the localization hypothesis and show that adjusting for the effect of conversation topic, toxicity becomes less localized.

📄 PDF Abstract BibTeX arXiv:2412.02893

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Influence Tuning: Demoting Spurious Correlations via Instance Attribution and Instance-Driven Updates

2021-10-07 · Findings (EMNLP) 2021 11 · Xiaochuang Han, Yulia Tsvetkov

Among the most critical limitations of deep learning NLP models are their lack of interpretability, and their reliance on spurious correlations. Prior work proposed various approaches to interpreting the black-box models…

Seeing What's Not There: Spurious Correlation in Multimodal LLMs

2025-03-11 · Parsa Hosseini, Sumit Nawathe, Mazda Moayeri, Sriram Balasubramanian 외

Unimodal vision models are known to rely on spurious correlations, but it remains unclear to what extent Multimodal Large Language Models (MLLMs) exhibit similar biases despite language supervision. In this paper, we inv…

HallucinationObjectObject HallucinationObject Recognition

Removing Spurious Concepts from Neural Network Representations via Joint Subspace Estimation

2023-10-18 · Floris Holstege, Bram Wouters, Noud van Giersbergen, Cees Diks

Out-of-distribution generalization in neural networks is often hampered by spurious correlations. A common strategy is to mitigate this by removing spurious concepts from the neural network representation of the data. Ex…

Out-of-Distribution Generalization

Identifying and Disentangling Spurious Features in Pretrained Image Representations

2023-06-22 · Rafayel Darbinyan, Hrayr Harutyunyan, Aram H. Markosyan, Hrant Khachatrian

Neural networks employ spurious correlations in their predictions, resulting in decreased performance when these correlations do not hold. Recent works suggest fixing pretrained representations and training a classificat…

Reducing Spurious Correlations for Answer Selection by Feature Decorrelation and Language Debiasing

2022-10-01 · COLING 2022 10 · Zeyi Zhong, Min Yang, Ruifeng Xu

Deep neural models have become the mainstream in answer selection, yielding state-of-the-art performance. However, these models tend to rely on spurious correlations between prediction labels and input features, which in…

Answer SelectionContrastive Learning