paper-with-me

Papers

Interventional Probing in High Dimensions: An NLI Case Study

2023-04-20 · Julia Rozanova, Marco Valentino, Lucas Cordeiro, Andre Freitas

Probing strategies have been shown to detect the presence of various linguistic features in large language models; in particular, semantic features intermediate to the "natural logic" fragment of the Natural Language Inference task (NLI). In the case of natural logic, the relation between the intermediate features and the entailment label is explicitly known: as such, this provides a ripe setting for interventional studies on the NLI models' representations, allowing for stronger causal conjectures and a deeper critical analysis of interventional probing methods. In this work, we carry out new and existing representation-level interventions to investigate the effect of these semantic features on NLI classification: we perform amnesic probing (which removes features as directed by learned linear probes) and introduce the mnestic probing variation (which forgets all dimensions except the probe-selected ones). Furthermore, we delve into the limitations of these methods and outline some pitfalls have been obscuring the effectivity of interventional probing studies.

📄 PDF Abstract BibTeX arXiv:2304.10346

Code (0)

등록된 구현이 없습니다.

Tasks

Natural Language InferenceVocal Bursts Intensity Prediction

Similar Papers 제목 키워드 기반

Ancestral causal learning in high dimensions with a human genome-wide application

2019-05-27 · Umberto Noè, Bernd Taschler, Joachim Täger, Peter Heutink 외

We consider learning ancestral causal relationships in high dimensions. Our approach is driven by a supervised learning perspective, with discrete indicators of causal relationships treated as labels to be learned from a…

Vocal Bursts Intensity Prediction

Everything that can be learned about a causal structure with latent variables by observational and interventional probing schemes

2024-07-01 · Marina Maciel Ansanelli, Elie Wolfe, Robert W. Spekkens

What types of differences among causal structures with latent variables are impossible to distinguish by statistical data obtained by probing each visible variable? If the probing scheme is simply passive observation, th…

Probing in the Wild: A Case Study of Self-Supervised Speech Representations on Mandarin Sub-dialects with Unsupervised Articulatory Analysis

2026-06-24 · Shu Shang, Fuliang Weng, Zeqian Hu, Yaqian Zhou arxiv

While self-supervised speech models have achieved strong performance across speech tasks, relatively little is known about how their internal phonetic representations behave under fine-grained dialect variation. Existing…

Discovering What You Can Control: Interventional Boundary Discovery for Reinforcement Learning

2026-03-18 · Jiaxin Liu, Anzhe Cheng, Paul Bogdan arxiv

When an RL agent's observations contain distractors driven by the same confounders as its true state, observational data alone cannot identify which dimensions the agent controls. In our benchmarks, even state-conditione…

Reinforcement Learning

Do We Know What LLMs Don't Know? A Study of Consistency in Knowledge Probing

2025-05-27 · Raoyuan Zhao, Abdullatif Köksal, Ali Modarressi, Michael A. Hedderich 외

The reliability of large language models (LLMs) is greatly compromised by their tendency to hallucinate, underscoring the need for precise identification of knowledge gaps within LLMs. Various methods for probing such ga…

Knowledge Probing