paper-with-me

홈 › Papers

Does My Representation Capture X? Probe-Ably

2021-04-12 · ACL 2021 5 · Deborah Ferreira, Julia Rozanova, Mokanarangan Thayaparan, Marco Valentino, André Freitas

Probing (or diagnostic classification) has become a popular strategy for investigating whether a given set of intermediate features is present in the representations of neural models. Probing studies may have misleading results, but various recent works have suggested more reliable methodologies that compensate for the possible pitfalls of probing. However, these best practices are numerous and fast-evolving. To simplify the process of running a set of probing experiments in line with suggested methodologies, we introduce Probe-Ably: an extendable probing framework which supports and automates the application of probing methods to the user's inputs.

📄 PDF Abstract BibTeX arXiv:2104.05807

Code (1)

ai-systems/Probe-Ably 공식 구현 pytorch

Tasks

Diagnostic

Similar Papers 제목 키워드 기반

Before the Last Token: Diagnosing Final-Token Safety Probe Failures

2026-05-12 · Shravan Doda arxiv

Final-token safety probes monitor a single hidden state after prompt prefill, but jailbreak prompts can contain probe-visible unsafe evidence distributed across earlier user-token representations that is missed by this r…

Rhetorical Questions in LLM Representations: A Linear Probing Study

2026-04-15 · Louie Hong Yao, Vishesh Anand, Yuan Zhuang, Tianyu Jiang arxiv

Rhetorical questions are asked not to seek information but to persuade or signal stance. How large language models internally represent them remains unclear. We analyze rhetorical questions in LLM representations using l…

Quantifying and Mitigating Korean Jamo-Level Typographical Vulnerabilities in Large Language Models

2026-08-31 · Seojin Lee, Hwanhee Lee arxiv

Korean introduces an additional typographical perturbation level not captured by ordinary character-level edit models: because syllable blocks are internally composed of sub-character units called jamo, keyboard-level er…

Grammatical Error Correction

When Does Syntax Mediate Neural Language Model Performance? Evidence from Dropout Probes

2022-04-20 · NAACL 2022 7 · Mycal Tucker, Tiwalayo Eisape, Peng Qian, Roger Levy 외

Recent causal probing literature reveals when language models and syntactic probes use similar representations. Such techniques may yield "false negative" causality results: models may use representations of syntax, but …

Language ModelingLanguage Modelling

When Does Syntax Mediate Neural Language Model Performance? Evidence from Dropout Probes

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Recent causal probing literature reveals when language models and syntactic probes use similar representations. Such techniques may yield ``false negative'' causality results: models may use representations of syntax, bu…

Language ModelingLanguage Modelling