paper-with-me

홈 › Papers

The Attentional White Bear Effect in Transformer Language Models

2026-05-27 · Rebecca Ramnauth, Brian Scassellati arxiv

Instruction-based suppression is widely used to prevent language models from generating prohibited content, yet it remains unclear whether suppression reduces internal representation or merely suppresses expression. We investigate this question through representational probing, attention analysis, and behavioral semantic leakage experiments across multiple transformer models. We find that prohibited concepts remain highly recoverable from hidden representations under suppression, continue to influence attention routing, and measurably shape downstream generations despite successful lexical avoidance. These effects persist across pooling strategies, indirect semantic controls, and multiple model families. Our results expose a fundamental gap between behavioral and representational alignment.

📄 PDF Abstract BibTeX arXiv:2605.28639

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Is Attentional Channel Processing Design Required? Comprehensive Analysis Of Robustness Between Vision Transformers And Fully Attentional Networks

2023-06-08 · Abhishri Ajit Medewar, Swanand Ashokrao Kavitkar

The robustness testing has been performed for standard CNN models and Vision Transformers, however there is a lack of comprehensive study between the robustness of traditional Vision Transformers without an extra attenti…

Local alterations of left arcuate fasciculus and transcallosal white matter microstructure in schizophrenia patients with medication-resistant auditory verbal hallucinations: A pilot study

2022-11-17 · Fanny Thomas, Cécile Gallea, Virginie Moulier, Noomane Bouaziz 외

Auditory verbal hallucinations (AVH) in schizophrenia (SZ) have been associated with abnormalities of the left arcuate fasciculus and transcallosal white matter projections linking homologous language areas of both hemis…

Co-Scale Conv-Attentional Image Transformers

2021-04-13 · ICCV 2021 10 · Weijian Xu, Yifan Xu, Tyler Chang, Zhuowen Tu

In this paper, we present Co-scale conv-attentional image Transformers (CoaT), a Transformer-based image classifier equipped with co-scale and conv-attentional mechanisms. First, the co-scale mechanism maintains the inte…

Instance Segmentationobject-detectionObject DetectionSemantic Segmentation

It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization

2025-04-17 · Ali Behrouz, Meisam Razaviyayn, Peilin Zhong, Vahab Mirrokni

Designing efficient and effective architectural backbones has been in the core of research efforts to enhance the capability of foundation models. Inspired by the human cognitive phenomenon of attentional bias-the natura…

AllLanguage ModelingLanguage ModellingMemorization

Do not think about pink elephant!

2024-04-22 · Kyomin Hwang, Suyoung Kim, JunHoo Lee, Nojun Kwak

Large Models (LMs) have heightened expectations for the potential of general AI as they are akin to human intelligence. This paper shows that recent large models such as Stable Diffusion and DALL-E3 also share the vulner…