paper-with-me

Papers

'Neural howlround' in large language models: a self-reinforcing bias phenomenon, and a dynamic attenuation solution

2025-04-07 · Seth Drake

Large language model (LLM)-driven AI systems may exhibit an inference failure mode we term neural howlround,' a self-reinforcing cognitive loop where certain highly weighted inputs become dominant, leading to entrenched response patterns resistant to correction. This paper explores the mechanisms underlying this phenomenon, which is distinct from model collapse and biased salience weighting. We propose an attenuation-based correction mechanism that dynamically introduces counterbalancing adjustments and can restore adaptive reasoning, even in locked-in' AI systems. Additionally, we discuss some other related effects arising from improperly managed reinforcement. Finally, we outline potential applications of this mitigation strategy for improving AI robustness in real-world decision-making tasks.

📄 PDF Abstract BibTeX arXiv:2504.07992

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingLanguage ModelingLanguage ModellingLarge Language Model

Similar Papers 제목 키워드 기반

Bias in the Mirror: Are LLMs opinions robust to their own adversarial attacks ?

2024-10-17 · Virgile Rennard, Christos Xypolopoulos, Michalis Vazirgiannis

Large language models (LLMs) inherit biases from their training data and alignment processes, influencing their responses in subtle ways. While many studies have examined these biases, little work has explored their robu…

Misinformation

Yet another algorithmic bias: A Discursive Analysis of Large Language Models Reinforcing Dominant Discourses on Gender and Race

2025-08-14 · Gustavo Bonil, Simone Hashiguti, Jhessica Silva, João Gondim 외 arxiv

With the advance of Artificial Intelligence (AI), Large Language Models (LLMs) have gained prominence and been applied in diverse contexts. As they evolve into more sophisticated versions, it is essential to assess wheth…

Bias Detection

Modeling Emergent Lexicon Formation with a Self-Reinforcing Stochastic Process

2022-06-22 · Brendon Boldt, David Mortensen

We introduce FiLex, a self-reinforcing stochastic process which models finite lexicons in emergent language experiments. The central property of FiLex is that it is a self-reinforcing process, parallel to the intuition t…

Internal Reasoning vs. External Control: A Thermodynamic Analysis of Sycophancy in Large Language Models

2025-12-16 · Edward Y. Chang arxiv

Large Language Models exhibit sycophancy: prioritizing agreeableness over correctness. Current remedies evaluate reasoning outcomes: RLHF rewards correct answers, self-correction critiques outputs. All require ground tru…

Clustering Discourses: Racial Biases in Short Stories about Women Generated by Large Language Models

2025-09-02 · Gustavo Bonil, João Gondim, Marina dos Santos, Simone Hashiguti 외 arxiv

This study investigates how large language models, in particular LLaMA 3.2-3B, construct narratives about Black and white women in short stories generated in Portuguese. From 2100 texts, we applied computational methods …