paper-with-me

홈 › Papers

Contextual Drag: How Errors in the Context Affect LLM Reasoning

2026-02-04 · Yun Cheng, Xingyu Zhu, Haoyu Zhao, Sanjeev Arora arxiv

Central to many self-improvement pipelines for large language models (LLMs) is the assumption that models can improve by reflecting on past mistakes. We study a phenomenon termed contextual drag: the presence of failed attempts in the context biases subsequent generations toward structurally similar errors. Across evaluations of 11 proprietary and open-weight models on 8 reasoning tasks, contextual drag induces 10-20% performance drops, and iterative self-refinement in models with severe contextual drag can collapse into self-deterioration. Structural analysis using tree edit distance reveals that subsequent reasoning trajectories inherit structurally similar error patterns from the context. We demonstrate that neither external feedback nor successful self-verification suffices to eliminate this effect. While mitigation strategies such as fallback-behavior fine-tuning and context denoising yield partial improvements, they fail to fully restore baseline performance, positioning contextual drag as a persistent failure mode in current reasoning architectures.

📄 PDF Abstract BibTeX arXiv:2602.04288

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Graph-Based Phonetic Error Correction of Noisy ASR

2026-04-29 · Pratik Rakesh Singh, Mohammadi Zaki, Aneesh Mukkamala, Pankaj Wasnik arxiv

Automatic speech recognition (ASR) systems, despite low overall word error rates, produce residual lexical errors that disproportionately affect semantically critical tokens such as named entities, negations, and sentime…

Graph Neural NetworkSpeech Recognition

ContextDrag: Precise Drag-Based Image Editing via Context-Preserving Token Injection and Position-Aligned Attention

2025-12-09 · Huiguo He, Pengyu Yan, Ziqi Yi, Weizhi Zhong 외 arxiv

Drag-based image editing enables intuitive visual manipulation through point-based drag operations. Existing methods mainly rely on diffusion inversion or pixel-space warping with inpainting. However, inversion inherentl…

Image Editing

Does Reasoning Help LLM Agents Play Dungeons and Dragons? A Prompt Engineering Experiment

2025-10-20 · Patricia Delafuente, Arya Honraopatil, Lara J. Martin arxiv

This paper explores the application of Large Language Models (LLMs) and reasoning to predict Dungeons & Dragons (DnD) player actions and format them as Avrae Discord bot commands. Using the FIREBALL dataset, we evaluated…

Prompt Engineering

What Did My Car Say? Impact of Autonomous Vehicle Explanation Errors and Driving Context On Comfort, Reliance, Satisfaction, and Driving Confidence

2024-09-09 · Robert Kaufman, Aaron Broukhim, David Kirsh, Nadir Weibel

Explanations for autonomous vehicle (AV) decisions may build trust, however, explanations can contain errors. In a simulated driving study (n = 232), we tested how AV explanation errors, driving context characteristics (…

Streaming Drag-Oriented Interactive Video Manipulation: Drag Anything, Anytime!

2025-10-03 · Junbao Zhou, Yuan Zhou, Kesen Zhao, Qingshan Xu 외 arxiv

Achieving streaming, fine-grained control over the outputs of autoregressive video diffusion models remains challenging, making it difficult to ensure that they consistently align with user expectations. To bridge this g…