paper-with-me

홈 › Papers

Circular Reasoning: Understanding Self-Reinforcing Loops in Large Reasoning Models

2026-01-09 · Zenghao Duan, Liang Pang, Zihao Wei, Wenbin Duan, Yuxin Tian, Shicheng Xu, Jingcheng Deng, Zhiyi Yin, Xueqi Cheng arxiv

Despite the success of test-time scaling, Large Reasoning Models (LRMs) frequently encounter repetitive loops that lead to computational waste and inference failure. In this paper, we identify a distinct failure mode termed Circular Reasoning. Unlike traditional model degeneration, this phenomenon manifests as a self-reinforcing trap where generated content acts as a logical premise for its own recurrence, compelling the reiteration of preceding text. To systematically analyze this phenomenon, we introduce LoopBench, a dataset designed to capture two distinct loop typologies: numerical loops and statement loops. Mechanistically, we characterize circular reasoning as a state collapse exhibiting distinct boundaries, where semantic repetition precedes textual repetition. We reveal that reasoning impasses trigger the loop onset, which subsequently persists as an inescapable cycle driven by a self-reinforcing V-shaped attention mechanism. Guided by these findings, we employ the Cumulative Sum (CUSUM) algorithm to capture these precursors for early loop prediction. Experiments across diverse LRMs validate its accuracy and elucidate the stability of long-chain reasoning.

📄 PDF Abstract BibTeX arXiv:2601.05693

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Closure of Self-Determining System Based on Causal and Constitutive Relations

2026-06-19 · Yoshiyuki Ohmura, Earnest Kota Carr, Yasuo Kuniyoshi arxiv

A self-determining system is defined as one in which causes originating within the system influence the system itself. This definition raises the question of how to specify system boundaries. Although the concept of "clo…

Towards Fair Personalization by Avoiding Feedback Loops

2020-12-20 · Gökhan Çapan, Özge Bozal, İlker Gündoğdu, Ali Taylan Cemgil

Self-reinforcing feedback loops are both cause and effect of over and/or under-presentation of some content in interactive recommender systems. This leads to erroneous user preference estimates, namely, overestimation of…

Recommendation Systems

A Bayesian Choice Model for Eliminating Feedback Loops

2019-08-15 · Gökhan Çapan, Ilker Gündoğdu, Ali Caner Türkmen, Çağrı Sofuoğlu 외

Self-reinforcing feedback loops in personalization systems are typically caused by users choosing from a limited set of alternatives presented systematically based on previous choices. We propose a Bayesian choice model …

Recommendation SystemsThompson Sampling

Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL

2026-08-18 · Yunhao Yang, Yuexin Bian, Yunjie Tian, Di Fu 외 arxiv

Reinforcement learning (RL) has emerged as a powerful approach for improving reasoning in language and vision-language models, yet its strongest successes still depend heavily on ground-truth supervision (e.g., verifiabl…

Reinforcement Learning

OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention

2026-02-05 · Zhangquan Chen, Jiale Tao, Ruihuang Li, Yihao Hu 외 arxiv

While humans perceive the world through diverse modalities that operate synergistically to support a holistic understanding of their surroundings, existing omnivideo models still face substantial challenges on audio-visu…

Self-Supervised LearningContrastive LearningVisual Reasoning