paper-with-me

홈 › Papers

Breaking the Correlation Plateau: On the Optimization and Capacity Limits of Attention-Based Regressors

2026-02-19 · Jingquan Yan, Yuwei Miao, Peiran Yu, Junzhou Huang arxiv

Attention-based regression models are often trained by jointly optimizing Mean Squared Error (MSE) loss and Pearson correlation coefficient (PCC) loss, emphasizing the magnitude of errors and the order or shape of targets, respectively. A common but poorly understood phenomenon during training is the PCC plateau: PCC stops improving early in training, even as MSE continues to decrease. We provide the first rigorous theoretical analysis of this behavior, revealing fundamental limitations in both optimization dynamics and model capacity. First, in regard to the flattened PCC curve, we uncover a critical conflict where lowering MSE (magnitude matching) can paradoxically suppress the PCC gradient (shape matching). This issue is exacerbated by the softmax attention mechanism, particularly when the data to be aggregated is highly homogeneous. Second, we identify a limitation in the model capacity: we derived a PCC improvement limit for any convex aggregator (including the softmax attention), showing that the convex hull of the inputs strictly bounds the achievable PCC gain. We demonstrate that data homogeneity intensifies both limitations. Motivated by these insights, we propose the Extrapolative Correlation Attention (ECA), which incorporates novel, theoretically-motivated mechanisms to improve the PCC optimization and extrapolate beyond the convex hull. Across diverse benchmarks, including challenging homogeneous data setting, ECA consistently breaks the PCC plateau, achieving significant improvements in correlation without compromising MSE performance.

📄 PDF Abstract BibTeX arXiv:2602.17898

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Routing Plateau: Understanding and Breaking the Accuracy Limits of LLM Routers

2026-05-27 · Yifan Lu, Qiyue Zhang, Shenrun Zhang, Zhibo Yu 외 arxiv

LLM routing has become a popular approach to improve the cost-quality trade-off of LLM services by dynamically selecting a model for each query. Recent work has explored a broad range of routing methods, including cluste…

Breaking Through Barren Plateaus: Reinforcement Learning Initializations for Deep Variational Quantum Circuits

2025-08-25 · Yifeng Peng, Xinyi Li, Zhemin Zhang, Samuel Yen-Chi Chen 외 arxiv

Variational Quantum Algorithms (VQAs) have gained prominence as a viable framework for exploiting near-term quantum devices in applications ranging from optimization and chemistry simulation to machine learning. However,…

Reinforcement Learning

Breaking the Precision Ceiling in Physics-Informed Neural Networks: A Hybrid Fourier-Neural Architecture for Ultra-High Accuracy

2025-07-28 · Wei Shan Lee, Chi Kiu Althina Chau, Kei Chon Sio, Kam Ian Leong arxiv

Physics-informed neural networks (PINNs) have plateaued at errors of $10^{-3}$-$10^{-4}$ for fourth-order partial differential equations, creating a perceived precision ceiling that limits their adoption in engineering a…

Symmetry Breaking in Neural Network Optimization: Insights from Input Dimension Expansion

2024-09-10 · Jun-Jie Zhang, Nan Cheng, Fu-Peng Li, Xiu-Cheng Wang 외

Understanding the mechanisms behind neural network optimization is crucial for improving network design and performance. While various optimization techniques have been developed, a comprehensive understanding of the und…

Limits of Difficulty Scaling: Hard Samples Yield Diminishing Returns in GRPO-Tuned SLMs

2026-04-07 · Suraj Yadav, Siddharth Yadav, Parth Goyal arxiv

Recent alignment work on Large Language Models (LLMs) suggests preference optimization can improve reasoning by shifting probability mass toward better solutions. We test this claim in a resource-constrained setting by a…