paper-with-me

홈 › Papers

Focus and Dilution: The Multi-stage Learning Process of Attention

2026-05-02 · Zheng-An Chen, Pengxiao Lin, Zhi-Qin John Xu, Tao Luo arxiv

Transformer-based models have achieved remarkable success across a wide range of domains, yet our understanding of their training dynamics remains limited. In this work, we identify a recurrent focus-dilution cycle in attention learning and provide a rigorous explanation in a one-layer Transformer setting for Markovian data via gradient-flow analysis. Using stage-wise linearization around critical points, we show that a single focus-dilution cycle can be decomposed into a sequence of distinct stages. First, embedding and projection rapidly condense to a rank-one structure, while attention parameters remain effectively frozen. Then, the attention parameters begin to increase, inducing a frequency-driven focus toward high-frequency tokens. As attention continues to evolve, it generates next-order perturbations in embeddings, leading to a mass-redistribution mechanism that progressively dilutes this focus. Finally, small asymmetries among low-frequency tokens lift a degenerate critical point, opening new embedding directions and initiating the next cycle. Experiments on synthetic Markovian data as well as WikiText and TinyStories corroborate the predicted stages and cyclical dynamics.

📄 PDF Abstract BibTeX arXiv:2605.01199

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Relieving the Over-Aggregating Effect in Graph Transformers

2025-10-24 · Junshu Sun, Wanxing Chang, Chenxue Yang, Qingming Huang 외 arxiv

Graph attention has demonstrated superior performance in graph learning tasks. However, learning from global interactions can be challenging due to the large number of nodes. In this paper, we discover a new phenomenon t…

Graph Learning

FocuSFT: Bilevel Optimization for Dilution-Aware Long-Context Fine-Tuning

2026-05-11 · Zehua Pei, Hui-Ling Zhen, Xianzhi Yu, Sinno Jialin Pan 외 arxiv

Large language models can now process increasingly long inputs, yet their ability to effectively use information spread across long contexts remains limited. We trace this gap to how attention budget is spent during supe…

Bilevel Optimization

Global Context-Aware Progressive Aggregation Network for Salient Object Detection

2020-03-02 · Zuyao Chen, Qianqian Xu, Runmin Cong, Qingming Huang

Deep convolutional neural networks have achieved competitive performance in salient object detection, in which how to learn effective and comprehensive features plays a critical role. Most of the previous works mainly ad…

Dichotomous Image Segmentationobject-detectionObject DetectionRGB Salient Object Detection+1

Predicting Invoice Dilution in Supply Chain Finance with Leakage Free Two Stage XGBoost, KAN (Kolmogorov Arnold Networks), and Ensemble Models

2026-02-16 · Pavel Koptev, Vishnu Kumar, Konstantin Malkov, George Shapiro 외 arxiv

Invoice or payment dilution is the gap between the approved invoice amount and the actual collection is a significant source of non credit risk and margin loss in supply chain finance. Traditionally, this risk is managed…

Zoom In, Reason Out: Efficient Far-field Anomaly Detection in Expressway Surveillance Videos via Focused VLM Reasoning Guided by Bayesian Inference

2026-04-26 · Xiaowei Mao, Bowen Sui, Weijie Zhang, Yawen Yang 외 arxiv

Expressway video anomaly detection is essential for safety management. However, identifying anomalies across diverse scenes remains challenging, particularly for far-field targets exhibiting subtle abnormal vehicle motio…

Video Anomaly DetectionBayesian Inference