paper-with-me

홈 › Papers

Modify Training Directions in Function Space to Reduce Generalization Error

2023-07-25 · Yi Yu, Wenlian Lu, BoYu Chen

We propose theoretical analyses of a modified natural gradient descent method in the neural network function space based on the eigendecompositions of neural tangent kernel and Fisher information matrix. We firstly present analytical expression for the function learned by this modified natural gradient under the assumptions of Gaussian distribution and infinite width limit. Thus, we explicitly derive the generalization error of the learned neural network function using theoretical methods from eigendecomposition and statistics theory. By decomposing of the total generalization error attributed to different eigenspace of the kernel in function space, we propose a criterion for balancing the errors stemming from training set and the distribution discrepancy between the training set and the true data. Through this approach, we establish that modifying the training direction of the neural network in function space leads to a reduction in the total generalization error. Furthermore, We demonstrate that this theoretical framework is capable to explain many existing results of generalization enhancing methods. These theoretical results are also illustrated by numerical examples on synthetic data.

📄 PDF Abstract BibTeX arXiv:2307.13290

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Natural Gradient Descent 설명 없음

Similar Papers 제목 키워드 기반

Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning

2025-07-22 · Helena Casademunt, Caden Juang, Adam Karvonen, Samuel Marks 외 arxiv

Fine-tuning large language models (LLMs) can lead to unintended out-of-distribution generalization. Standard approaches to this problem rely on modifying training data, for example by adding data that better specify the …

Gradient Descent for Low-Rank Functions

2022-06-16 · Romain Cosson, Ali Jadbabaie, Anuran Makur, Amirhossein Reisizadeh 외

Several recent empirical studies demonstrate that important machine learning tasks, e.g., training deep neural networks, exhibit low-rank structure, where the loss function varies significantly in only a few directions o…

Kill it with FIRE: On Leveraging Latent Space Directions for Runtime Backdoor Mitigation in Deep Neural Networks

2026-02-11 · Enrico Ahlers, Daniel Passon, Yannic Noller, Lars Grunske arxiv

Machine learning models are increasingly present in our everyday lives; as a result, they become targets of adversarial attackers seeking to manipulate the systems we interact with. A well-known vulnerability is a backdo…

Words in Motion: Extracting Interpretable Control Vectors for Motion Transformers

2024-06-17 · Omer Sahin Tas, Royden Wagner

Transformer-based models generate hidden states that are difficult to interpret. In this work, we analyze hidden states and modify them at inference, with a focus on motion forecasting. We use linear probing to analyze w…

Motion ForecastingZero-shot Generalization

When Shared Knowledge Hurts: Spectral Over-Accumulation in Model Merging

2026-02-05 · Yayuan Li, Ze Peng, Jian Zhang, Jintao Guo 외 arxiv

Model merging combines multiple fine-tuned models into a single model by adding their weight updates, providing a lightweight alternative to retraining. Existing methods primarily target resolving conflicts between task …