paper-with-me

홈 › Papers

Understanding Layer-wise Contributions in Deep Neural Networks through Spectral Analysis

2021-11-06 · Yatin Dandi, Arthur Jacot

Spectral analysis is a powerful tool, decomposing any function into simpler parts. In machine learning, Mercer's theorem generalizes this idea, providing for any kernel and input distribution a natural basis of functions of increasing frequency. More recently, several works have extended this analysis to deep neural networks through the framework of Neural Tangent Kernel. In this work, we analyze the layer-wise spectral bias of Deep Neural Networks and relate it to the contributions of different layers in the reduction of generalization error for a given target function. We utilize the properties of Hermite polynomials and Spherical Harmonics to prove that initial layers exhibit a larger bias towards high-frequency functions defined on the unit sphere. We further provide empirical results validating our theory in high dimensional datasets for Deep Neural Networks.

📄 PDF Abstract BibTeX arXiv:2111.03972

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SpectralKD: A Unified Framework for Interpreting and Distilling Vision Transformers via Spectral Analysis

2024-12-26 · Huiyuan Tian, Bonan Xu, Shijian Li, Gang Pan

Knowledge Distillation (KD) has achieved widespread success in compressing large Vision Transformers (ViTs), but a unified theoretical framework for both ViTs and KD is still lacking. In this paper, we propose SpectralKD…

Knowledge DistillationTransfer Learning

Deep Spectral Convolution Network for HyperSpectral Unmixing

2018-06-22 · Savas Ozkan, Gozde Bozdagi Akar

In this paper, we propose a novel hyperspectral unmixing technique based on deep spectral convolution networks (DSCN). Particularly, three important contributions are presented throughout this paper. First, fully-connect…

Hyperspectral Unmixing

Module-wise Adaptive Distillation for Multimodality Foundation Models

2023-10-06 · NeurIPS 2023 11

Pre-trained multimodal foundation models have demonstrated remarkable generalizability but pose challenges for deployment due to their large sizes. One effective approach to reducing their sizes is layerwise distillation…

Image CaptioningThompson Sampling

Mathematical Foundations of Neural Tangents and Infinite-Width Networks

2025-12-09 · Rachana Mysore, Preksha Girish, Kavitha Jayaram, Shrey Kumar 외 arxiv

We investigate the mathematical foundations of neural networks in the infinite-width regime through the Neural Tangent Kernel (NTK). We propose the NTK-Eigenvalue-Controlled Residual Network (NTK-ECRN), an architecture i…

Comprehensive Layer-wise Analysis of SSL Models for Audio Deepfake Detection

2025-02-05 · Yassine El Kheir, Youness Samih, Suraj Maharjan, Tim Polzehl 외

This paper conducts a comprehensive layer-wise analysis of self-supervised learning (SSL) models for audio deepfake detection across diverse contexts, including multilingual datasets (English, Chinese, Spanish), partial,…

Audio Deepfake DetectionDeepFake DetectionFace SwappingSelf-Supervised Learning