paper-with-me

Papers

VeLU: Variance-enhanced Learning Unit for Deep Neural Networks

2025-04-21 · Ashkan Shakarami, Yousef Yeganeh, Azade Farshad, Lorenzo Nicolè, Stefano Ghidoni, Nassir Navab

Activation functions are fundamental in deep neural networks and directly impact gradient flow, optimization stability, and generalization. Although ReLU remains standard because of its simplicity, it suffers from vanishing gradients and lacks adaptability. Alternatives like Swish and GELU introduce smooth transitions, but fail to dynamically adjust to input statistics. We propose VeLU, a Variance-enhanced Learning Unit as an activation function that dynamically scales based on input variance by integrating ArcTan-Sin transformations and Wasserstein-2 regularization, effectively mitigating covariate shifts and stabilizing optimization. Extensive experiments on ViT_B16, VGG19, ResNet50, DenseNet121, MobileNetV2, and EfficientNetB3 confirm VeLU's superiority over ReLU, ReLU6, Swish, and GELU on six vision benchmarks. The codes of VeLU are publicly available on GitHub.

📄 PDF Abstract BibTeX arXiv:2504.15051

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Pointwise Convolution Pointwise Convolution is a type of convolution that uses a 1x1 kernel: a kernel that iterates through every single point. This…
Depthwise Convolution Depthwise Convolution is a type of convolution where we apply a single convolutional filter for each input channel. In the regular 2D…
Depthwise Separable Convolution While standard convolution performs the channelwise and spatial-wise computation in one step, Depthwise Separable Convolution
Batch Normalization 설명 없음
Average Pooling 설명 없음
Inverted Residual Block 설명 없음
Sigmoid Activation 설명 없음

Similar Papers 제목 키워드 기반

Enhancing Acoustic-to-Articulatory Speech Inversion by Incorporating Nasality

2025-06-10 · Saba Tabatabaee, Suzanne Boyce, Liran Oren, Mark Tiede 외

Speech is produced through the coordination of vocal tract constricting organs: lips, tongue, velum, and glottis. Previous works developed Speech Inversion (SI) systems to recover acoustic-to-articulatory mappings for li…

Design and Optimization of Reinforcement Learning-Based Agents in Text-Based Games

2025-09-03 · Haonan Wang, Mingjia Zhao, Junfeng Sun, Wei Liu arxiv

As AI technology advances, research in playing text-based games with agents has becomeprogressively popular. In this paper, a novel approach to agent design and agent learning ispresented with the context of reinforcemen…

Reinforcement Learning

An error correction scheme for improved air-tissue boundary in real-time MRI video for speech production

2022-03-09 · Anwesha Roy, Varun Belagali, Prasanta Kumar Ghosh

The best performance in Air-tissue boundary (ATB) segmentation of real-time Magnetic Resonance Imaging (rtMRI) videos in speech production is known to be achieved by a 3-dimensional convolutional neural network (3D-CNN) …

Dynamic Time WarpingSegmentation

Neural Network-Driven Volatility Drag Mitigation under Aggressive Leverage

2026-07-25 · Christian Bongiorno, Efstratios Manolakis, Rosario Nunzio Mantegna arxiv

This paper introduces a compact reformulation of a modular end-to-end neural network for global minimum-variance portfolio optimization that decouples model complexity from both look-back window length and universe size.…

Portfolio Optimization

Speaker-independent Speech Inversion for Estimation of Nasalance

2023-05-31 · Yashish M. Siriwardena, Carol Espy-Wilson, Suzanne Boyce, Mark K. Tiede 외

The velopharyngeal (VP) valve regulates the opening between the nasal and oral cavities. This valve opens and closes through a coordinated motion of the velum and pharyngeal walls. Nasalance is an objective measure deriv…