paper-with-me

홈 › Papers

Normalized Attention Without Probability Cage

2020-05-19 · Oliver Richter, Roger Wattenhofer

Attention architectures are widely used; they recently gained renewed popularity with Transformers yielding a streak of state of the art results. Yet, the geometrical implications of softmax-attention remain largely unexplored. In this work we highlight the limitations of constraining attention weights to the probability simplex and the resulting convex hull of value vectors. We show that Transformers are sequence length dependent biased towards token isolation at initialization and contrast Transformers to simple max- and sum-pooling - two strong baselines rarely reported. We propose to replace the softmax in self-attention with normalization, yielding a hyperparameter and data-bias robust, generally applicable architecture. We support our insights with empirical results from more than 25,000 trained models. All results and implementations are made available.

📄 PDF Abstract BibTeX arXiv:2005.09561

Code (2)

OliverRichter/normalized-attention 공식 구현 tf
lucidrains/all-normalization-transformer pytorch

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…

Similar Papers 제목 키워드 기반

UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation

2025-08-07 · Wonjun Kang, Byeongkeun Ahn, Minjae Lee, Kevin Galim 외 arxiv

Text-to-image (T2I) generation has been actively studied using Diffusion Models and Autoregressive Models. Recently, Masked Generative Transformers have gained attention as an alternative to Autoregressive Models to over…

Text-to-Image Generation

Neural Cages for Detail-Preserving 3D Deformations

2019-12-13 · CVPR 2020 6 · Wang Yifan, Noam Aigerman, Vladimir G. Kim, Siddhartha Chaudhuri 외

We propose a novel learnable representation for detail-preserving shape deformation. The goal of our method is to warp a source shape to match the general structure of a target shape, while preserving the surface details…

UnCageNet: Tracking and Pose Estimation of Caged Animal

2025-12-08 · Sayak Dutta, Harish Katti, Shashikant Verma, Shanmuganathan Raman arxiv

Animal tracking and pose estimation systems, such as STEP (Simultaneous Tracking and Pose Estimation) and ViTPose, experience substantial performance drops when processing images and videos with cage structures and syste…

Keypoint DetectionPose Estimation

Zero Shot Deformation Reconstruction for Soft Robots Using a Flexible Sensor Array and Cage Based 3D Gaussian Modeling

2026-03-20 · Linrui Shou, Zilang Chen, Wenjia Xu, Yiyue Luo 외 arxiv

We present a zero-shot deformation reconstruction framework for soft robots that operates without any visual supervision at inference time. In this work, zero-shot deformation reconstruction is defined as the ability to …

Zero-shot Generalization

CAGE-GS: High-fidelity Cage Based 3D Gaussian Splatting Deformation

2025-04-17 · Yifei Tong, Runze Tian, Xiao Han, Dingyao Liu 외

As 3D Gaussian Splatting (3DGS) gains popularity as a 3D representation of real scenes, enabling user-friendly deformation to create novel scenes while preserving fine details from the original 3DGS has attracted signifi…

3DGS