paper-with-me

Papers

$\boldsymbolλ$-Orthogonality Regularization for Compatible Representation Learning

2025-09-20 · Simone Ricci, Niccolò Biondi, Federico Pernici, Ioannis Patras, Alberto Del Bimbo arxiv

Retrieval systems rely on representations learned by increasingly powerful models. However, due to the high training cost and inconsistencies in learned representations, there is significant interest in facilitating communication between representations and ensuring compatibility across independently trained neural networks. In the literature, two primary approaches are commonly used to adapt different learned representations: affine transformations, which adapt well to specific distributions but can significantly alter the original representation, and orthogonal transformations, which preserve the original structure with strict geometric constraints but limit adaptability. A key challenge is adapting the latent spaces of updated models to align with those of previous models on downstream distributions while preserving the newly learned representation spaces. In this paper, we impose a relaxed orthogonality constraint, namely $λ$-Orthogonality regularization, while learning an affine transformation, to obtain distribution-specific adaptation while retaining the original learned representations. Extensive experiments across various architectures and datasets validate our approach, demonstrating that it preserves the model's zero-shot performance and ensures compatibility across model updates. Code available at: \href{https://github.com/miccunifi/lambda_orthogonality.git}{https://github.com/miccunifi/lambda\_orthogonality}.

📄 PDF Abstract BibTeX arXiv:2509.16664

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning

Similar Papers 제목 키워드 기반

Max-Margin Token Selection in Attention Mechanism

2023-06-23 · NeurIPS 2023 11 · Davoud Ataee Tarzanagh, Yingcong Li, Xuechen Zhang, Samet Oymak

Attention mechanism is a central component of the transformer architecture which led to the phenomenal success of large language models. However, the theoretical principles underlying the attention mechanism are poorly u…

Towards Better Orthogonality Regularization with Disentangled Norm in Training Deep CNNs

2023-06-16 · Changhao Wu, Shenan Zhang, Fangsong Long, Ziliang Yin 외

Orthogonality regularization has been developed to prevent deep CNNs from training instability and feature redundancy. Among existing proposals, kernel orthogonality regularization enforces orthogonality by minimizing th…

Approximate Leave-One-Out for Fast Parameter Tuning in High Dimensions

2018-07-07 · ICML 2018 7 · Shuaiwen Wang, Wenda Zhou, Haihao Lu, Arian Maleki 외

Consider the following class of learning schemes: $$\hat{\boldsymbol{\beta}} := \arg\min_{\boldsymbol{\beta}}\;\sum_{j=1}^n \ell(\boldsymbol{x}_j^\top\boldsymbol{\beta}; y_j) + \lambda R(\boldsymbol{\beta}),\qquad\qquad …

Vocal Bursts Intensity Prediction

Orthogonality Constrained Multi-Head Attention For Keyword Spotting

2019-10-10 · Mingu Lee, Jinkyu Lee, Hye Jin Jang, Byeonggeun Kim 외

Multi-head attention mechanism is capable of learning various representations from sequential data while paying attention to different subsequences, e.g., word-pieces or syllables in a spoken word. From the subsequences,…

Keyword Spotting

Approximate Leave-One-Out for High-Dimensional Non-Differentiable Learning Problems

2018-10-04 · Shuaiwen Wang, Wenda Zhou, Arian Maleki, Haihao Lu 외

Consider the following class of learning schemes: \begin{equation} \label{eq:main-problem1} \hat{\boldsymbol{\beta}} := \underset{\boldsymbol{\beta} \in \mathcal{C}}{\arg\min} \;\sum_{j=1}^n \ell(\boldsymbol{x}_j^\top\…

Vocal Bursts Intensity Prediction