paper-with-me

홈 › Papers

Overcoming Rank Collapse in Feedback Alignment

2026-06-09 · Gauthier Boeshertz, Razvan Pascanu, Claudia Clopath arxiv

Backpropagation (BP) is widely viewed as biologically implausible, in part because it requires feedback weights to be the transpose of forward weights for error propagation. Interestingly, when training a network with fixed random feedback weights to circumvent this issue, learning aligns the forward weights with the feedback weights, leading the backpropagated error signal to become an approximation of the standard gradient used by BP. This process, called Feedback Alignment (FA), occurs in MLPs and very shallow CNNs but does not scale well to deeper architectures. In this work, we first investigated differences between BP and FA models, trained on CIFAR10, specifically focusing on the effective rank of the signal. We found that the FA error has a considerably lower rank and hence is constrained to a lower-dimensional subspace compared to BP, limiting exploration of the parameter space. Motivated by this observation, we evaluated two mechanisms for increasing the effective dimensionality of FA: Muon, an optimiser that orthogonalises weight updates; and hidden activity normalisation, which promotes activation orthogonality. Across larger architectures and benchmarks, we find that these methods consistently improve over FA baselines, for example, on CIFAR100 with a Resnet-18, accuracy increases by 9 percentage points. Our results identify low-dimensional gradient dynamics as a key obstacle to scaling FA and suggest that inducing higher-dimensional update geometry is a promising route toward scaling alternatives to backpropagation.

📄 PDF Abstract BibTeX arXiv:2606.11123

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Pistis-RAG: Enhancing Retrieval-Augmented Generation with Human Feedback

2024-06-21 · Yu Bai, Yukai Miao, Li Chen, Dawei Wang 외

RAG systems face limitations when semantic relevance alone does not guarantee improved generation quality. This issue becomes particularly evident due to the sensitivity of large language models (LLMs) to the ordering of…

Information RetrievalLearning-To-RankMMLURAG+2

Explaining and Preventing Alignment Collapse in Iterative RLHF

2026-05-05 · Etienne Gauthier, Francis Bach, Michael I. Jordan arxiv

Reinforcement learning from human feedback (RLHF) typically assumes a static or non-strategic reward model (RM). In iterative deployment, however, the policy generates the data on which the RM is retrained, creating a fe…

Reinforcement Learning

PRISMR: Overcoming Parse Collapse in Multimodal Listwise Ranking via Parameterized Representation Internalization

2026-06-11 · Hao Jiang, Xin Li, Annan Wang, Zhi Yang 외 arxiv

Generative listwise ranking with Large Multimodal Models (LMMs) aims to capture global list context in a single forward pass, but its effectiveness degrades in long-context multimodal scenarios. We identify a recurring f…

Prompt Engineering

Overcoming Prior Misspecification in Online Learning to Rank

2023-01-25 · Javad Azizi, Ofer Meshi, Masrour Zoghi, Maryam Karimzadehgan

The recent literature on online learning to rank (LTR) has established the utility of prior knowledge to Bayesian ranking bandit algorithms. However, a major limitation of existing work is the requirement for the prior u…

Learning-To-Rank

What Accuracy and Gradient Cosine Miss: Evaluating Feedback Alignment via Scale Stability, Reference Validity, and Depth Utility

2026-06-19 · Yuren Hao, Xiang Wan, ChengXiang Zhai arxiv

Despite the success of deep learning, training deep networks in biologically plausible and hardware-efficient ways remains an open challenge. Feedback alignment (FA) methods address this by replacing backpropagation's sy…