paper-with-me

Papers

Fresh-CL: Feature Realignment through Experts on Hypersphere in Continual Learning

2025-01-04 · Zhongyi Zhou, Yaxin Peng, Pin Yi, Minjie Zhu, Chaomin Shen

Continual Learning enables models to learn and adapt to new tasks while retaining prior knowledge. Introducing new tasks, however, can naturally lead to feature entanglement across tasks, limiting the model's capability to distinguish between new domain data. In this work, we propose a method called Feature Realignment through Experts on hyperSpHere in Continual Learning (Fresh-CL). By leveraging predefined and fixed simplex equiangular tight frame (ETF) classifiers on a hypersphere, our model improves feature separation both intra and inter tasks. However, the projection to a simplex ETF shifts with new tasks, disrupting structured feature representation of previous tasks and degrading performance. Therefore, we propose a dynamic extension of ETF through mixture of experts, enabling adaptive projections onto diverse subspaces to enhance feature representation. Experiments on 11 datasets demonstrate a 2% improvement in accuracy compared to the strongest baseline, particularly in fine-grained datasets, confirming the efficacy of combining ETF and MoE to improve feature distinction in continual learning scenarios.

📄 PDF Abstract BibTeX arXiv:2501.02198

Code (0)

등록된 구현이 없습니다.

Tasks

Continual LearningMixture-of-Experts

Methods 이 논문이 사용한 방법론

MoE 설명 없음

Similar Papers 제목 키워드 기반

GLOVE: Global Verifier for LLM Memory-Environment Realignment

2026-01-27 · Xingkun Yin, Hongyang Du arxiv

Most existing memory-enhanced Large Language Model (LLM) approaches implicitly assume that memory validity can be established either through external evaluators that provide task-specific success signals or through inter…

The Realignment Problem: When Right becomes Wrong in LLMs

2025-11-04 · Aakash Sen Sharma, Debdeep Sanyal, Manodeep Ray, Vivek Srivastava 외 arxiv

Post-training alignment of large language models (LLMs) relies on large-scale human annotations guided by policy specifications that change over time. Cultural shifts, value reinterpretations, and regulatory or industria…

An Emergent Mirage: Is Emergent Misalignment and Realignment Indeed a Robust Phenomenon?

2026-07-10 · Abhinav Rao, Liancheng Gong, Bin Hu, Atharva Naik arxiv

Recent work has reported Emergent Misalignment (EM), where language models fine-tuned on narrow, domain-specific misaligned datasets abruptly acquire broadly misaligned behavior, alongside evidence that this behavior can…

Rethinking Language Model Scaling under Transferable Hypersphere Optimization

2026-03-30 · Liliang Ren, Yang Liu, Yelong Shen, Weizhu Chen arxiv

Scaling laws for large language models depend critically on the optimizer and parameterization. Existing hyperparameter transfer laws are mainly developed for first-order optimizers, and they do not structurally prevent …

Exploring the Relationship between Alignment and Cross-lingual Transfer in Multilingual Transformers

2023-06-05 · Félix Gaschi, Patricio Cerda, Parisa Rastin, Yannick Toussaint

Without any explicit cross-lingual training data, multilingual language models can achieve cross-lingual transfer. One common way to improve this transfer is to perform realignment steps before fine-tuning, i.e., to trai…

Cross-Lingual TransferPOSPOS TaggingXLM-R