paper-with-me

Papers

Hierarchical-Task-Aware Multi-modal Mixture of Incremental LoRA Experts for Embodied Continual Learning

2025-06-05 · Ziqi Jia, Anmin Wang, Xiaoyang Qu, Xiaowen Yang, Jianzong Wang

Previous continual learning setups for embodied intelligence focused on executing low-level actions based on human commands, neglecting the ability to learn high-level planning and multi-level knowledge. To address these issues, we propose the Hierarchical Embodied Continual Learning Setups (HEC) that divide the agent's continual learning process into two layers: high-level instructions and low-level actions, and define five embodied continual learning sub-setups. Building on these setups, we introduce the Task-aware Mixture of Incremental LoRA Experts (Task-aware MoILE) method. This approach achieves task recognition by clustering visual-text embeddings and uses both a task-level router and a token-level router to select the appropriate LoRA experts. To effectively address the issue of catastrophic forgetting, we apply Singular Value Decomposition (SVD) to the LoRA parameters obtained from prior tasks, preserving key components while orthogonally training the remaining parts. The experimental results show that our method stands out in reducing the forgetting of old tasks compared to other methods, effectively supporting agents in retaining prior knowledge while continuously learning new tasks.

📄 PDF Abstract BibTeX arXiv:2506.04595

Code (0)

등록된 구현이 없습니다.

Tasks

Continual Learning

Similar Papers 제목 키워드 기반

Hierarchical Time-Aware Mixture of Experts for Multi-Modal Sequential Recommendation

2025-01-24 · Shengzhe Zhang, Liyi Chen, Dazhong Shen, Chao Wang 외

Multi-modal sequential recommendation (SR) leverages multi-modal data to learn more comprehensive item features and user preferences than traditional SR methods, which has become a critical topic in both academia and ind…

Contrastive LearningMixture-of-ExpertsMulti-Task LearningSequential Recommendation

EEG-Based Multimodal Learning via Hyperbolic Mixture-of-Curvature Experts

2026-04-14 · Runhe Zhou, Shanglin Li, Guanxiang Huang, Xinliang Zhou 외 arxiv

Electroencephalography (EEG)-based multimodal learning integrates brain signals with complementary modalities to improve mental state assessment, providing great clinical potential. The effectiveness of such paradigms la…

Representation LearningEmotion Recognition

M$^4$-SAM: Multi-Modal Mixture-of-Experts with Memory-Augmented SAM for RGB-D Video Salient Object Detection

2026-05-12 · Jiyuan Liu, Jia Lin, Xiaofei Zhou, Runmin Cong 외 arxiv

The Segment Anything Model 2 (SAM2) has emerged as a foundation model for universal segmentation. Owing to its generalizable visual representations, SAM2 has been successfully applied to various downstream tasks. However…

Video Salient Object Detection

Multimodal Mixture-of-Experts with Retrieval Augmentation for Protein Active Site Identification

2026-03-02 · Jiayang Wu, Jiale Zhou, Rubo Wang, Xingyi Zhang 외 arxiv

Accurate identification of protein active sites at the residue level is crucial for understanding protein function and advancing drug discovery. However, current methods face two critical challenges: vulnerability in sin…

Drug Discovery

ESG-Net: Event-Aware Semantic Guided Network for Dense Audio-Visual Event Localization

2025-07-14 · Huilai Li, Yonghao Dang, Ying Xing, Yiming Wang 외 arxiv

Dense audio-visual event localization (DAVE) aims to identify event categories and locate the temporal boundaries in untrimmed videos. Most studies only employ event-related semantic constraints on the final outputs, lac…

audio-visual event localization