MAL: Cluster-Masked and Multi-Task Pretraining for Enhanced xLSTM Vision Performance
The Long Short-Term Memory (LSTM) networks have traditionally faced challenges in scaling and effectively capturing complex dependencies in visual tasks. The xLSTM architecture has emerged to address these limitations, incorporating exponential gating and a parallel matrix memory structure to enhance performance and scalability. Despite these advancements, the potential of xLSTM in visual computing has not been fully realized, particularly in leveraging autoregressive techniques for improved feature extraction. In this paper, we introduce MAL (Cluster-Masked and Multi-Task Pretraining for Enhanced xLSTM Vision Performance), a novel framework that enhances xLSTM's capabilities through innovative pretraining strategies. We propose a cluster-masked masking method that significantly improves local feature capture and optimizes image scanning efficiency. Additionally, our universal encoder-decoder pretraining approach integrates multiple tasks, including image autoregression, depth estimation, and image segmentation, thereby enhancing the model's adaptability and robustness across diverse visual tasks. Our experimental results demonstrate that MAL surpasses traditional supervised models and fully leverages the scaling potential of xLSTM, setting a new benchmark in visual task performance.
Code (0)
등록된 구현이 없습니다.
Tasks
DecoderDepth EstimationImage SegmentationSemantic SegmentationSimilar Papers 제목 키워드 기반
BPDec: Unveiling the Potential of Masked Language Modeling Decoder in BERT pretraining
BERT (Bidirectional Encoder Representations from Transformers) has revolutionized the field of natural language processing through its exceptional performance on numerous tasks. Yet, the majority of researchers have main…
DecoderLanguage ModelingLanguage ModellingMasked Language Modeling+1Look Ahead or Look Around? A Theoretical Comparison Between Autoregressive and Masked Pretraining
In recent years, the rise of generative self-supervised learning (SSL) paradigms has exhibited impressive performance across visual, language, and multi-modal domains. While the varied designs of generative SSL objective…
Self-Supervised LearningSelf-distilled Masked Attention guided masked image modeling with noise Regularized Teacher (SMART) for medical image analysis
Pretraining vision transformers (ViT) with attention guided masked image modeling (MIM) has shown to increase downstream accuracy for natural image analysis. Hierarchical shifted window (Swin) transformer, often used in …
Computed Tomography (CT)Medical Image AnalysisCluster-Wise Spatio-Temporal Masking for Efficient Video-Language Pretraining
Large-scale video-language pretraining enables strong generalization across multimodal tasks but often incurs prohibitive computational costs. Although recent advances in masked visual modeling help mitigate this issue, …
Video Question AnsweringVideo-Text RetrievalVideo CaptioningThe Hidden Uniform Cluster Prior in Self-Supervised Learning
A successful paradigm in representation learning is to perform self-supervised pretraining using tasks based on mini-batch statistics (e.g., SimCLR, VICReg, SwAV, MSN). We show that in the formulation of all these method…
ClusteringRepresentation LearningSelf-Supervised Learning