paper-with-me

홈 › Papers

MAL: Cluster-Masked and Multi-Task Pretraining for Enhanced xLSTM Vision Performance

2024-12-14 · Wenjun Huang, Jianguo Hu

The Long Short-Term Memory (LSTM) networks have traditionally faced challenges in scaling and effectively capturing complex dependencies in visual tasks. The xLSTM architecture has emerged to address these limitations, incorporating exponential gating and a parallel matrix memory structure to enhance performance and scalability. Despite these advancements, the potential of xLSTM in visual computing has not been fully realized, particularly in leveraging autoregressive techniques for improved feature extraction. In this paper, we introduce MAL (Cluster-Masked and Multi-Task Pretraining for Enhanced xLSTM Vision Performance), a novel framework that enhances xLSTM's capabilities through innovative pretraining strategies. We propose a cluster-masked masking method that significantly improves local feature capture and optimizes image scanning efficiency. Additionally, our universal encoder-decoder pretraining approach integrates multiple tasks, including image autoregression, depth estimation, and image segmentation, thereby enhancing the model's adaptability and robustness across diverse visual tasks. Our experimental results demonstrate that MAL surpasses traditional supervised models and fully leverages the scaling potential of xLSTM, setting a new benchmark in visual task performance.

📄 PDF Abstract BibTeX arXiv:2412.10730

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderDepth EstimationImage SegmentationSemantic Segmentation

Similar Papers 제목 키워드 기반

BPDec: Unveiling the Potential of Masked Language Modeling Decoder in BERT pretraining

2024-01-29 · Wen Liang, Youzhi Liang

BERT (Bidirectional Encoder Representations from Transformers) has revolutionized the field of natural language processing through its exceptional performance on numerous tasks. Yet, the majority of researchers have main…

DecoderLanguage ModelingLanguage ModellingMasked Language Modeling+1

Look Ahead or Look Around? A Theoretical Comparison Between Autoregressive and Masked Pretraining

2024-07-01 · Qi Zhang, Tianqi Du, Haotian Huang, Yifei Wang 외

In recent years, the rise of generative self-supervised learning (SSL) paradigms has exhibited impressive performance across visual, language, and multi-modal domains. While the varied designs of generative SSL objective…

Self-Supervised Learning

Self-distilled Masked Attention guided masked image modeling with noise Regularized Teacher (SMART) for medical image analysis

2023-10-02 · Jue Jiang, Aneesh Rangnekar, Chloe Min Seo Choi, Harini Veeraraghavan

Pretraining vision transformers (ViT) with attention guided masked image modeling (MIM) has shown to increase downstream accuracy for natural image analysis. Hierarchical shifted window (Swin) transformer, often used in …

Computed Tomography (CT)Medical Image Analysis

Cluster-Wise Spatio-Temporal Masking for Efficient Video-Language Pretraining

2026-03-24 · Weijun Zhuang, Yuqing Huang, Weikang Meng, Xin Li 외 arxiv

Large-scale video-language pretraining enables strong generalization across multimodal tasks but often incurs prohibitive computational costs. Although recent advances in masked visual modeling help mitigate this issue, …

Video Question AnsweringVideo-Text RetrievalVideo Captioning

The Hidden Uniform Cluster Prior in Self-Supervised Learning

2022-10-13 · Mahmoud Assran, Randall Balestriero, Quentin Duval, Florian Bordes 외

A successful paradigm in representation learning is to perform self-supervised pretraining using tasks based on mini-batch statistics (e.g., SimCLR, VICReg, SwAV, MSN). We show that in the formulation of all these method…

ClusteringRepresentation LearningSelf-Supervised Learning