paper-with-me

Papers

Matrix Information Theory for Self-Supervised Learning

2023-05-27 · Yifan Zhang, Zhiquan Tan, Jingqin Yang, Weiran Huang, Yang Yuan

The maximum entropy encoding framework provides a unified perspective for many non-contrastive learning methods like SimSiam, Barlow Twins, and MEC. Inspired by this framework, we introduce Matrix-SSL, a novel approach that leverages matrix information theory to interpret the maximum entropy encoding loss as matrix uniformity loss. Furthermore, Matrix-SSL enhances the maximum entropy encoding method by seamlessly incorporating matrix alignment loss, directly aligning covariance matrices in different branches. Experimental results reveal that Matrix-SSL outperforms state-of-the-art methods on the ImageNet dataset under linear evaluation settings and on MS-COCO for transfer learning tasks. Specifically, when performing transfer learning tasks on MS-COCO, our method outperforms previous SOTA methods such as MoCo v2 and BYOL up to 3.3% with only 400 epochs compared to 800 epochs pre-training. We also try to introduce representation learning into the language modeling regime by fine-tuning a 7B model using matrix cross-entropy loss, with a margin of 3.1% on the GSM8K dataset over the standard cross-entropy loss. Code available at https://github.com/yifanzhang-pro/Matrix-SSL.

📄 PDF Abstract BibTeX arXiv:2305.17326

Code (3)

yifanzhang-pro/matrix-llm 공식 구현
yifanzhang-pro/matrix-ssl 공식 구현 pytorch
huang-research-group/Matrix-SSL pytorch

Tasks

Contrastive LearningGSM8KLanguage ModelingLanguage ModellingLinear evaluationRepresentation LearningSelf-Supervised LearningTransfer Learning

Methods 이 논문이 사용한 방법론

Bitcoin Customer Service Number +1-833-534-1729 설명 없음
ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
MoCo v2 MoCo v2 is an improved version of the Momentum Contrast self-supervised learning algorithm. Motivated by the findings presented in…
InfoNCE 설명 없음
MoCo 설명 없음
1x1 Convolution A 1 x 1 Convolution is a convolution with some special properties in that it can be used for dimensionality reduction,…
Average Pooling 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…

Similar Papers 제목 키워드 기반

Unveiling the Dynamics of Information Interplay in Supervised Learning

2024-06-06 · Kun Song, Zhiquan Tan, Bochao Zou, Huimin Ma 외

In this paper, we use matrix information theory as an analytical tool to analyze the dynamics of the information interplay between data representations and classification head vectors in the supervised learning process. …

Linear Mode Connectivity

Self-information Domain-based Neural CSI Compression with Feature Coupling

2023-04-30 · Ziqing Yin, Renjie Xie, Wei Xu, Zhaohui Yang 외

Deep learning (DL)-based channel state information (CSI) feedback methods compressed the CSI matrix by exploiting its delay and angle features straightforwardly, while the measure in terms of information contained in the…

To Compress or Not to Compress- Self-Supervised Learning and Information Theory: A Review

2023-04-19 · Ravid Shwartz-Ziv, Yann Lecun

Deep neural networks excel in supervised learning tasks but are constrained by the need for extensive labeled data. Self-supervised learning emerges as a promising alternative, allowing models to learn without explicit l…

Self-Supervised Learning

Matrix Completion with Noisy Side Information

2015-12-01 · NeurIPS 2015 12 · Kai-Yang Chiang, Cho-Jui Hsieh, Inderjit S. Dhillon

We study matrix completion problem with side information. Side information has been considered in several matrix completion applications, and is generally shown to be useful empirically. Recently, Xu et al. studied the…

ClusteringMatrix Completion

Exploring Information-Theoretic Metrics Associated with Neural Collapse in Supervised Training

2024-09-25 · Kun Song, Zhiquan Tan, Bochao Zou, Jiansheng Chen 외

In this paper, we utilize information-theoretic metrics like matrix entropy and mutual information to analyze supervised learning. We explore the information content of data representations and classification head weight…

Classificationcross-modal alignment