paper-with-me

홈 › Papers

Fast, Effective, and Self-Supervised: Transforming Masked Language Models into Universal Lexical and Sentence Encoders

2021-04-16 · EMNLP 2021 11 · Fangyu Liu, Ivan Vulić, Anna Korhonen, Nigel Collier

Pretrained Masked Language Models (MLMs) have revolutionised NLP in recent years. However, previous work has indicated that off-the-shelf MLMs are not effective as universal lexical or sentence encoders without further task-specific fine-tuning on NLI, sentence similarity, or paraphrasing tasks using annotated task data. In this work, we demonstrate that it is possible to turn MLMs into effective universal lexical and sentence encoders even without any additional data and without any supervision. We propose an extremely simple, fast and effective contrastive learning technique, termed Mirror-BERT, which converts MLMs (e.g., BERT and RoBERTa) into such encoders in 20-30 seconds without any additional external knowledge. Mirror-BERT relies on fully identical or slightly modified string pairs as positive (i.e., synonymous) fine-tuning examples, and aims to maximise their similarity during identity fine-tuning. We report huge gains over off-the-shelf MLMs with Mirror-BERT in both lexical-level and sentence-level tasks, across different domains and different languages. Notably, in the standard sentence semantic similarity (STS) tasks, our self-supervised Mirror-BERT model even matches the performance of the task-tuned Sentence-BERT models from prior work. Finally, we delve deeper into the inner workings of MLMs, and suggest some evidence on why this simple approach can yield effective universal lexical and sentence encoders.

📄 PDF Abstract BibTeX arXiv:2104.08027

Code (1)

cambridgeltl/mirror-bert 공식 구현 pytorch

Tasks

Contrastive LearningCross-Lingual Semantic Textual SimilarityEntity LinkingSemantic SimilaritySemantic Textual SimilaritySentenceSentence SimilaritySTS

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Contrastive Learning 설명 없음
Mirror-BERT Mirror-BERT converts pretrained language models into effective universal text encoders without any supervision, in 20-30 seconds. It is an extremely simple, fast, and effective…
Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Adam 설명 없음

Similar Papers 제목 키워드 기반

Exploring The Role of Mean Teachers in Self-supervised Masked Auto-Encoders

2022-10-05 · Youngwan Lee, Jeffrey Willette, Jonghee Kim, Juho Lee 외

Masked image modeling (MIM) has become a popular strategy for self-supervised learning~(SSL) of visual representations with Vision Transformers. A representative MIM model, the masked auto-encoder (MAE), randomly masks a…

ClassificationInstance Segmentationobject-detectionObject Detection+2

Time Series Generation with Masked Autoencoder

2022-01-14 · Mengyue Zha, SiuTim Wong, Mengqi Liu, Tong Zhang 외

This paper shows that masked autoencoder with extrapolator (ExtraMAE) is a scalable self-supervised model for time series generation. ExtraMAE randomly masks some patches of the original time series and learns temporal d…

Data AugmentationDecoderImputationManagement+5

Efficient Self-supervised Vision Pretraining with Local Masked Reconstruction

2022-06-01 · Jun Chen, Ming Hu, Boyang Li, Mohamed Elhoseiny

Self-supervised learning for computer vision has achieved tremendous progress and improved many downstream vision tasks such as image classification, semantic segmentation, and object detection. Among these, generative s…

image-classificationImage ClassificationInstance SegmentationObject Detection+2

FILS: Self-Supervised Video Feature Prediction In Semantic Language Space

2024-06-05 · Mona Ahmadian, Frank Guerin, Andrew Gilbert

This paper demonstrates a self-supervised approach for learning semantic video representations. Recent vision studies show that a masking strategy for vision and natural language supervision has contributed to developing…

Action RecognitionDecoder

Structure is Supervision: Multiview Masked Autoencoders for Radiology

2025-11-27 · Sonia Laguna, Andrea Agostini, Alain Ryser, Samuel Ruiperez-Campillo 외 arxiv

Building robust medical machine learning systems requires pretraining strategies that exploit the intrinsic structure present in clinical data. We introduce Multiview Masked Autoencoder (MVMAE), a self-supervised framewo…

Image Reconstruction