Learning Universal Sentence Embeddings with Large-scale Parallel Translation Datasets
Although contrastive learning has greatly improved sentence representation, its performance is still limited by the size of monolingual sentence-pair datasets. Meanwhile, there exist large-scale parallel translation pairs (100x larger than monolingual pairs) that are highly correlated in semantic, but have not been utilized for learning universal sentence representation. Furthermore, given parallel translation pairs, previous contrastive learning frameworks can not well balance the monolingual embeddings’ alignment and uniformity which represent the quality of embeddings. In this paper, we build on the top of dual encoder and propose to freeze the source language encoder, utilizing its consistent embeddings to supervise the target language encoder via contrastive learning, where source-target translation pairs are regarded as positives. We provide the first exploration of utilizing parallel translation sentence pairs to learn universal sentence embeddings and show superior performance to balance the alignment and uniformity. We achieve a new state-of-the-art performance on the average score of standard semantic textual similarity (STS), outperforming both SimCSE and Sentence-T5, and the best performance in corresponding tracks on transfer tasks.
Code (0)
등록된 구현이 없습니다.
Tasks
Contrastive LearningSemantic Textual SimilaritySentenceSentence EmbeddingsSTSTranslationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
English Contrastive Learning Can Learn Universal Cross-lingual Sentence Embeddings
Universal cross-lingual sentence embeddings map semantically similar cross-lingual sentences into a shared embedding space. Aligning cross-lingual sentence embeddings usually requires supervised cross-lingual parallel se…
Contrastive LearningRetrievalSentenceSentence Embedding+3Learning Monolingual Sentence Embeddings with Large-scale Parallel Translation Datasets
Although contrastive learning has greatly improved sentence representation, its performance is still limited by the size of monolingual sentence-pair datasets. Meanwhile, there exist large-scale parallel translation pair…
Contrastive LearningSemantic Textual SimilaritySentenceSentence Embeddings+2A simple method for domain adaptation of sentence embeddings
Pre-trained sentence embeddings have been shown to be very useful for a variety of NLP tasks. Due to the fact that training such embeddings requires a large amount of data, they are commonly trained on a variety of text …
Domain AdaptationSentenceSentence EmbeddingsModel-Based Quality Assessment for Massively Multilingual Parallel Data
Large-scale multilingual bitext often contains two distinct problems: non-parallel sentence pairs and low-quality translations. We decompose model-based assessment for such data into two independent components: paralleli…
Exploring Multilingual Syntactic Sentence Representations
We study methods for learning sentence embeddings with syntactic structure. We focus on methods of learning syntactic sentence-embeddings by using a multilingual parallel-corpus augmented by Universal Parts-of-Speech tag…
SentenceSentence EmbeddingsTransfer Learning