Exploring Multilingual Syntactic Sentence Representations
We study methods for learning sentence embeddings with syntactic structure. We focus on methods of learning syntactic sentence-embeddings by using a multilingual parallel-corpus augmented by Universal Parts-of-Speech tags. We evaluate the quality of the learned embeddings by examining sentence-level nearest neighbours and functional dissimilarity in the embedding space. We also evaluate the ability of the method to learn syntactic sentence-embeddings for low-resource languages and demonstrate strong evidence for transfer learning. Our results show that syntactic sentence-embeddings can be learned while using less training data, fewer model parameters, and resulting in better evaluation metrics than state-of-the-art language models.
Code (1)
Tasks
SentenceSentence EmbeddingsTransfer LearningSimilar Papers 제목 키워드 기반
Exploring syntactic information in sentence embeddings through multilingual subject-verb agreement
In this paper, our goal is to investigate to what degree multilingual pretrained language models capture cross-linguistically valid abstract linguistic representations. We take the approach of developing curated syntheti…
Multiple-choiceSentenceSentence EmbeddingsvalidVyākarana: A Colorless Green Benchmark for Syntactic Evaluation in Indic Languages
While there has been significant progress towards developing NLU resources for Indic languages, syntactic evaluation has been relatively less explored. Unlike English, Indic languages have rich morphosyntax, grammatical …
Depth EstimationDepth PredictionPOSPOS Tagging+2What does it mean to be language-agnostic? Probing multilingual sentence encoders for typological properties
Multilingual sentence encoders have seen much success in cross-lingual model transfer for downstream NLP tasks. Yet, we know relatively little about the properties of individual languages or the general patterns of lingu…
SentenceXLM-RExploring Semantic Properties of Sentence Embeddings
Neural vector representations are ubiquitous throughout all subfields of NLP. While word vectors have been studied in much detail, thus far only little light has been shed on the properties of sentence embeddings. In thi…
Machine TranslationReading ComprehensionSemantic Textual SimilaritySentence+3Exploring Anisotropy and Outliers in Multilingual Language Models for Cross-Lingual Semantic Sentence Similarity
Previous work has shown that the representations output by contextual language models are more anisotropic than static type embeddings, and typically display outlier dimensions. This seems to be true for both monolingual…
Semantic SimilaritySemantic Textual SimilaritySentenceSentence Similarity