Exploring syntactic information in sentence embeddings through multilingual subject-verb agreement
In this paper, our goal is to investigate to what degree multilingual pretrained language models capture cross-linguistically valid abstract linguistic representations. We take the approach of developing curated synthetic data on a large scale, with specific properties, and using them to study sentence representations built using pretrained language models. We use a new multiple-choice task and datasets, Blackbird Language Matrices (BLMs), to focus on a specific grammatical structural phenomenon -- subject-verb agreement across a variety of sentence structures -- in several languages. Finding a solution to this task requires a system detecting complex linguistic patterns and paradigms in text representations. Using a two-level architecture that solves the problem in two steps -- detect syntactic objects and their properties in individual sentences, and find patterns across an input sequence of sentences -- we show that despite having been trained on multilingual texts in a consistent manner, multilingual pretrained language models have language-specific differences, and syntactic structure is not shared, even across closely related languages.
Code (0)
등록된 구현이 없습니다.
Tasks
Multiple-choiceSentenceSentence EmbeddingsvalidMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Exploring Italian sentence embeddings properties through multi-tasking
We investigate to what degree existing LLMs encode abstract linguistic information in Italian in a multi-task setting. We exploit curated synthetic data on a large scale -- several Blackbird Language Matrices (BLMs) prob…
SentenceSentence EmbeddingsExploring Multilingual Syntactic Sentence Representations
We study methods for learning sentence embeddings with syntactic structure. We focus on methods of learning syntactic sentence-embeddings by using a multilingual parallel-corpus augmented by Universal Parts-of-Speech tag…
SentenceSentence EmbeddingsTransfer LearningExploring Semantic Properties of Sentence Embeddings
Neural vector representations are ubiquitous throughout all subfields of NLP. While word vectors have been studied in much detail, thus far only little light has been shed on the properties of sentence embeddings. In thi…
Machine TranslationReading ComprehensionSemantic Textual SimilaritySentence+3Disentangling Semantics and Syntax in Sentence Embeddings with Pre-trained Language Models
Pre-trained language models have achieved huge success on a wide range of NLP tasks. However, contextual representations from pre-trained models contain entangled semantic and syntactic information, and therefore cannot …
Semantic SimilaritySemantic Textual SimilaritySentenceSentence Embedding+2Syntactic representation learning for neural network based TTS with syntactic parse tree traversal
Syntactic structure of a sentence text is correlated with the prosodic structure of the speech that is crucial for improving the prosody and naturalness of a text-to-speech (TTS) system. Nowadays TTS systems usually try …
DiversityRepresentation LearningSentencetext-to-speech+1