Exploring Italian sentence embeddings properties through multi-tasking
We investigate to what degree existing LLMs encode abstract linguistic information in Italian in a multi-task setting. We exploit curated synthetic data on a large scale -- several Blackbird Language Matrices (BLMs) problems in Italian -- and use them to study how sentence representations built using pre-trained language models encode specific syntactic and semantic information. We use a two-level architecture to model separately a compression of the sentence embeddings into a representation that contains relevant information for a task, and a BLM task. We then investigate whether we can obtain compressed sentence representations that encode syntactic and semantic information relevant to several BLM tasks. While we expected that the sentence structure -- in terms of sequence of phrases/chunks -- and chunk properties could be shared across tasks, performance and error analysis show that the clues for the different tasks are encoded in different manners in the sentence embeddings, suggesting that abstract linguistic notions such as constituents or thematic roles does not seem to be present in the pretrained sentence embeddings.
Code (1)
Tasks
SentenceSentence EmbeddingsSimilar Papers 제목 키워드 기반
Exploring Semantic Properties of Sentence Embeddings
Neural vector representations are ubiquitous throughout all subfields of NLP. While word vectors have been studied in much detail, thus far only little light has been shed on the properties of sentence embeddings. In thi…
Machine TranslationReading ComprehensionSemantic Textual SimilaritySentence+3Exploring Sentence Vectors Through Automatic Summarization
Vector semantics, especially sentence vectors, have recently been used successfully in many areas of natural language processing. However, relatively little work has explored the internal structure and properties of spac…
SentenceSentence EmbeddingsExploring Sentence Vector Spaces through Automatic Summarization
Given vector representations for individual words, it is necessary to compute vector representations of sentences for many applications in a compositional manner, often using artificial neural networks. Relatively litt…
SentenceSentence EmbeddingsAs Good as New. How to Successfully Recycle English GPT-2 to Make Models for Other Languages
Large generative language models have been very successful for English, but other languages lag behind, in part due to data and computational limitations. We propose a method that may overcome these problems by adapting …
Discovering the Italian literature: interactive access to audio indexed text resources
In this paper we present a web interface to study Italian through the access to read Italian literature. The system allows to browse the content, search for specific words and listen to the correct pronunciation produced…
Cultural Vocal Bursts Intensity PredictionSentencetext-to-speechText to Speech