Open Sentence Embeddings for Portuguese with the Serafim PT* encoders family
Sentence encoder encode the semantics of their input, enabling key downstream applications such as classification, clustering, or retrieval. In this paper, we present Serafim PT*, a family of open-source sentence encoders for Portuguese with various sizes, suited to different hardware/compute budgets. Each model exhibits state-of-the-art performance and is made openly available under a permissive license, allowing its use for both commercial and research purposes. Besides the sentence encoders, this paper contributes a systematic study and lessons learned concerning the selection criteria of learning objectives and parameters that support top-performing encoders.
Code (0)
등록된 구현이 없습니다.
Tasks
ClusteringRetrievalSentenceSentence EmbeddingsSimilar Papers 제목 키워드 기반
Fast developing of a Natural Language Interface for a Portuguese WordNet: Leveraging on Sentence Embeddings
We describe how a natural language interface can be developed for a wordnet with a small set of handcrafted templates, leveraging on sentence embeddings. The proposed approach does not use rules for parsing natural langu…
Natural Language QueriesSentenceSentence EmbeddingsUnsupervised Transfer Learning in Multilingual Neural Machine Translation with Cross-Lingual Word Embeddings
In this work we look into adding a new language to a multilingual NMT system in an unsupervised fashion. Under the utilization of pre-trained cross-lingual word embeddings we seek to exploit a language independent multil…
Cross-Lingual Word EmbeddingsMachine TranslationNMTSentence+3Fostering the Ecosystem of Open Neural Encoders for Portuguese with Albertina PT* Family
To foster the neural encoding of Portuguese, this paper contributes foundation encoder models that represent an expansion of the still very scarce ecosystem of large language models specifically developed for this langua…
Beyond Multilingual Averages: MTEB-PT, a Benchmark for Portuguese Sentence Encoders
Portuguese remains underrepresented in text embedding evaluation, despite being one of the most widely spoken languages in the world. As a result, embedding models are often selected based on English or multilingual metr…
Semantic Textual SimilarityRepresentation LearningDeBERTinha: A Multistep Approach to Adapt DebertaV3 XSmall for Brazilian Portuguese Natural Language Processing Task
This paper presents an approach for adapting the DebertaV3 XSmall model pre-trained in English for Brazilian Portuguese natural language processing (NLP) tasks. A key aspect of the methodology involves a multistep traini…
named-entity-recognitionNamed Entity RecognitionSentenceSentiment Analysis