paper-with-me

홈 › Papers

Open Sentence Embeddings for Portuguese with the Serafim PT* encoders family

2024-07-28 · Luís Gomes, António Branco, João Silva, João Rodrigues, Rodrigo Santos

Sentence encoder encode the semantics of their input, enabling key downstream applications such as classification, clustering, or retrieval. In this paper, we present Serafim PT*, a family of open-source sentence encoders for Portuguese with various sizes, suited to different hardware/compute budgets. Each model exhibits state-of-the-art performance and is made openly available under a permissive license, allowing its use for both commercial and research purposes. Besides the sentence encoders, this paper contributes a systematic study and lessons learned concerning the selection criteria of learning objectives and parameters that support top-performing encoders.

📄 PDF Abstract BibTeX arXiv:2407.19527

Code (0)

등록된 구현이 없습니다.

Tasks

ClusteringRetrievalSentenceSentence Embeddings

Similar Papers 제목 키워드 기반

Fast developing of a Natural Language Interface for a Portuguese WordNet: Leveraging on Sentence Embeddings

2019-07-01 · GWC 2019 7 · Hugo Gonçalo Oliveira, Alexandre Rademaker

We describe how a natural language interface can be developed for a wordnet with a small set of handcrafted templates, leveraging on sentence embeddings. The proposed approach does not use rules for parsing natural langu…

Natural Language QueriesSentenceSentence Embeddings

Unsupervised Transfer Learning in Multilingual Neural Machine Translation with Cross-Lingual Word Embeddings

2021-03-11 · Carlos Mullov, Ngoc-Quan Pham, Alexander Waibel

In this work we look into adding a new language to a multilingual NMT system in an unsupervised fashion. Under the utilization of pre-trained cross-lingual word embeddings we seek to exploit a language independent multil…

Cross-Lingual Word EmbeddingsMachine TranslationNMTSentence+3

Fostering the Ecosystem of Open Neural Encoders for Portuguese with Albertina PT* Family

2024-03-04 · Rodrigo Santos, João Rodrigues, Luís Gomes, João Silva 외

To foster the neural encoding of Portuguese, this paper contributes foundation encoder models that represent an expansion of the still very scarce ecosystem of large language models specifically developed for this langua…

Beyond Multilingual Averages: MTEB-PT, a Benchmark for Portuguese Sentence Encoders

2026-07-05 · Lucas Hideki Takeuchi Okamura, Alexandre Alcoforado, Anna Helena Reali Costa arxiv

Portuguese remains underrepresented in text embedding evaluation, despite being one of the most widely spoken languages in the world. As a result, embedding models are often selected based on English or multilingual metr…

Semantic Textual SimilarityRepresentation Learning

DeBERTinha: A Multistep Approach to Adapt DebertaV3 XSmall for Brazilian Portuguese Natural Language Processing Task

2023-09-28 · Israel Campiotti, Matheus Rodrigues, Yuri Albuquerque, Rafael Azevedo 외

This paper presents an approach for adapting the DebertaV3 XSmall model pre-trained in English for Brazilian Portuguese natural language processing (NLP) tasks. A key aspect of the methodology involves a multistep traini…

named-entity-recognitionNamed Entity RecognitionSentenceSentiment Analysis