paper-with-me

Papers

Arctic-Embed 2.0: Multilingual Retrieval Without Compromise

2024-12-03 · Puxuan Yu, Luke Merrick, Gaurav Nuti, Daniel Campos

This paper presents the training methodology of Arctic-Embed 2.0, a set of open-source text embedding models built for accurate and efficient multilingual retrieval. While prior works have suffered from degraded English retrieval quality, Arctic-Embed 2.0 delivers competitive retrieval quality on multilingual and English-only benchmarks, and supports Matryoshka Representation Learning (MRL) for efficient embedding storage with significantly lower compressed quality degradation compared to alternatives. We detail the design and implementation, presenting several important open research questions that arose during model development. We conduct experiments exploring these research questions and include extensive discussion aimed at fostering further discussion in this field.

📄 PDF Abstract BibTeX arXiv:2412.04506

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningRetrieval

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Arctic-Embed: Scalable, Efficient, and Accurate Text Embedding Models

2024-05-08 · Luke Merrick, Danmei Xu, Gaurav Nuti, Daniel Campos

This report describes the training dataset creation and recipe behind the family of \texttt{arctic-embed} text embedding models (a set of five models ranging from 22 to 334 million parameters with weights open-sourced un…

Retrieval

Optimized Text Embedding Models and Benchmarks for Amharic Passage Retrieval

2025-05-25 · Kidist Amde Mekonnen, Yosef Worku Alemneh, Maarten de Rijke

Neural retrieval methods using transformer-based pre-trained language models have advanced multilingual and cross-lingual retrieval. However, their effectiveness for low-resource, morphologically rich languages such as A…

Passage RetrievalRetrieval

IRSC: A Zero-shot Evaluation Benchmark for Information Retrieval through Semantic Comprehension in Retrieval-Augmented Generation Scenarios

2024-09-24 · Hai Lin, Shaoxiong Zhan, Junyou Su, Haitao Zheng 외

In Retrieval-Augmented Generation (RAG) tasks using Large Language Models (LLMs), the quality of retrieved information is critical to the final output. This paper introduces the IRSC benchmark for evaluating the performa…

Information RetrievalRAGRetrievalRetrieval-augmented Generation

ORPHEAS: A Cross-Lingual Greek-English Embedding Model for Retrieval-Augmented Generation

2026-04-22 · Ioannis E. Livieris, Athanasios Koursaris, Alexandra Apostolopoulou, Konstantinos Kanaris Dimitris Tsakalidis 외 arxiv

Effective retrieval-augmented generation across bilingual Greek--English applications requires embedding models capable of capturing both domain-specific semantic relationships and cross-lingual semantic alignment. Exist…

jina-embeddings-v3: Multilingual Embeddings With Task LoRA

2024-09-16 · Saba Sturua, Isabelle Mohr, Mohammad Kalim Akram, Michael Günther 외

We introduce jina-embeddings-v3, a novel text embedding model with 570 million parameters, achieves state-of-the-art performance on multilingual data and long-context retrieval tasks, supporting context lengths of up to …

MTEB BenchmarkRepresentation LearningRetrievalText Matching