paper-with-me

홈 › Papers

AfriMTEB and AfriE5: Benchmarking and Adapting Text Embedding Models for African Languages

2025-10-27 · Kosei Uemura, Miaoran Zhang, David Ifeoluwa Adelani arxiv

Text embeddings are an essential building component of several NLP tasks such as retrieval-augmented generation which is crucial for preventing hallucinations in LLMs. Despite the recent release of massively multilingual MTEB (MMTEB), African languages remain underrepresented, with existing tasks often repurposed from translation benchmarks such as FLORES clustering or SIB-200. In this paper, we introduce AfriMTEB -- a regional expansion of MMTEB covering 59 languages, 14 tasks, and 38 datasets, including six newly added datasets. Unlike many MMTEB datasets that include fewer than five languages, the new additions span 14 to 56 African languages and introduce entirely new tasks, such as hate speech detection, intent detection, and emotion classification, which were not previously covered. Complementing this, we present AfriE5, an adaptation of the instruction-tuned mE5 model to African languages through cross-lingual contrastive distillation. Our evaluation shows that AfriE5 achieves state-of-the-art performance, outperforming strong baselines such as Gemini-Embeddings and mE5.

📄 PDF Abstract BibTeX arXiv:2510.23896

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion ClassificationHate Speech DetectionIntent Detection

Similar Papers 제목 키워드 기반

AfriEconQA: A Benchmark Dataset for African Economic Analysis based on World Bank Reports

2026-01-06 · Edward Ajayi arxiv

We introduce AfriEconQA, a specialized benchmark dataset for African economic analysis grounded in a comprehensive corpus of 236 World Bank reports. The task of AfriEconQA is to answer complex economic queries that requi…

Information Retrieval

Refining Joint Text and Source Code Embeddings for Retrieval Task with Parameter-Efficient Fine-Tuning

2024-05-07 · Karim Galliamov, Leila Khaertdinova, Karina Denisova

The latest developments in Natural Language Processing (NLP) have demonstrated remarkable progress in a code-text retrieval problem. As the Transformer-based models used in this task continue to increase in size, the com…

BenchmarkingContrastive Learningparameter-efficient fine-tuningRetrieval+1

Choosing a Text Embedding Model: A Practical Benchmarking and Decision Framework

2026-07-26 · Madhav S Baidya arxiv

Choosing the right text embedding model is one of the most consequential -- and most frequently under-examined -- decisions in building a retrieval or search system, yet the model that tops a leaderboard is rarely the be…

Semantic Similarity

Adapting Deep Learning for Sentiment Classification of Code-Switched Informal Short Text

2020-01-04 · Muhammad Haroon Shakeel, Asim Karim

Nowadays, an abundance of short text is being generated that uses nonstandard writing styles influenced by regional languages. Such informal and code-switched content are under-resourced in terms of labeled datasets and …

ClassificationGeneral ClassificationLexical NormalizationSentiment Analysis+2

Benchmarking pre-trained text embedding models in aligning built asset information

2024-11-18 · Mehrzad Shahinmoghadam, Ali Motamedi

Accurate mapping of the built asset information to established data classification systems and taxonomies is crucial for effective asset management, whether for compliance at project handover or ad-hoc data integration s…

Asset ManagementBenchmarkingData IntegrationDomain Adaptation+2