paper-with-me

MTEB

Massive Text Embedding Benchmark

홈페이지 · 논문 155편

MTEB is a benchmark that spans 8 embedding tasks covering a total of 56 datasets and 112 languages. The 8 task types are Bitext mining, Classification, Clustering, Pair Classification, Reranking, Retrieval, Semantic Textual Similarity and Summarisation. The 56 datasets contain varying text lengths and they are grouped into three categories: Sentence to sentence, Paragraph to paragraph, and Sentence to paragraph. Check the latest leaderboards at HuggingFace.

Texts

벤치마크

Text Classification on MTEB 결과 62개
Text Clustering on MTEB 결과 31개
Semantic Textual Similarity on MTEB 결과 30개
Text Retrieval on MTEB 결과 30개
Text Summarization on MTEB 결과 26개
Information Retrieval on MTEB 결과 2개