paper-with-me

Papers

The Russian-focused embedders' exploration: ruMTEB benchmark and Russian embedding model design

2024-08-22 · Artem Snegirev, Maria Tikhonova, Anna Maksimova, Alena Fenogenova, Alexander Abramov

Embedding models play a crucial role in Natural Language Processing (NLP) by creating text embeddings used in various tasks such as information retrieval and assessing semantic text similarity. This paper focuses on research related to embedding models in the Russian language. It introduces a new Russian-focused embedding model called ru-en-RoSBERTa and the ruMTEB benchmark, the Russian version extending the Massive Text Embedding Benchmark (MTEB). Our benchmark includes seven categories of tasks, such as semantic textual similarity, text classification, reranking, and retrieval.The research also assesses a representative set of Russian and multilingual models on the proposed benchmark. The findings indicate that the new model achieves results that are on par with state-of-the-art models in Russian. We release the model ru-en-RoSBERTa, and the ruMTEB framework comes with open-source code, integration into the original framework and a public leaderboard.

📄 PDF Abstract BibTeX arXiv:2408.12503

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalRerankingRetrievalSemantic Textual Similaritytext-classificationText Classificationtext similarity

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

GigaEmbeddings: Efficient Russian Language Embedding Model

2025-10-25 · Egor Kolodin, Daria Khomich, Nikita Savushkin, Anastasia Ianina 외 arxiv

We introduce GigaEmbeddings, a novel framework for training high-performance Russian-focused text embeddings through hierarchical instruction tuning of the decoder-only LLM designed specifically for Russian language (Gig…

Synthetic Data Generation

Situating Sentence Embedders with Nearest Neighbor Overlap

2019-09-24 · ICLR 2020 1 · Lucy H. Lin, Noah A. Smith

As distributed approaches to natural language semantics have developed and diversified, embedders for linguistic units larger than words have come to play an increasingly important role. To date, such embedders have been…

Sentence

LLM-based Embedders for Prior Case Retrieval

2025-07-24 · Damith Premasiri, Tharindu Ranasinghe, Ruslan Mitkov arxiv

In common law systems, legal professionals such as lawyers and judges rely on precedents to build their arguments. As the volume of cases has grown massively over time, effectively retrieving prior cases has become essen…

Information Retrieval

OBSR: Open Benchmark for Spatial Representations

2025-10-07 · Julia Moska, Oleksii Furman, Kacper Kozaczko, Szymon Leszkiewicz 외 arxiv

GeoAI is evolving rapidly, fueled by diverse geospatial datasets like traffic patterns, environmental data, and crowdsourced OpenStreetMap (OSM) information. While sophisticated AI models are being developed, existing be…

Omni-Interactive Universal Embedder

2026-08-27 · Wei-Yao Wang, Kazuya Tateishi, Shuyang Cui, Christian Simon 외 arxiv

Multimodal representation learning has been shifting from traditional two-tower architectures to large language model (LLM)-based embedders due to their strong instruction-following capabilities. Despite this progress, e…

Representation Learning