paper-with-me

홈 › Papers

SEA-LION-Embedding: Open and Reproducible Text Embeddings for Southeast Asia

2026-06-02 · Peerat Limkonchotiwat, Raymond Ng, Sarana Nutanong, Jian Gang Ngui arxiv

Text embeddings are fundamental to many downstream applications, making robustness important for real-world NLP. However, most recent state-of-the-art embedding models are not reproducible because they rely on closed or undisclosed training data, and they remain insufficiently robust for Southeast Asian languages. We present SEA-LION-Embedding, a fully open and reproducible text-embedding pipeline for Southeast Asian languages trained only on publicly available data, and use it to study three core factors of robust embedding design: data composition, training objective, and base encoder initialization. SEA-LION-Embedding achieves state-of-the-art results on SEA-BED while enabling systematic and reproducible analysis of robust text embeddings for the region.

📄 PDF Abstract BibTeX arXiv:2606.03027

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Nomic Embed: Training a Reproducible Long Context Text Embedder

2024-02-02 · Zach Nussbaum, John X. Morris, Brandon Duderstadt, Andriy Mulyar

This technical report describes the training of nomic-embed-text-v1, the first fully reproducible, open-source, open-weights, open-data, 8192 context length English text embedding model that outperforms both OpenAI Ada-0…

Incorporating LLM Embeddings for Variation Across the Human Genome

2025-09-25 · Hongqian Niu, Jordan Bryan, Jacob Williams, Hufeng Zhou 외 arxiv

Recent advances in large language model (LLM) embeddings have enabled powerful representations for biological data, but most applications to date focus on gene-level information. We present one of the first systematic fr…

LMEB: Long-horizon Memory Embedding Benchmark

2026-03-13 · Xinping Zhao, Xinshuo Hu, Jiaxin Xu, Danyu Tang 외 arxiv

Memory embeddings are crucial for memory-augmented systems, such as OpenClaw, but their evaluation is underexplored in current text embedding benchmarks, which narrowly focus on traditional passage retrieval and fail to …

Passage Retrieval

jina-embeddings-v3: Multilingual Embeddings With Task LoRA

2024-09-16 · Saba Sturua, Isabelle Mohr, Mohammad Kalim Akram, Michael Günther 외

We introduce jina-embeddings-v3, a novel text embedding model with 570 million parameters, achieves state-of-the-art performance on multilingual data and long-context retrieval tasks, supporting context lengths of up to …

MTEB BenchmarkRepresentation LearningRetrievalText Matching

ChEmbed: Enhancing Chemical Literature Search Through Domain-Specific Text Embeddings

2025-08-03 · Ali Shiraee Kasmaee, Mohammad Khodadad, Mehdi Astaraki, Mohammad Arshi Saloot 외 arxiv

Retrieval-Augmented Generation (RAG) systems in chemistry heavily depend on accurate and relevant retrieval of chemical literature. However, general-purpose text embedding models frequently fail to adequately represent c…