paper-with-me

Papers

Towards Better Monolingual Japanese Retrievers with Multi-Vector Models

2023-12-26 · Benjamin Clavié

As language-specific training data tends to be sparsely available compared to English, document retrieval in many languages has been largely relying on multilingual models. In Japanese, the best performing deep-learning based retrieval approaches rely on multilingual dense embedders, with Japanese-only models lagging far behind. However, multilingual models require considerably more compute and data to train and have higher computational and memory requirements while often missing out on culturally-relevant information. In this paper, we introduce JaColBERT, a family of multi-vector retrievers trained on two magnitudes fewer data than their multilingual counterparts while reaching competitive performance. Our strongest model largely outperform all existing monolingual Japanese retrievers on all dataset, as well as the strongest existing multilingual models on all out-of-domain tasks, highlighting the need for specialised models able to handle linguistic specificities. These results are achieved using a model with only 110 million parameters, considerably smaller than all multilingual models, and using only a limited Japanese-language. We believe our results show great promise to support Japanese retrieval-enhanced application pipelines in a wide variety of domains.

📄 PDF Abstract BibTeX arXiv:2312.16144

Code (0)

등록된 구현이 없습니다.

Tasks

AllRetrieval

Similar Papers 제목 키워드 기반

JaColBERTv2.5: Optimising Multi-Vector Retrievers to Create State-of-the-Art Japanese Retrievers with Constrained Resources

2024-07-30 · Benjamin Clavié

Neural Information Retrieval has advanced rapidly in high-resource languages, but progress in lower-resource ones such as Japanese has been hindered by data scarcity, among other challenges. Consequently, multilingual mo…

Information RetrievalRetrieval

Pre-training via Leveraging Assisting Languages for Neural Machine Translation

2020-07-01 · ACL 2020 6 · Haiyue Song, Raj Dabre, Zhuoyuan Mao, Fei Cheng 외

Sequence-to-sequence (S2S) pre-training using large monolingual data is known to improve performance for various S2S NLP tasks. However, large monolingual corpora might not always be available for the languages of intere…

Machine TranslationNMTTranslation

Exploration of Language Dependency for Japanese Self-Supervised Speech Representation Models

2023-05-09 · Takanori Ashihara, Takafumi Moriya, Kohei Matsuura, Tomohiro Tanaka

Self-supervised learning (SSL) has been dramatically successful not only in monolingual but also in cross-lingual settings. However, since the two settings have been studied individually in general, there has been little…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Self-Supervised Learningspeech-recognition+1

Are Girls Neko or Shōjo? Cross-Lingual Alignment of Non-Isomorphic Embeddings with Iterative Normalization

2019-06-04 · Mozhi Zhang, Keyulu Xu, Ken-ichi Kawarabayashi, Stefanie Jegelka 외

Cross-lingual word embeddings (CLWE) underlie many multilingual natural language processing systems, often through orthogonal transformations of pre-trained monolingual embeddings. However, orthogonal mapping only works …

Cross-Lingual Word EmbeddingsTranslationWord EmbeddingsWord Translation

Are Girls Neko or Sh\=ojo? Cross-Lingual Alignment of Non-Isomorphic Embeddings with Iterative Normalization

2019-07-01 · ACL 2019 7 · Mozhi Zhang, Keyulu Xu, Ken-ichi Kawarabayashi, Stefanie Jegelka 외

Cross-lingual word embeddings (CLWE) underlie many multilingual natural language processing systems, often through orthogonal transformations of pre-trained monolingual embeddings. However, orthogonal mapping only works …

Cross-Lingual Word EmbeddingsTranslationWord EmbeddingsWord Translation