paper-with-me

Papers

Benchmarking and Building Zero-Shot Hindi Retrieval Model with Hindi-BEIR and NLLB-E5

2024-09-09 · Arkadeep Acharya, Rudra Murthy, Vishwajeet Kumar, Jaydeep Sen

Given the large number of Hindi speakers worldwide, there is a pressing need for robust and efficient information retrieval systems for Hindi. Despite ongoing research, comprehensive benchmarks for evaluating retrieval models in Hindi are lacking. To address this gap, we introduce the Hindi-BEIR benchmark, comprising 15 datasets across seven distinct tasks. We evaluate state-of-the-art multilingual retrieval models on the Hindi-BEIR benchmark, identifying task and domain-specific challenges that impact Hindi retrieval performance. Building on the insights from these results, we introduce NLLB-E5, a multilingual retrieval model that leverages a zero-shot approach to support Hindi without the need for Hindi training data. We believe our contributions, which include the release of the Hindi-BEIR benchmark and the NLLB-E5 model, will prove to be a valuable resource for researchers and promote advancements in multilingual retrieval models.

📄 PDF Abstract BibTeX arXiv:2409.05401

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingInformation RetrievalRetrievalText Retrieval

Similar Papers 제목 키워드 기반

Bilingual BSARD: Extending Statutory Article Retrieval to Dutch

2024-12-10 · Ehsan Lotfi, Nikolay Banar, Nerses Yuzbashyan, Walter Daelemans

Statutory article retrieval plays a crucial role in making legal information more accessible to both laypeople and legal professionals. Multilingual countries like Belgium present unique challenges for retrieval models d…

ArticlesBenchmarkingRetrieval

Cross-Lingual Relevance Transfer for Document Retrieval

2019-11-08 · Peng Shi, Jimmy Lin

Recent work has shown the surprising ability of multi-lingual BERT to serve as a zero-shot cross-lingual transfer model for a number of language processing tasks. We combine this finding with a similarly-recently proposa…

Cross-Lingual TransferRetrievalSentenceZero-Shot Cross-Lingual Transfer

Zero-shot Disfluency Detection for Indian Languages

2022-10-01 · COLING 2022 10 · Rohit Kundu, Preethi Jyothi, Pushpak Bhattacharyya

Disfluencies that appear in the transcriptions from automatic speech recognition systems tend to impair the performance of downstream NLP tasks. Disfluency correction models can help alleviate this problem. However, the …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Unsupervised Approach for Zero-Shot Experiments: Bhojpuri–Hindi and Magahi–Hindi@LoResMT 2020

2020-12-01 · loresmt (AACL) 2020 12 · Amit Kumar, Rajesh Kumar Mundotiya, Anil Kumar Singh

This paper reports a Machine Translation (MT) system submitted by the NLPRL team for the Bhojpuri–Hindi and Magahi–Hindi language pairs at LoResMT 2020 shared task. We used an unsupervised domain adaptation approach that…

Domain AdaptationMachine TranslationTranslationUnsupervised Domain Adaptation

Cross-Lingual Training with Dense Retrieval for Document Retrieval

2021-09-03 · Peng Shi, Rui Zhang, He Bai, Jimmy Lin

Dense retrieval has shown great success in passage ranking in English. However, its effectiveness in document retrieval for non-English languages remains unexplored due to the limitation in training resources. In this wo…

Document RankingPassage RankingRetrieval