The 2019 BBN Cross-lingual Information Retrieval System
In this paper, we describe a cross-lingual information retrieval (CLIR) system that, given a query in English, and a set of audio and text documents in a foreign language, can return a scored list of relevant documents, and present findings in a summary form in English. Foreign audio documents are first transcribed by a state-of-the-art pretrained multilingual speech recognition model that is finetuned to the target language. For text documents, we use multiple multilingual neural machine translation (MT) models to achieve good translation results, especially for low/medium resource languages. The processed documents and queries are then scored using a probabilistic CLIR model that makes use of the probability of translation from GIZA translation tables and scores from a Neural Network Lexical Translation Model (NNLTM). Additionally, advanced score normalization, combination, and thresholding schemes are employed to maximize the Average Query Weighted Value (AQWV) scores. The CLIR output, together with multiple translation renderings, are selected and translated into English snippets via a summarization model. Our turnkey system is language agnostic and can be quickly trained for a new low-resource language in few days.
Code (0)
등록된 구현이 없습니다.
Tasks
Cross-Lingual Information RetrievalInformation RetrievalMachine TranslationRetrievalspeech-recognitionSpeech RecognitionTranslationSimilar Papers 제목 키워드 기반
The Cross-Lingual Arabic Information REtrieval (CLAIRE) System
Despite advances in neural machine translation, cross-lingual retrieval tasks in which queries and documents live in different natural language spaces remain challenging. Although neural translation models may provide an…
Information RetrievalMachine TranslationRetrievalTranslationBridging Language Gaps: Advances in Cross-Lingual Information Retrieval with Multilingual LLMs
Cross-lingual information retrieval (CLIR) addresses the challenge of retrieving relevant documents written in languages different from that of the original query. Research in this area has typically framed the task as m…
Information RetrievalQuestion AnsweringAnswer GenerationA Cross-Lingual Statutory Article Retrieval Dataset for Taiwan Legal Studies
This paper introduces a cross-lingual statutory article retrieval (SAR) dataset designed to enhance legal information retrieval in multilingual settings. Our dataset features spoken-language-style legal inquiries in Engl…
Information RetrievalRetrievalMMed-Bench-IR: A Heterogeneous Benchmark for Multilingual Medical Information Retrieval
Retrieval-augmented generation (RAG) in clinical settings increasingly requires multilingual retrieval against predominantly English evidence corpora. Multilingual medical retrieval demands three capabilities: cross-ling…
Information RetrievalCONCRETE: Improving Cross-lingual Fact-checking with Cross-lingual Retrieval
Fact-checking has gained increasing attention due to the widespread of falsified information. Most fact-checking approaches focus on claims made in English only due to the data scarcity issue in other languages. The lack…
Cross-lingual Fact-checkingCross-Lingual Information RetrievalCross-Lingual TransferFact Checking+3