paper-with-me

홈 › Papers

SARAL: A Low-Resource Cross-Lingual Domain-Focused Information Retrieval System for Effective Rapid Document Triage

2019-07-01 · ACL 2019 7 · Elizabeth Boschee, Joel Barry, Jayadev Billa, Marjorie Freedman, Thamme Gowda, Constantine Lignos, Chester Palen-Michel, Michael Pust, Banriskhem Kayang Khonglah, Srikanth Madikeri, Jonathan May, Scott Miller

With the increasing democratization of electronic media, vast information resources are available in less-frequently-taught languages such as Swahili or Somali. That information, which may be crucially important and not available elsewhere, can be difficult for monolingual English speakers to effectively access. In this paper we present an end-to-end cross-lingual information retrieval (CLIR) and summarization system for low-resource languages that 1) enables English speakers to search foreign language repositories of text and audio using English queries, 2) summarizes the retrieved documents in English with respect to a particular information need, and 3) provides complete transcriptions and translations as needed. The SARAL system achieved the top end-to-end performance in the most recent IARPA MATERIAL CLIR+summarization evaluations. Our demonstration system provides end-to-end open query retrieval and summarization capability, and presents the original source text or audio, speech transcription, and machine translation, for two low resource languages.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Lingual Information RetrievalInformation RetrievalMachine TranslationRetrievalTranslation

Similar Papers 제목 키워드 기반

Beyond Ranked Lists: The SARAL Framework for Cross-Lingual Document Set Retrieval

2025-11-05 · Shantanu Agarwal, Joel Barry, Elizabeth Boschee, Scott Miller arxiv

Machine Translation for English Retrieval of Information in Any Language (MATERIAL) is an IARPA initiative targeted to advance the state of cross-lingual information retrieval (CLIR). This report provides a detailed desc…

Information RetrievalMachine Translation

NICT's participation to WAT 2019: Multilingualism and Multi-step Fine-Tuning for Low Resource NMT

2019-11-01 · WS 2019 11 · Raj Dabre, Eiichiro Sumita

In this paper we describe our submissions to WAT 2019 for the following tasks: English{--}Tamil translation and Russian{--}Japanese translation. Our team,{``}NICT-5{''}, focused on multilingual domain adaptation and back…

Domain AdaptationLow Resource NMTNMTTranslation

MDIA: A Benchmark for Multilingual Dialogue Generation in 46 Languages

2022-08-27 · Qingyu Zhang, Xiaoyu Shen, Ernie Chang, Jidong Ge 외

Owing to the lack of corpora for low-resource languages, current works on dialogue generation have mainly focused on English. In this paper, we present mDIA, the first large-scale multilingual benchmark for dialogue gene…

ChatbotDialogue GenerationDiversity

Monolingual and Multilingual Misinformation Detection for Low-Resource Languages: A Comprehensive Survey

2024-10-24 · Xinyu Wang, Wenbo Zhang, Sarah Rajtmajer

In today's global digital landscape, misinformation transcends linguistic boundaries, posing a significant challenge for moderation systems. Most approaches to misinformation detection are monolingual, focused on high-re…

Misinformation

XWikiGen: Cross-lingual Summarization for Encyclopedic Text Generation in Low Resource Languages

2023-03-22 · Dhaval Taunk, Shivprasad Sagare, Anupam Patil, Shivansh Subramanian 외

Lack of encyclopedic text contributors, especially on Wikipedia, makes automated text generation for low resource (LR) languages a critical problem. Existing work on Wikipedia text generation has focused on English only …

ArticlesCross-Lingual Abstractive SummarizationDocument SummarizationExtractive Summarization+3