Cross-Dialect Information Retrieval: Information Access in Low-Resource and High-Variance Languages
A large amount of local and culture-specific knowledge (e.g., people, traditions, food) can only be found in documents written in dialects. While there has been extensive research conducted on cross-lingual information retrieval (CLIR), the field of cross-dialect retrieval (CDIR) has received limited attention. Dialect retrieval poses unique challenges due to the limited availability of resources to train retrieval models and the high variability in non-standardized languages. We study these challenges on the example of German dialects and introduce the first German dialect retrieval dataset, dubbed WikiDIR, which consists of seven German dialects extracted from Wikipedia. Using WikiDIR, we demonstrate the weakness of lexical methods in dealing with high lexical variation in dialects. We further show that commonly used zero-shot cross-lingual transfer approach with multilingual encoders do not transfer well to extremely low-resource setups, motivating the need for resource-lean and dialect-specific retrieval models. We finally demonstrate that (document) translation is an effective way to reduce the dialect gap in CDIR.
Code (1)
Tasks
Cross-Lingual Information RetrievalCross-Lingual TransferDocument TranslationInformation RetrievalRetrievalZero-Shot Cross-Lingual TransferSimilar Papers 제목 키워드 기반
Resource-Lean Lexicon Induction for German Dialects
Automatic induction of high-quality dictionaries is essential for building lexical resources, yet low-resource languages and dialects pose several challenges: limited access to annotators, high degree of spelling variati…
Information RetrievalMessIRve: A Large-Scale Spanish Information Retrieval Dataset
Information retrieval (IR) is the task of finding relevant documents in response to a user query. Although Spanish is the second most spoken native language, current IR benchmarks lack Spanish data, hindering the develop…
Information RetrievalRetrievalThe Ontology of Bulgarian Dialects -- Architecture and Information Retrieval
Following a concise description of the structure, the paper focuses on the potential of the Ontology of the Bulgarian Dialects, which demonstrates a novel usage of the ontological modelling for the purposes of dialect di…
DiagnosticInformation RetrievalRetrievalDIRA: Dialectal Arabic Information Retrieval Assistant
A Multi-Agent Framework for Mitigating Dialect Biases in Privacy Policy Question-Answering Systems
Privacy policies inform users about data collection and usage, yet their complexity limits accessibility for diverse populations. Existing Privacy Policy Question Answering (QA) systems exhibit performance disparities ac…
Question Answering