paper-with-me

홈 › Papers

Two LRL \& Distractor Corpora from Web Information Retrieval and a Small Case Study in Language Identification without Training Corpora

2020-05-01 · LREC 2020 5 · Armin Hoenen, Cemre Koc, Marc Rahn

In recent years, low resource languages (LRLs) have seen a surge in interest after certain tasks have been solved for larger ones and as they present various challenges (data sparsity, sparsity of experts and expertise, unusual structural properties etc.). For a larger number of them in the wake of this interest resources and technologies have been created. However, there are very small languages for which this has not yet led to a significant change. We focus here one such language (Nogai) and one larger small language (Maori). Since especially smaller languages often face the situation of having very similar siblings or a larger small sister language which is more accessible, the rate of noise in data gathered on them so far is often high. Therefore, we present small corpora for our 2 case study languages which we obtained through web information retrieval and likewise for their noise inducing distractor languages and conduct a small language identification experiment where we identify documents in a boolean way as either belonging or not to the target language. We release our test corpora for two such scenarios in the format of the An Crubadan project (Scannell, 2007) and a tool for unsupervised language identification using alphabet and toponym information.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalLanguage IdentificationRetrieval

Similar Papers 제목 키워드 기반

The First Drop of Ink: Nonlinear Impact of Distracting Information in Long-Context Reasoning

2026-05-11 · Muhan Gao, Zih-Ching Chen, Kuan-Hao Huang arxiv

As large language models are increasingly deployed in retrieval-augmented generation and agentic systems that accumulate extensive context, understanding how distracting information affects long-context performance becom…

SPECTRA: Synthetic IR Test Collections with Relevance Oracles and Controlled Distractor Diagnostics

2026-05-29 · Eric Liang arxiv

Scalable information retrieval testing needs corpora that are large enough to stress index construction, ranking latency, query routing, and evaluation tooling, yet human-judged test collections remain expensive and may …

Information Retrieval

Automatic Distractor Suggestion for Multiple-Choice Tests Using Concept Embeddings and Information Retrieval

2018-06-01 · WS 2018 6 · Le An Ha, Victoria Yaneva

Developing plausible distractors (wrong answer options) when writing multiple-choice questions has been described as one of the most challenging and time-consuming parts of the item-writing process. In this paper we prop…

Information RetrievalMultiple-choiceRetrieval

RAGShaper: Eliciting Sophisticated Agentic RAG Skills via Automated Data Synthesis

2026-01-13 · Zhengwei Tao, Bo Li, Jialong Wu, Guochen Yan 외 arxiv

Agentic Retrieval-Augmented Generation (RAG) empowers large language models to autonomously plan and retrieve information for complex problem-solving. However, the development of robust agents is hindered by the scarcity…

Knowledge-Driven Distractor Generation for Cloze-style Multiple Choice Questions

2020-04-21 · Siyu Ren, Kenny Q. Zhu

In this paper, we propose a novel configurable framework to automatically generate distractive choices for open-domain cloze-style multiple-choice questions, which incorporates a general-purpose knowledge base to effecti…

Distractor GenerationLearning-To-RankMultiple-choice