paper-with-me

홈 › Papers

DRAGOn: Designing RAG On Periodically Updated Corpus

2025-07-08 · Fedor Chernogorskii, Sergei Averkiev, Liliya Kudraleeva, Zaven Martirosian, Maria Tikhonova, Valentin Malykh, Alena Fenogenova arxiv

This paper introduces DRAGOn, method to design a RAG benchmark on a regularly updated corpus. It features recent reference datasets, a question generation framework, an automatic evaluation pipeline, and a public leaderboard. Specified reference datasets allow for uniform comparison of RAG systems, while newly generated dataset versions mitigate data leakage and ensure that all models are evaluated on unseen, comparable data. The pipeline for automatic question generation extracts the Knowledge Graph from the text corpus and produces multiple question-answer pairs utilizing modern LLM capabilities. A set of diverse LLM-as-Judge metrics is provided for a comprehensive model evaluation. We used Russian news outlets to form the datasets and demonstrate our methodology. We launch a public leaderboard to track the development of RAG systems and encourage community participation.

📄 PDF Abstract BibTeX arXiv:2507.05713

Code (0)

등록된 구현이 없습니다.

Tasks

Question Generation

Similar Papers 제목 키워드 기반

Collection and Annotation of the Romanian Legal Corpus

2020-05-01 · LREC 2020 5 · Dan Tufi{\textcommabelow{s}}, Maria Mitrofan, Vasile P{\u{a}}i{\textcommabelow{s}}, Radu Ion 외

We present the Romanian legislative corpus which is a valuable linguistic asset for the development of machine translation systems, especially for under-resourced languages. The knowledge that can be extracted from this …

Machine TranslationPOSTranslation

Synthesis and Evaluation of a Domain-specific Large Data Set for Dungeons & Dragons

2022-12-18 · Akila Peiris, Nisansa de Silva

This paper introduces the Forgotten Realms Wiki (FRW) data set and domain specific natural language generation using FRW along with related analyses. Forgotten Realms is the de-facto default setting of the popular open e…

ArticlesText Generation

TemporalWiki: A Lifelong Benchmark for Training and Evaluating Ever-Evolving Language Models

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Language Models (LMs) become outdated as the world changes; they often fail to perform tasks requiring recent factual information which was absent or different during training, a phenomenon called temporal misalignment. …

Continual Learning

DRAGON: Distributional Rewards Optimize Diffusion Generative Models

2025-04-21 · Yatong Bai, Jonah Casebeer, Somayeh Sojoudi, Nicholas J. Bryan

We present Distributional RewArds for Generative OptimizatioN (DRAGON), a versatile framework for fine-tuning media generation models towards a desired outcome. Compared with traditional reinforcement learning with human…

FAD

TemporalWiki: A Lifelong Benchmark for Training and Evaluating Ever-Evolving Language Models

2022-04-29 · Joel Jang, Seonghyeon Ye, Changho Lee, Sohee Yang 외

Language Models (LMs) become outdated as the world changes; they often fail to perform tasks requiring recent factual information which was absent or different during training, a phenomenon called temporal misalignment. …

Continual Learning