paper-with-me

홈 › Papers

LIT-RAGBench: Benchmarking Generator Capabilities of Large Language Models in Retrieval-Augmented Generation

2026-03-06 · Koki Itai, Shunichi Hasegawa, Yuta Yamamoto, Gouki Minegishi, Masaki Otsuki arxiv

Retrieval-Augmented Generation (RAG) is a framework in which a Generator, such as a Large Language Model (LLM), produces answers by retrieving documents from an external collection using a Retriever. In practice, Generators must integrate evidence from long contexts, perform multi-step reasoning, interpret tables, and abstain when evidence is missing. However, existing benchmarks for Generators provide limited coverage, with none enabling simultaneous evaluation of multiple capabilities under unified conditions. To bridge the gap between existing evaluations and practical use, we introduce LIT-RAGBench (the Logic, Integration, Table, Reasoning, and Abstention RAG Generator Benchmark), which defines five categories: Integration, Reasoning, Logic, Table, and Abstention, each further divided into practical evaluation aspects. LIT-RAGBench systematically covers patterns combining multiple aspects across categories. By using fictional entities and scenarios, LIT-RAGBench evaluates answers grounded in the provided external documents. The dataset consists of 114 human-constructed Japanese questions and an English version generated by machine translation with human curation. We use LLM-as-a-Judge for scoring and report category-wise and overall accuracy. Across API-based and open-weight models, no model exceeds 90% overall accuracy. By making strengths and weaknesses measurable within each category, LIT-RAGBench serves as a valuable metric for model selection in practical RAG deployments and for building RAG-specialized models. We release LIT-RAGBench, including the dataset and evaluation code, at https://github.com/Koki-Itai/LIT-RAGBench.

📄 PDF Abstract BibTeX arXiv:2603.06198

Code (0)

등록된 구현이 없습니다.

Tasks

Machine Translation

Similar Papers 제목 키워드 기반

RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems

2024-06-25 · Robert Friel, Masha Belyi, Atindriyo Sanyal

Retrieval-Augmented Generation (RAG) has become a standard architectural pattern for incorporating domain-specific knowledge into user-facing chat applications powered by Large Language Models (LLMs). RAG systems are cha…

BenchmarkingRAGRetrievalRetrieval-augmented Generation

FragBench: Cross-Session Attacks Hidden in Benign-Looking Fragments

2026-05-10 · Astha Mehta, Niruthiha Selvanayagam, Cedric Lam, Hengxu Li 외 arxiv

An attacker can split a malicious goal into sub-prompts that each look benign on their own and only become harmful in combination. Existing LLM safety benchmarks evaluate prompts one at a time, or across turns of a singl…

FinRAGBench-V: A Benchmark for Multimodal RAG with Visual Citation in the Financial Domain

2025-05-23 · Suifeng Zhao, Zhuoran Jin, Sujian Li, Jun Gao

Retrieval-Augmented Generation (RAG) plays a vital role in the financial domain, powering applications such as real-time market analysis, trend forecasting, and interest rate computation. However, most existing RAG resea…

Question AnsweringRAGRetrievalRetrieval-augmented Generation

On the Benchmarking of LLMs for Open-Domain Dialogue Evaluation

2024-07-04 · John Mendonça, Alon Lavie, Isabel Trancoso

Large Language Models (LLMs) have showcased remarkable capabilities in various Natural Language Processing tasks. For automatic open-domain dialogue evaluation in particular, LLMs have been seamlessly integrated into eva…

BenchmarkingChatbotDialogue Evaluation

Rethinking RAG in Long Videos: What to Retrieve and How to Use It?

2026-06-11 · Yuho Lee, Jisu Shin, Nicole Hee-Yeon Kim, Jihwan Bang 외 arxiv

Retrieval-augmented generation is moving beyond text into long, egocentric video, where systems must select query-relevant chunks across multiple modalities and temporal granularities. Yet progress in VideoRAG is limited…