paper-with-me

홈 › Papers

LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes

2026-07-14 · Michael Solodko, Steven Gong, Guangwei Yu, Satya Krishna Gorti, Jesse C. Cresswell, Victor Zhong arxiv

While modern question answering (QA) systems excel on clean, schema-aligned corpora, real-world knowledge is rarely so neatly packaged. Answering questions over enterprise and scientific data lakes requires systems to navigate heterogeneous, weakly structured collections of tables, passages, and linked metadata. Current benchmarks abstract away this noisy discovery process, failing to evaluate end-to-end performance. To bridge this gap, we introduce LakeQuest, a human-validated benchmark of 9,846 QA pairs designed to evaluate the end-to-end retrieve-and-synthesize pipeline over realistic data lakes. LakeQuest spans three diverse domains (AI/ML metadata, retail banking, and multimodal biomedical drug information) and pairs every question with exact, modality-aware evidence pointers. By isolating source discovery from cross-modal synthesis, LakeQuest exposes critical failure modes in modern QA systems. Our baseline evaluations, including standard Retrieval-Augmented Generation (RAG) and agentic tool-use methods, reveal that high-quality retrieval does not guarantee correct reasoning. Systems consistently struggle with relation chaining in metadata graphs, policy grounding in bank ledgers, and joint tabular QA in biomedical contexts, highlighting the need for robust discovery and faithful cross-file composition mechanisms in future agentic QA systems.

📄 PDF Abstract BibTeX arXiv:2607.12310

Code (2)

Tavish9/awesome-daily-AI-arxiv ★ 111
arxivsub/arXivSub_daily_arxiv ★ 3

Tasks

Question Answering

Similar Papers 제목 키워드 기반

Interpreting Answers to Yes-No Questions in Dialogues from Multiple Domains

2024-04-25 · Zijie Wang, Farzana Rashid, Eduardo Blanco

People often answer yes-no questions without explicitly saying yes, no, or similar polar keywords. Figuring out the meaning of indirect answers is challenging, even for large language models. In this paper, we investigat…

KoBLEX: Open Legal Question Answering with Multi-hop Reasoning

2025-09-01 · Jihyung Lee, Daehui Kim, Seonjeong Hwang, Hyounghun Kim 외 arxiv

Large Language Models (LLM) have achieved remarkable performances in general domains and are now extending into the expert domain of law. Several benchmarks have been proposed to evaluate LLMs' legal capabilities. Howeve…

Question AnsweringLegal Reasoning

WixQA: A Multi-Dataset Benchmark for Enterprise Retrieval-Augmented Generation

2025-05-13 · Dvir Cohen, Lin Burg, Sviatoslav Pykhnivskyi, Hagit Gur 외

Retrieval-Augmented Generation (RAG) is a cornerstone of modern question answering (QA) systems, enabling grounded answers based on external knowledge. Although recent progress has been driven by open-domain datasets, en…

Question AnsweringRAGRetrievalRetrieval-augmented Generation

Automated Benchmark Generation from Domain Guidelines Informed by Bloom's Taxonomy

2026-01-28 · Si Chen, Le Huy Khiem, Annalisa Szymanski, Ronald Metoyer 외 arxiv

Open-ended question answering (QA) evaluates a model's ability to perform contextualized reasoning beyond factual recall. This challenge is especially acute in practice-based domains, where knowledge is procedural and gr…

Question Answering

ESGBench: A Benchmark for Explainable ESG Question Answering in Corporate Sustainability Reports

2025-11-20 · Sherine George, Nithish Saji arxiv

We present ESGBench, a benchmark dataset and evaluation framework designed to assess explainable ESG question answering systems using corporate sustainability reports. The benchmark consists of domain-grounded questions …

Question Answering