paper-with-me

홈 › Papers

LakeQA: An Exploratory QA Benchmark over a Million-Scale Data Lake

2026-06-09 · Haonan Wang, Jiaxiang Liu, Yurong Liu, Austin Senna Wijaya, Tianle Zhou, Eden Wu, Yijia Chen, Wanting You, Reya Vir, Daniela Pinto, Grace Fan, Yusen Zhang, Juliana Freire, Eugene Wu arxiv

Recent large language models (LLMs) have shown rapid progress in reading-based question answering (QA), where evidence is explicitly provided or can be trivially retrieved. In contrast, real-world questions are often not paired with accurate evidence documents. The useful evidence resides in massive data lakes, making search a prerequisite for answering. However, there is a lack of comprehensive benchmarks that require both searching and reasoning over large data lakes. To this end, we introduce LakeQA, a comprehensive benchmark for search-centric question answering over data lakes that jointly emphasizes searching and reasoning capabilities. LakeQA is built on a heterogeneous collection of approximately 9.5 TB of text resources from Wikipedia and open-source government data, spanning structured and unstructured data. To ensure task quality, each sample is annotated by at least one Ph.D.-level expert. Each task requires long-horizon multi-hop reasoning with implicit intermediate steps: agents need to discover the correct documents and then compose evidence across sources to produce the answer. Experimental results on seven frontier LLMs demonstrate that LakeQA is challenging. For instance, GPT-5.2 achieves only an exact-match score of 18.37% on LakeQA. Overall, LakeQA provides a realistic testbed for developing LLM agents that can both find and analyze data in modern data lakes.

📄 PDF Abstract BibTeX arXiv:2606.10460

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

SANA: What Matters for QA Agents over Massive Data Lakes?

2026-06-11 · Austin Senna Wijaya, Jiaxiang Liu, Haonan Wang, Eugene Wu arxiv

Exploratory question answering (EQA) over data lakes requires an LLM agent to discover relevant sources, analyze retrieved data, and adapt its actions based on intermediate results. End-to-end accuracy alone cannot disti…

Question Answering

DataClawBench: An Agent Benchmark for Exploratory Real-World Financial Data Analysis

2026-05-04 · Qiaohong Zhang, Weihao Ye, Jialong Chen, Yi Luo 외 arxiv

Autonomous data analysis agents are increasingly expected to conduct exploratory analysis with limited human guidance about data. However, existing benchmarks typically evaluate such agents in prior-guided settings, prov…

Lookup or Exploratory: What is Your Search Intent?

2021-10-09 · Manoj K. Agarwal, Tezan Sahu

Search query specificity is broadly divided into two categories - Exploratory or Lookup. If a query specificity can be identified at the run time, it can be used to significantly improve the search results as well as qua…

SpecificityTriplet

ClusterTalk: Corpus Exploration Framework using Multi-Dimensional Exploratory Search

2024-12-19 · Ashish Chouhan, Saifeldin Mandour, Michael Gertz

Exploratory search of large text corpora is essential in domains like biomedical research, where large amounts of research literature are continuously generated. This paper presents ClusterTalk (The demo video and source…

Clustering

Twitter Dataset on the Russo-Ukrainian War

2022-04-07 · Alexander Shevtsov, Christos Tzagkarakis, Despoina Antonakaki, Polyvios Pratikakis 외

On 24 February 2022, Russia invaded Ukraine, also known now as Russo-Ukrainian War. We have initiated an ongoing dataset acquisition from Twitter API. Until the day this paper was written the dataset has reached the amou…

Sentiment Analysis