paper-with-me

홈 › Papers

MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of Documents

2025-08-15 · Tomer Wolfson, Harsh Trivedi, Mor Geva, Yoav Goldberg, Dan Roth, Tushar Khot, Ashish Sabharwal, Reut Tsarfaty arxiv

Automated agents, powered by Large language models (LLMs), are emerging as the go-to tool for querying information. However, evaluation benchmarks for LLM agents rarely feature natural questions that are both information-seeking and genuinely time-consuming for humans. To address this gap we introduce MoNaCo, a benchmark of 1,315 natural and time-consuming questions that require dozens, and at times hundreds, of intermediate steps to solve -- far more than any existing QA benchmark. To build MoNaCo, we developed a decomposed annotation pipeline to elicit and manually answer real-world time-consuming questions at scale. Frontier LLMs evaluated on MoNaCo achieve at most 61.2% F1, hampered by low recall and hallucinations. Our results underscore the limitations of LLM-powered agents in handling the complexity and sheer breadth of real-world information-seeking tasks -- with MoNaCo providing an effective resource for tracking such progress. The MoNaCo benchmark, codebase, prompts and models predictions are all publicly available at: https://tomerwolgithub.github.io/monaco

📄 PDF Abstract BibTeX arXiv:2508.11133

Code (0)

등록된 구현이 없습니다.

Tasks

Natural Questions

Similar Papers 제목 키워드 기반

Reducing Hallucinations in Complex Question Answering using Simple Graph-based Retrieval-Augmented Generation (long version)

2026-06-04 · Christopher J. Wedge, Joshua Stutter, Danny Dixon, Jacek Cała arxiv

Large language models (LLMs) have fundamentally transformed the landscape of Natural Language Processing (NLP), although they remain susceptible to errors. Retrieval-augmented generation (RAG) systems have emerged as a c…

Complex Query AnsweringQuestion Answering

MonaCoBERT: Monotonic attention based ConvBERT for Knowledge Tracing

2022-08-19 · Unggi Lee, Yonghyun Park, Yujin Kim, Seongyune Choi 외

Knowledge tracing (KT) is a field of study that predicts the future performance of students based on prior performance datasets collected from educational applications such as intelligent tutoring systems, learning manag…

Knowledge TracingManagement

NaturalReasoning: Reasoning in the Wild with 2.8M Challenging Questions

2025-02-18 · Weizhe Yuan, Jane Yu, Song Jiang, Karthik Padthe 외

Scaling reasoning capabilities beyond traditional domains such as math and coding is hindered by the lack of diverse and high-quality questions. To overcome this limitation, we introduce a scalable approach for generatin…

Knowledge DistillationMath

Understanding Unnatural Questions Improves Reasoning over Text

2020-10-19 · COLING 2020 8 · Xiao-Yu Guo, Yuan-Fang Li, Gholamreza Haffari

Complex question answering (CQA) over raw text is a challenging task. A prominent approach to this task is based on the programmer-interpreter framework, where the programmer maps the question into a sequence of reasonin…

DiversityNatural QuestionsQuestion Answering

A deep cut into Split Federated Self-supervised Learning

2024-06-12 · Marcin Przewięźlikowski, Marcin Osial, Bartosz Zieliński, Marek Śmieja

Collaborative self-supervised learning has recently become feasible in highly distributed environments by dividing the network layers between client devices and a central server. However, state-of-the-art methods, such a…

Federated LearningSelf-Supervised Learning