paper-with-me

홈 › Papers

Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document Collections

2026-03-12 · Łukasz Borchmann, Jordy Van Landeghem, Michał Turski, Shreyansh Padarha, Ryan Othniel Kearns, Adam Mahdi, Niels Rogge, Clémentine Fourrier, Siwei Han, Huaxiu Yao, Artemis Llabrés, Yiming Xu, Dimosthenis Karatzas, Hao Zhang, Anupam Datta arxiv

Multimodal agents offer a promising path to automating complex document-intensive workflows. Yet, a critical question remains: do these agents demonstrate genuine strategic reasoning, or merely stochastic trial-and-error search? To address this, we introduce MADQA, a benchmark of 2,250 human-authored questions grounded in 800 heterogeneous PDF documents. Guided by Classical Test Theory, we design it to maximize discriminative power across varying levels of agentic abilities. To evaluate agentic behaviour, we introduce a novel evaluation protocol measuring the accuracy-effort trade-off. Using this framework, we show that while the best agents can match human searchers in raw accuracy, they succeed on largely different questions and rely on brute-force search to compensate for weak strategic planning. They fail to close the nearly 20% gap to oracle performance, persisting in unproductive loops. We release the dataset and evaluation harness to help facilitate the transition from brute-force retrieval to calibrated, efficient reasoning.

📄 PDF Abstract BibTeX arXiv:2603.12180

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Investigating Navigation Strategies in the Morris Water Maze through Deep Reinforcement Learning

2023-06-01 · Andrew Liu, Alla Borisyuk

Navigation is a complex skill with a long history of research in animals and humans. In this work, we simulate the Morris Water Maze in 2D to train deep reinforcement learning agents. We perform automatic classification …

Deep Reinforcement Learningreinforcement-learning

An LLM-Powered Cooperative Framework for Large-Scale Multi-Vehicle Navigation

2025-10-09 · Yuping Zhou, Siqi Lai, Jindong Han, Hao Liu arxiv

The rise of Internet of Vehicles (IoV) technologies is transforming traffic management from isolated control to a collective, multi-vehicle process. At the heart of this shift is multi-vehicle dynamic navigation, which r…

Reinforcement Learning

Balancing Two-Player Stochastic Games with Soft Q-Learning

2018-02-09 · Jordi Grau-Moya, Felix Leibfried, Haitham Bou-Ammar

Within the context of video games the notion of perfectly rational agents can be undesirable as it leads to uninteresting situations, where humans face tough adversarial decision makers. Current frameworks for stochastic…

Q-LearningReinforcement LearningReinforcement Learning (RL)Vocal Bursts Valence Prediction

Plan-MCTS: Plan Exploration for Action Exploitation in Web Navigation

2026-02-15 · Weiming Zhang, Jihong Wang, Jiamu Zhou, Qingyao Li 외 arxiv

Large Language Models (LLMs) have empowered autonomous agents to handle complex web navigation tasks. While recent studies integrate tree search to enhance long-horizon reasoning, applying these algorithms in web navigat…

Parallel Algorithm for Approximating Nash Equilibrium in Multiplayer Stochastic Games with Application to Naval Strategic Planning

2019-10-01 · Sam Ganzfried, Conner Laughlin, Charles Morefield

Many real-world domains contain multiple agents behaving strategically with probabilistic transitions and uncertain (potentially infinite) duration. Such settings can be modeled as stochastic games. While algorithms have…