paper-with-me

홈 › Papers

Don't Let Me Ask for It: LLMs Show Deficiencies in Active Multi-Turn Information Acquisition for Abductive Inference

2026-08-04 · Shahrukh Mohiuddin, Chalamalasetti Kranti, Sherzod Hakimov, David Schlangen arxiv

Abductive reasoning requires forming hypotheses that explain observed evidence and revising them as new evidence becomes available. While large language models (LLMs) are often evaluated on whether they solve abductive reasoning tasks correctly, less is known about how they acquire evidence, update their hypotheses, and decide when to stop. We introduce Alien Abduction game, an interactive probe for studying these behaviours under different interaction modes. The modes vary in whether evidence is provided upfront or across turns, and whether queries are selected by the model or examples are provided by the oracle. Across models, providing evidence upfront leads to higher success rates than distributing it across turns. In multi-turn settings, some models commit before using the available evidence, while others exhaust the turn budget without converging. Models also achieve higher success rates when examples are provided by the oracle than when they select their own queries, although their final hypotheses are more consistent with the evidence they selected. These findings suggest that models may form hypotheses that fit self-selected evidence without sufficiently distinguishing them from alternatives, and may struggle to validate and refine their hypotheses or determine when to stop.

📄 PDF Abstract BibTeX arXiv:2608.03388

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

One Battle After Another: Probing LLMs' Limits on Multi-Turn Instruction Following with a Benchmark Evolving Framework

2025-11-05 · Qi Jia, Ye Shen, Xiujie Song, Kaiwei Zhang 외 arxiv

Evaluating LLMs' instruction-following ability in multi-topic dialogues is essential yet challenging. Existing benchmarks are limited to a fixed number of turns, susceptible to saturation and failing to account for users…

Instruction Following

Agent-Based Detection and Resolution of Incompleteness and Ambiguity in Interactions with Large Language Models

2025-07-04 · Riya Naik, Ashwin Srinivasan, Swati Agarwal, Estrid He

Many of us now treat LLMs as modern-day oracles asking it almost any kind of question. However, consulting an LLM does not have to be a single turn activity. But long multi-turn interactions can get tedious if it is simp…

Question Answering

PlotCraft: Pushing the Limits of LLMs for Complex and Interactive Data Visualization

2025-10-15 · Jiajun Zhang, Jianke Zhang, Zeyu Cui, Jiaxi Yang 외 arxiv

Recent Large Language Models (LLMs) have demonstrated remarkable proficiency in code generation. However, their ability to create complex visualizations for scaled and structured data remains largely unevaluated and unde…

Code Generation

StructFlowBench: A Structured Flow Benchmark for Multi-turn Instruction Following

2025-02-20 · Jinnan Li, Jinzhe Li, Yue Wang, Yi Chang 외

Multi-turn instruction following capability constitutes a core competency of large language models (LLMs) in real-world applications. Existing evaluation benchmarks predominantly focus on fine-grained constraint satisfac…

Instruction Following

IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams

2026-05-26 · Jinzhao Li, Yinuo Chen, Wenxuan Song, Yijia Lei 외 arxiv

Recent multimodal large language models (MLLMs) achieve strong performance on reactive question answering, but real-world streaming assistants require proactive reasoning over continuous visual inputs. Existing benchmark…

Question Answering