When Contextual Inference Fails: Cancelability in Interactive Instruction Following
We investigate the separation of literal interpretation from contextual inference in a collaborative block-building tasks, where an agent must resolve underspecified instructions using context. We adapt an existing two-speaker psycholinguistic paradigm into an interactive benchmark called Build What I Mean (BWIM). This setup contrasts a pragmatically cooperative speaker with one who is only literally reliable. In BWIM, models face underspecified instructions and must choose between making a contextual inference or requesting clarification at a small communication cost. Evaluating several state-of-the-art LLMs, we find a clear dissociation between judgment and action. Although models successfully detect speaker unreliability in explicit confidence ratings, they fail to leverage this awareness when taking action. Instead of deploying efficient clarification strategies, models default to suboptimal behaviors. These include partner-blind over-clarification and question-averse guessing under uncertainty. BWIM provides a controlled environment to evaluate online partner adaptation and contextual reasoning in interactive settings.
Code (0)
등록된 구현이 없습니다.
Tasks
Instruction FollowingSimilar Papers 제목 키워드 기반
Greedy Algorithm for Structured Bandits: A Sharp Characterization of Asymptotic Success / Failure
We study the greedy (exploitation-only) algorithm in bandit problems with a known reward structure. We allow arbitrary finite reward structures, while prior work focused on a few specific ones. We fully characterize when…
Decision MakingMulti-Armed BanditsSequential Synthetic Difference in Differences
We propose the Sequential Synthetic Difference-in-Differences (Sequential SDiD) estimator for event studies with staggered treatment adoption, particularly when the parallel trends assumption fails. The method uses an it…
ImputationCan LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
The interactive use of large language models (LLMs) in AI assistants (at work, home, etc.) introduces a new set of inference-time privacy risks: LLMs are fed different types of information from multiple sources in their …
Privacy PreservingBeyond Parallel Tracking: Interactive Multi-Feature Fusion Drives Semantic Reconstruction from Non-invasive Brain Recordings
Continuous semantic reconstruction from non-invasive neural recordings remains limited by the representational mismatch between semantic feature spaces and neural coding patterns, which severely impedes cross-modal align…
Text GenerationJointly Learning Aspect-Focused and Inter-Aspect Relations with Graph Convolutional Networks for Aspect Sentiment Analysis
In this paper, we explore a novel solution of constructing a heterogeneous graph for each instance by leveraging aspect-focused and inter-aspect contextual dependencies for the specific aspect and propose an Interactive …
SentenceSentiment Analysis