Beneath the Surface: Investigating LLMs' Capabilities for Communicating with Subtext
Human communication is fundamentally creative, and often makes use of subtext -- implied meaning that goes beyond the literal content of the text. Here, we systematically study whether language models can use subtext in communicative settings, and introduce four new evaluation suites to assess these capabilities. Our evaluation settings range from writing & interpreting allegories to playing multi-agent and multi-modal games inspired by the rules of board games like Dixit. We find that frontier models generally exhibit a strong bias towards overly literal, explicit communication, and thereby fail to account for nuanced constraints -- even the best performing models generate literal clues 60% of times in one of our environments -- Visual Allusions. However, we find that some models can sometimes make use of common ground with another party to help them communicate with subtext, achieving 30%-50% reduction in overly literal clues; but they struggle at inferring presence of a common ground when not explicitly stated. For allegory understanding, we find paratextual and persona conditions to significantly shift the interpretation of subtext. Overall, our work provides quantifiable measures for an inherently complex and subjective phenomenon like subtext and reveals many weaknesses and idiosyncrasies of current LLMs. We hope this research to inspire future work towards socially grounded creative communication and reasoning.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Connotation Lexicon: A Dash of Sentiment Beneath the Surface Meaning
Capacitive Sensor Based 2D Subsurface Imaging Technology for Non Destructive Evaluation of Building Surfaces
Understanding the underlying structure of building surfaces like walls and floors is essential when carrying out building maintenance and modification work. To facilitate such work, this paper introduces a capacitive sen…
Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs
Reasoning in large language models unfolds through diverse functional operations, such as problem formulation, goal decomposition, and deduction. Although these operations are explicitly distinguished in text, little is …
Holographic MIMO Communications exploiting the Orbital Angular Momentum
This study delves into the potential of harnessing the orbital angular momentum (OAM) property of electromagnetic waves in near-field and line-of-sight scenarios by utilizing large intelligent surfaces, in the context of…
Beneath the Surface: Unveiling Harmful Memes with Multimodal Reasoning Distilled from Large Language Models
The age of social media is rife with memes. Understanding and detecting harmful memes pose a significant challenge due to their implicit meaning that is not explicitly conveyed through the surface text and image. However…
Multimodal Reasoning