Talking to a Know-It-All GPT or a Second-Guesser Claude? How Repair reveals unreliable Multi-Turn Behavior in LLMs
Repair, an important resource for resolving trouble in human-human conversation, remains underexplored in human-LLM interaction. In this study, we investigate how LLMs engage in the interactive process of repair in multi-turn dialogues around solvable and unsolvable math questions. We examine whether models initiate repair themselves and how they respond to user-initiated repair. Our results show strong differences across models: reactions range from being almost completely resistant to (appropriate) repair attempts to being highly susceptible and easily manipulated. We further demonstrate that once conversations extend beyond a single turn, model behavior becomes more distinctive and less predictable across systems. Overall, our findings indicate that each tested LLM exhibits its own characteristic form of unreliability in the context of repair.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
DAVD-Net: Deep Audio-Aided Video Decompression of Talking Heads
Close-up talking heads are among the most common and salient object in video contents, such as face-to-face conversations in social media, teleconferences, news broadcasting, talk shows, etc. Due to the high sensitivity …
Video CompressionVideo ReconstructionFaithful Autoformalization via Roundtrip Verification and Repair
When an LLM formalizes natural language, how do we know the output is faithful? We propose a roundtrip verification approach which does not require ground-truth annotations: formalize a statement, translate the result ba…
Learning Better Visual Dialog Agents with Pretrained Visual-Linguistic Representation
GuessWhat?! is a two-player visual dialog guessing game where player A asks a sequence of yes/no questions (Questioner) and makes a final guess (Guesser) about a target object in an image, based on answers from player B …
Referring ExpressionReferring Expression ComprehensionVisual DialogVisual GroundingRepoRepair: Leveraging Code Documentation for Repository-Level Automated Program Repair
Automated program repair (APR) struggles to scale from isolated functions to full repositories, as it demands a global, task-aware understanding to locate necessary changes. Current methods, limited by context and relian…
Program RepairA Zero-Shot Classification Approach for a Word-Guessing Challenge
The Taboo Challenge competition, a task based on the well-known Taboo game, has been proposed to stimulate research in the AI field. The challenge requires building systems able to comprehend the implied inferences betwe…
ClassificationLanguage ModelingLanguage Modellingzero-shot-classification+1