paper-with-me

홈 › Papers

SayNext-Bench: Why Do LLMs Struggle with Next-Utterance Anticipation?

2026-01-30 · Yueyi Yang, Haotian Liu, Fang Kang, Mengqi Zhang, Zheng Lian, Hao Tang, Haoyu Chen arxiv

We explore the use of large language models (LLMs) for next-utterance anticipation in human dialogue. Despite recent advances in LLMs demonstrating their ability to engage in natural conversations with users, we show that even leading models surprisingly struggle to anticipate a human speaker's next utterance. Instead, humans can readily anticipate forthcoming utterances based on multi-modal cues -- such as gestures, gaze, and emotional tone -- from the context. To systematically examine this gap, we propose SayNext-Bench, a benchmark evaluating MLLMs on anticipating context-conditioned responses across diverse real-world scenarios. To support it, we build SayNext-PC, a large-scale multimodal dialogue dataset, and carefully design a multi-level evaluation framework spanning lexical similarity, emotion-intention consistency, and LLM-based overall alignment. Building on this, we develop SayNext-Chat, a cognitively inspired dual-route MLLM that incorporates learnable priming tokens to fuse perceptual cues with anticipatory priors. Extensive experiments demonstrate that SayNext-Chat consistently outperforms state-of-the-art MLLMs across all evaluation levels, corroborated by user studies and LLM-as-Judge evaluations. Our results emphasize the (i) indispensable role of multimodal cues and (ii) active anticipatory processing as foundations of natural human interaction currently missing in MLLMs.

📄 PDF Abstract BibTeX arXiv:2602.00327

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LLM-Driven Preference Data Synthesis for Proactive Prediction of the Next User Utterance in Human-Machine Dialogue

2025-12-24 · Jinqiang Wang, Huansheng Ning, Jianguo Ding, Tao Zhu 외 arxiv

Proactively predicting a users next utterance in human-machine dialogue can streamline interaction and improve user experience. Existing commercial API-based solutions are subject to privacy concerns while deploying gene…

Benchmarking LLMs for Mimicking Child-Caregiver Language in Interaction

2024-12-12 · Jing Liu, Abdellah Fourtassi

LLMs can generate human-like dialogues, yet their ability to simulate early child-adult interactions remains largely unexplored. In this paper, we examined how effectively LLMs can capture the distinctive features of chi…

BenchmarkingDiversity

Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

2023-08-18 · Rishabh Bhardwaj, Soujanya Poria

Larger language models (LLMs) have taken the world by storm with their massive multi-tasking capabilities simply by optimizing over a next-word prediction objective. With the emergence of their properties and encoded kno…

MMLURed TeamingSafety AlignmentText Generation+1

Evaluating the Semantic Profiling Abilities of LLMs for Natural Language Utterances in Data Visualization

2024-07-08 · Hannah K. Bako, Arshnoor Bhutani, Xinyi Liu, Kwesi A. Cobbina 외

Automatically generating data visualizations in response to human utterances on datasets necessitates a deep semantic understanding of the data utterance, including implicit and explicit references to data attributes, vi…

Data Visualization

Diverse In-Context Example Selection After Decomposing Programs and Aligned Utterances Improves Semantic Parsing

2025-04-04 · Mayank Kothyari, Sunita Sarawagi, Soumen Chakrabarti, Gaurav Arora 외

LLMs are increasingly used as seq2seq translators from natural language utterances to structured programs, a process called semantic interpretation. Unlike atomic labels or token sequences, programs are naturally represe…

Semantic Parsing