paper-with-me

홈 › Papers

SAGE: Steering and Refining Dialog Generation with State-Action Augmentation

2025-03-04 · Yizhe Zhang, Navdeep Jaitly

Recent advances in large language models have demonstrated impressive capabilities in task-oriented applications, yet building emotionally intelligent chatbots that can engage in natural, strategic conversations remains a challenge. We present a novel approach called SAGE that uses latent variables to control long-horizon behavior in dialogue generation. At the core of our method is the State-Action Chain (SAC), which augments standard language model fine-tuning by introducing latent variables that encapsulate emotional states and conversational strategies between dialogue turns. During inference, these variables are generated before each response, enabling coarse-grained control over dialogue progression while maintaining natural interaction patterns. We also introduce a self-improvement pipeline that leverages dialogue tree search, LLM-based reward modeling, and targeted fine-tuning to optimize conversational trajectories. Our experimental results show that models trained with this approach demonstrate improved performance in emotional intelligence metrics while maintaining strong capabilities on LLM benchmarks. The discrete nature of our latent variables facilitates search-based strategies and provides a foundation for future applications of reinforcement learning to dialogue systems, where learning can occur at the state level rather than the token level.

📄 PDF Abstract BibTeX arXiv:2503.03040

Code (1)

apple/ml-sage-dialog-gen 공식 구현

Tasks

Dialogue GenerationEmotional Intelligence

Similar Papers 제목 키워드 기반

SDialog: A Python Toolkit for End-to-End Agent Building, User Simulation, Dialog Generation, and Evaluation

2025-12-09 · Sergio Burdisso, Séverin Baroudi, Yanis Labrak, David Grunert 외 arxiv

We present SDialog, an MIT-licensed open-source Python toolkit that unifies dialog generation, evaluation and mechanistic interpretability into a single end-to-end framework for building and analyzing LLM-based conversat…

Audio Generation

Emergent Crowds Dynamics from Language-Driven Multi-Agent Interactions

2025-08-20 · Yibo Liu, Liam Shatzel, Brandon Haworth, Teseo Schneider arxiv

Animating and simulating crowds using an agent-based approach is a well-established area where every agent in the crowd is individually controlled such that global human-like behaviour emerges. We observe that human navi…

TeleMem: Building Long-Term and Multimodal Memory for Agentic AI

2025-12-12 · Chunliang Chen, Ming Guan, Xiao Lin, Jiaxu Li 외 arxiv

Large language models (LLMs) excel at many NLP tasks but struggle to sustain long-term interactions due to limited attention over extended dialogue histories. Retrieval-augmented generation (RAG) mitigates this issue but…

Multimodal Reasoning

Enhancing Medical Dialogue Generation through Knowledge Refinement and Dynamic Prompt Adjustment

2025-06-12 · Hongda Sun, Jiaren Peng, Wenzhong Yang, Liang He 외

Medical dialogue systems (MDS) have emerged as crucial online platforms for enabling multi-turn, context-aware conversations with patients. However, existing MDS often struggle to (1) identify relevant medical knowledge …

Dialogue GenerationTriplet

Controlling Pretrained Language Generation Models by Learning to Focus

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Transformer-based language models, which are pretrained on large-scale unsupervised data and then finetuned on task-specific datasets, have become the dominant paradigm for various natural language generation tasks. The …

Abstractive Text SummarizationResponse GenerationText Generation