paper-with-me

Papers

Learning to Generate Context-Sensitive Backchannel Smiles for Embodied AI Agents with Applications in Mental Health Dialogues

2024-02-13 · Maneesh Bilalpur, Mert Inan, Dorsa Zeinali, Jeffrey F. Cohn, Malihe Alikhani

Addressing the critical shortage of mental health resources for effective screening, diagnosis, and treatment remains a significant challenge. This scarcity underscores the need for innovative solutions, particularly in enhancing the accessibility and efficacy of therapeutic support. Embodied agents with advanced interactive capabilities emerge as a promising and cost-effective supplement to traditional caregiving methods. Crucial to these agents' effectiveness is their ability to simulate non-verbal behaviors, like backchannels, that are pivotal in establishing rapport and understanding in therapeutic contexts but remain under-explored. To improve the rapport-building capabilities of embodied agents we annotated backchannel smiles in videos of intimate face-to-face conversations over topics such as mental health, illness, and relationships. We hypothesized that both speaker and listener behaviors affect the duration and intensity of backchannel smiles. Using cues from speech prosody and language along with the demographics of the speaker and listener, we found them to contain significant predictors of the intensity of backchannel smiles. Based on our findings, we introduce backchannel smile production in embodied agents as a generation problem. Our attention-based generative model suggests that listener information offers performance improvements over the baseline speaker-centric generation approach. Conditioned generation using the significant predictors of smile intensity provides statistically significant improvements in empirical measures of generation quality. Our user study by transferring generated smiles to an embodied agent suggests that agent with backchannel smiles is perceived to be more human-like and is an attractive alternative for non-personal conversations over agent without backchannel smiles.

📄 PDF Abstract BibTeX arXiv:2402.08837

Code (1)

bmaneesh/generating-context-sensitive-backchannel-smiles 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Aligning Backchannel and Dialogue Context Representations via Contrastive LLM Fine-Tuning

2026-04-17 · Livia Qian, Gabriel Skantze arxiv

Backchannels (e.g., `yeah', `mhm', and `right') are short, non-interruptive feedback signals whose lexical form and prosody jointly convey pragmatic meaning. While prior computational research has largely focused on pred…

Robotic Backchanneling in Online Conversation Facilitation: A Cross-Generational Study

2024-09-25 · Sota Kobuki, Katie Seaborn, Seiki Tokunaga, Kosuke Fukumori 외

Japan faces many challenges related to its aging society, including increasing rates of cognitive decline in the population and a shortage of caregivers. Efforts have begun to explore solutions using artificial intellige…

CMIS-Net: A Cascaded Multi-Scale Individual Standardization Network for Backchannel Agreement Estimation

2025-10-15 · Yuxuan Huang, Kangzhong Wang, Eugene Yujun Fu, Grace Ngai 외 arxiv

Backchannels are subtle listener responses, such as nods, smiles, or short verbal cues like "yes" or "uh-huh," which convey understanding and agreement in conversations. These signals provide feedback to speakers, improv…

Emotion RecognitionData Augmentation

ECHO: A Matched-Contrast Benchmark for Context-Sensitive Turn-Taking in Full-Duplex Dialogue

2026-09-15 · Shuofeng Zhao, Hongwei Cai, Wenke Fan, Qingxiang Guo 외 arxiv

Full-duplex spoken dialogue systems must distinguish interruptions that require yielding the floor from backchannels that permit continued speaking. Existing benchmarks typically evaluate events independently and may the…

Multilingual and Continuous Backchannel Prediction: A Cross-lingual Study

2025-12-16 · Koji Inoue, Mikey Elmers, Yahui Fu, Zi Haur Pang 외 arxiv

We present a multilingual, continuous backchannel prediction model for Japanese, English, and Chinese, and use it to investigate cross-linguistic timing behavior. The model is Transformer-based and operates at the frame …