HUMBO: Bridging Response Generation and Facial Expression Synthesis
Spoken dialogue systems that assist users to solve complex tasks such as movie ticket booking have become an emerging research topic in artificial intelligence and natural language processing areas. With a well-designed dialogue system as an intelligent personal assistant, people can accomplish certain tasks more easily via natural language interactions. Today there are several virtual intelligent assistants in the market; however, most systems only focus on textual or vocal interaction. In this paper, we present HUMBO, a system aiming at generating dialogue responses and simultaneously synthesize corresponding visual expressions on faces for better multimodal interaction. HUMBO can (1) let users determine the appearances of virtual assistants by a single image, and (2) generate coherent emotional utterances and facial expressions on the user-provided image. This is not only a brand new research direction but more importantly, an ultimate step toward more human-like virtual assistants.
Code (0)
등록된 구현이 없습니다.
Tasks
Dialogue Generationmultimodal interactionResponse GenerationSpoken Dialogue SystemsSimilar Papers 제목 키워드 기반
HAPI: A Model for Learning Robot Facial Expressions from Human Preferences
Automatic robotic facial expression generation is crucial for human-robot interaction, as handcrafted methods based on fixed joint configurations often yield rigid and unnatural behaviors. Although recent automated techn…
Bayesian OptimizationFacial expression generationLearning-To-RankMagPlus: Bridging Micro-to-Regular Facial Expressions through Learnable Magnification
Facial micro-expressions are subtle and short-lived facial movements that provide important cues about genuine human emotions. However, modeling and generating them remains difficult because annotated micro-expression da…
Facial Expression Generation Aligned with Human Preference for Natural Dyadic Interaction
Achieving natural dyadic interaction requires generating facial expressions that are emotionally appropriate and socially aligned with human preference. Human feedback offers a compelling mechanism to guide such alignmen…
Reinforcement LearningExpCLIP: Bridging Text and Facial Expressions via Semantic Alignment
The objective of stylized speech-driven facial animation is to create animations that encapsulate specific emotional expressions. Existing methods often depend on pre-established emotional labels or facial expression tem…
Do Facial Expressions Predict Ad Sharing? A Large-Scale Observational Study
People often share news and information with their social connections, but why do some advertisements get shared more than others? A large-scale test examines whether facial responses predict sharing. Facial expressions …