paper-with-me

홈 › Papers

A$^2$-LLM: An End-to-end Conversational Audio Avatar Large Language Model

2026-02-04 · Xiaolin Hu, Hang Yuan, Xinzhu Sang, Binbin Yan, Zhou Yu, Cong Huang, Kai Chen arxiv

Developing expressive and responsive conversational digital humans is a cornerstone of next-generation human-computer interaction. While large language models (LLMs) have significantly enhanced dialogue capabilities, most current systems still rely on cascaded architectures that connect independent modules. These pipelines are often plagued by accumulated errors, high latency, and poor real-time performance. Lacking access to the underlying conversational context, these pipelines inherently prioritize rigid lip-sync over emotional depth. To address these challenges, we propose A$^2$-LLM, an end-to-end conversational audio avatar large language model that jointly reasons about language, audio prosody, and 3D facial motion within a unified framework. To facilitate training, we introduce FLAME-QA, a high-quality multimodal dataset designed to align semantic intent with expressive facial dynamics within a QA format. By leveraging deep semantic understanding, A$^2$-LLM generates emotionally rich facial movements beyond simple lip-synchronization. Experimental results demonstrate that our system achieves superior emotional expressiveness while maintaining real-time efficiency (500 ms latency, 0.7 RTF).

📄 PDF Abstract BibTeX arXiv:2602.04913

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens

2026-05-29 · Qingcheng Zhao, Yifang Pan, Karan Singh arxiv

Recent advances in Audio-LLMs like GPT-4o have ushered in an era of conversational interaction with language models. Conversational avatars however, still seem robotic in facial expression and conversational flow, in par…

Speech RecognitionSpeech SynthesisText Generation

Allo-AVA: A Large-Scale Multimodal Conversational AI Dataset for Allocentric Avatar Gesture Animation

2024-10-21 · Saif Punjwani, Larry Heck

The scarcity of high-quality, multimodal training data severely hinders the creation of lifelike avatar animations for conversational AI in virtual environments. Existing datasets often lack the intricate synchronization…

EchoAvatar: Real-time Generative Avatar Animation from Audio Streams

2026-05-27 · Bohong Chen, Yumeng Li, Yinglin Xu, Youyi Zheng 외 arxiv

Real-time synthesis of high-fidelity 3D character motion from audio is a pivotal component for next-generation interactive avatars and virtual assistants. However, most existing approaches are limited to offline processi…

Reinforcement Learning

ICo3D: An Interactive Conversational 3D Virtual Human

2026-01-19 · Richard Shaw, Youngkyoon Jang, Athanasios Papaioannou, Arthur Moreau 외 arxiv

This work presents Interactive Conversational 3D Virtual Human (ICo3D), a method for generating an interactive, conversational, and photorealistic 3D human avatar. Based on multi-view captures of a subject, we create an …

From Audio to Photoreal Embodiment: Synthesizing Humans in Conversations

2024-01-03 · CVPR 2024 1 · Evonne Ng, Javier Romero, Timur Bagautdinov, Shaojie Bai 외

We present a framework for generating full-bodied photorealistic avatars that gesture according to the conversational dynamics of a dyadic interaction. Given speech audio, we output multiple possibilities of gestural mot…

DiversityQuantization