paper-with-me

Papers

Kling-Avatar: Grounding Multimodal Instructions for Cascaded Long-Duration Avatar Animation Synthesis

2025-09-11 · Yikang Ding, Jiwen Liu, Wenyuan Zhang, Zekun Wang, Wentao Hu, Liyuan Cui, Mingming Lao, Yingchao Shao, Hui Liu, Xiaohan Li, Ming Chen, Xiaoqiang Liu, Yu-Shen Liu, Pengfei Wan arxiv

Recent advances in audio-driven avatar video generation have significantly enhanced audio-visual realism. However, existing methods treat instruction conditioning merely as low-level tracking driven by acoustic or visual cues, without modeling the communicative purpose conveyed by the instructions. This limitation compromises their narrative coherence and character expressiveness. To bridge this gap, we introduce Kling-Avatar, a novel cascaded framework that unifies multimodal instruction understanding with photorealistic portrait generation. Our approach adopts a two-stage pipeline. In the first stage, we design a multimodal large language model (MLLM) director that produces a blueprint video conditioned on diverse instruction signals, thereby governing high-level semantics such as character motion and emotions. In the second stage, guided by blueprint keyframes, we generate multiple sub-clips in parallel using a first-last frame strategy. This global-to-local framework preserves fine-grained details while faithfully encoding the high-level intent behind multimodal instructions. Our parallel architecture also enables fast and stable generation of long-duration videos, making it suitable for real-world applications such as digital human livestreaming and vlogging. To comprehensively evaluate our method, we construct a benchmark of 375 curated samples covering diverse instructions and challenging scenarios. Extensive experiments demonstrate that Kling-Avatar is capable of generating vivid, fluent, long-duration videos at up to 1080p and 48 fps, achieving superior performance in lip synchronization accuracy, emotion and dynamic expressiveness, instruction controllability, identity preservation, and cross-domain generalization. These results establish Kling-Avatar as a new benchmark for semantically grounded, high-fidelity audio-driven avatar synthesis.

📄 PDF Abstract BibTeX arXiv:2509.09595

Code (0)

등록된 구현이 없습니다.

Tasks

Domain GeneralizationVideo Generation

Similar Papers 제목 키워드 기반

AgileAvatar: Stylized 3D Avatar Creation via Cascaded Domain Bridging

2022-11-15 · Shen Sang, Tiancheng Zhi, Guoxian Song, Minghao Liu 외

Stylized 3D avatars have become increasingly prominent in our modern life. Creating these avatars manually usually involves laborious selection and adjustment of continuous and discrete parameters and is time-consuming f…

Self-Supervised Learning

KlingAvatar 2.0 Technical Report

2025-12-15 · Kling Team, Jialu Chen, Yikang Ding, Zhixue Fang 외 arxiv

Avatar video generation models have achieved remarkable progress in recent years. However, prior work exhibits limited efficiency in generating long-duration high-resolution videos, suffering from temporal drifting, qual…

Instruction FollowingVideo Generation

JoyStreamer: Unlocking Highly Expressive Avatars via Harmonized Text-Audio Conditioning

2026-01-31 · Ruikui Wang, Jinheng Feng, Lang Tian, Huaishao Luo 외 arxiv

Existing video avatar models have demonstrated impressive capabilities in scenarios such as talking, public speaking, and singing. However, the majority of these methods exhibit limited alignment with respect to text ins…

Kling-Omni Technical Report

2025-12-18 · Kling Team, Jialu Chen, Yuanzheng Ci, Xiangyu Du 외 arxiv

We present Kling-Omni, a generalist generative framework designed to synthesize high-fidelity videos directly from multimodal visual language inputs. Adopting an end-to-end perspective, Kling-Omni bridges the functional …

Instruction FollowingVideo Generation

Instruct-Video2Avatar: Video-to-Avatar Generation with Instructions

2023-06-05 · Shaoxu Li

We propose a method for synthesizing edited photo-realistic digital avatars with text instructions. Given a short monocular RGB video and text instructions, our method uses an image-conditioned diffusion model to edit on…