paper-with-me

홈 › Papers

CogPortrait: Fine-Grained Eye-Region Control in Portrait Animation via Hierarchical Agent Planning

2026-05-27 · He Feng, Yongjia Ma, Donglin Di, Lei Fan, Tonghua Su arxiv

Portrait animation methods have achieved substantial visual quality and lip synchronization, but fine-grained manipulation of the eye region still faces a trade-off between input granularity and motion accuracy. Existing methods using emotion labels or coarse text prompts are insufficient for describing subtle ocular dynamics, whereas approaches based on Action Units or driving videos provide higher fidelity at the cost of a heavier input burden. These limitations are still restrictive for beyond-emotion states (e.g., thinking) and drowsiness. In light of the above, we propose CogPortrait, a two-stage framework that generates portrait animations from high-level labels. In the first stage, three chain-of-thought Multimodal Large Language Models (MLLMs) agents compile high-level labels into facial keypoints through temporal event planning, prototype retrieval, and composition from a real-behavior library, and semantic-physiological constraint enforcement. In the second stage, a DiT-based video generation backbone synthesizes the final animation conditioned on the keypoints, reference portrait, audio, and text prompt, enhanced by a dynamic classifier-free guidance strategy with eye-region-aware reweighting and KTO-based refinement for boundary cases. We further introduce the EMH benchmark covering diverse emotions and beyond-emotion categories with two AU-level metrics for evaluating fine-grained eye-region and head-motion control. Extensive experiments on HDTF and the EMH benchmark demonstrate that CogPortrait achieves more precise eye-region control than existing methods while maintaining supe- rior visual quality and identity consistency

📄 PDF Abstract BibTeX arXiv:2605.28056

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

ConsistentID: Portrait Generation with Multimodal Fine-Grained Identity Preserving

2024-04-25 · Jiehui Huang, Xiao Dong, Wenhui Song, Zheng Chong 외

Diffusion-based technologies have made significant strides, particularly in personalized and customized facialgeneration. However, existing methods face challenges in achieving high-fidelity and detailed identity (ID)con…

Diversity

LIA-X: Interpretable Latent Portrait Animator

2025-08-13 · Yaohui Wang, Di Yang, Xinyuan Chen, Francois Bremond 외 arxiv

We introduce LIA-X, a novel interpretable portrait animator designed to transfer facial dynamics from a driving video to a source portrait with fine-grained control. LIA-X is an autoencoder that models motion transfer as…

PortraitCraft: A Benchmark for Portrait Composition Understanding and Generation

2026-04-04 · Yuyang Sha, Zijie Lou, Youyun Tang, Xiaochao Qu 외 arxiv

Portrait composition plays a central role in portrait aesthetics and visual communication, yet existing datasets and benchmarks mainly focus on coarse aesthetic scoring, generic image aesthetics, or unconstrained portrai…

Visual Question Answering

3D-SSGAN: Lifting 2D Semantics for 3D-Aware Compositional Portrait Synthesis

2024-01-08 · Ruiqi Liu, Peng Zheng, Ye Wang, Rui Ma

Existing 3D-aware portrait synthesis methods can generate impressive high-quality images while preserving strong 3D consistency. However, most of them cannot support the fine-grained part-level control over synthesized i…

DisentanglementImage Generation

StyleAvatar: Real-time Photo-realistic Portrait Avatar from a Single Video

2023-05-01 · Lizhen Wang, Xiaochen Zhao, Jingxiang Sun, Yuxiang Zhang 외

Face reenactment methods attempt to restore and re-animate portrait videos as realistically as possible. Existing methods face a dilemma in quality versus controllability: 2D GAN-based methods achieve higher image qualit…

Face ReenactmentTranslationVideo Generation