paper-with-me

Papers

Synergizing Motion and Appearance: Multi-Scale Compensatory Codebooks for Talking Head Video Generation

2024-12-01 · CVPR 2025 1 · Shuling Zhao, Fa-Ting Hong, Xiaoshui Huang, Dan Xu

Talking head video generation aims to generate a realistic talking head video that preserves the person's identity from a source image and the motion from a driving video. Despite the promising progress made in the field, it remains a challenging and critical problem to generate videos with accurate poses and fine-grained facial details simultaneously. Essentially, facial motion is often highly complex to model precisely, and the one-shot source face image cannot provide sufficient appearance guidance during generation due to dynamic pose changes. To tackle the problem, we propose to jointly learn motion and appearance codebooks and perform multi-scale codebook compensation to effectively refine both the facial motion conditions and appearance features for talking face image decoding. Specifically, the designed multi-scale motion and appearance codebooks are learned simultaneously in a unified framework to store representative global facial motion flow and appearance patterns. Then, we present a novel multi-scale motion and appearance compensation module, which utilizes a transformer-based codebook retrieval strategy to query complementary information from the two codebooks for joint motion and appearance compensation. The entire process produces motion flows of greater flexibility and appearance features with fewer distortions across different scales, resulting in a high-quality talking head video generation framework. Extensive experiments on various benchmarks validate the effectiveness of our approach and demonstrate superior generation results from both qualitative and quantitative perspectives when compared to state-of-the-art competitors.

📄 PDF Abstract BibTeX arXiv:2412.00719

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Simoun: Synergizing Interactive Motion-appearance Understanding for Vision-based Reinforcement Learning

2023-01-01 · ICCV 2023 1 · Yangru Huang, Peixi Peng, Yifan Zhao, Yunpeng Zhai 외

Efficient motion and appearance modeling are critical for vision-based Reinforcement Learning (RL). However, existing methods struggle to reconcile motion and appearance information within the state representations l…

reinforcement-learningReinforcement Learning (RL)

Dynamical Non-compensatory Multidimensional IRT Model Using Variational Approximation

2026-09-09 · Hiroshi Tamano, Daichi Mochihashi arxiv

Multidimensional item response theory (MIRT) is a statistical test theory that precisely estimates multiple latent skills of learners from the responses in a test. Both compensatory and non-compensatory models have been …

Empowering Resampling Operation for Ultra-High-Definition Image Enhancement with Model-Aware Guidance

2024-01-01 · CVPR 2024 1 · Wei Yu, Jie Huang, Bing Li, Kaiwen Zheng 외

Image enhancement algorithms have made remarkable advancements in recent years but directly applying them to Ultra-high-definition (UHD) images presents intractable computational overheads. Therefore previous straigh…

Image Enhancement

Robust Facial Reactions Generation: An Emotion-Aware Framework with Modality Compensation

2024-07-22 · Guanyu Hu, Jie Wei, Siyang Song, Dimitrios Kollias 외

The objective of the Multiple Appropriate Facial Reaction Generation (MAFRG) task is to produce contextually appropriate and diverse listener facial behavioural responses based on the multimodal behavioural data of the c…

Misspecifying non-compensatory as compensatory IRT: analysis of estimated skills and variance

2025-07-21 · Hiroshi Tamano, Hideitsu Hino, Daichi Mochihashi arxiv

Multidimensional item response theory is a statistical test theory used to estimate the latent skills of learners and the difficulty levels of problems based on test results. Both compensatory and non-compensatory models…