paper-with-me

Papers

SelfTalk: A Self-Supervised Commutative Training Diagram to Comprehend 3D Talking Faces

2023-06-19 · Ziqiao Peng, Yihao Luo, Yue Shi, Hao Xu, Xiangyu Zhu, Jun He, Hongyan Liu, Zhaoxin Fan

Speech-driven 3D face animation technique, extending its applications to various multimedia fields. Previous research has generated promising realistic lip movements and facial expressions from audio signals. However, traditional regression models solely driven by data face several essential problems, such as difficulties in accessing precise labels and domain gaps between different modalities, leading to unsatisfactory results lacking precision and coherence. To enhance the visual accuracy of generated lip movement while reducing the dependence on labeled data, we propose a novel framework SelfTalk, by involving self-supervision in a cross-modals network system to learn 3D talking faces. The framework constructs a network system consisting of three modules: facial animator, speech recognizer, and lip-reading interpreter. The core of SelfTalk is a commutative training diagram that facilitates compatible features exchange among audio, text, and lip shape, enabling our models to learn the intricate connection between these factors. The proposed framework leverages the knowledge learned from the lip-reading interpreter to generate more plausible lip shapes. Extensive experiments and user studies demonstrate that our proposed approach achieves state-of-the-art performance both qualitatively and quantitatively. We recommend watching the supplementary video.

📄 PDF Abstract BibTeX arXiv:2306.10799

Code (1)

psyai-net/SelfTalk_release 공식 구현 pytorch

Tasks

3D Face AnimationLip Reading

Similar Papers 제목 키워드 기반

Self-Supervised Prime-Dual Networks for Few-Shot Image Classification

2021-09-29 · Wenming Cao, Qifan Liu, Guang Liu, Zhihai He

We construct a prime-dual network structure for few-shot learning which establishes a commutative relationship between the support set and the query set, as well as a new self- supervision constraint for highly effective…

Few-Shot Image ClassificationFew-Shot Learningimage-classificationImage Classification+1

Learning Disordered Topological Phases by Statistical Recovery of Symmetry

2017-09-18 · Nobuyuki Yoshioka, Yutaka Akagi, Hosho Katsura

In this letter, we apply the artificial neural network in a supervised manner to map out the quantum phase diagram of disordered topological superconductor in class DIII. Given the disorder that keeps the discrete symmet…

Meaning updating of density matrices

2020-01-03 · Bob Coecke, Konstantinos Meichanetzidis

The DisCoCat model of natural language meaning assigns meaning to a sentence given: (i) the meanings of its words, and, (ii) its grammatical structure. The recently introduced DisCoCirc model extends this to text consist…

Sentence

Training Transitive and Commutative Multimodal Transformers with LoReTTa

2023-05-23 · NeurIPS 2023 11 · Manuel Tran, Yashin Dicente Cid, Amal Lahiani, Fabian J. Theis 외

Training multimodal foundation models is challenging due to the limited availability of multimodal datasets. While many public datasets pair images with text, few combine images with audio or text with audio. Even rarer …

Triplet

CommVQ: Commutative Vector Quantization for KV Cache Compression

2025-06-23 · Junyan Li, Yang Zhang, Muhammad Yusuf Hassan, Talha Chafekar 외

Large Language Models (LLMs) are increasingly used in applications requiring long context lengths, but the key-value (KV) cache often becomes a memory bottleneck on GPUs as context grows. To address this, we propose Comm…

GPUGSM8KQuantization