paper-with-me

Papers

Voice2Mesh: Cross-Modal 3D Face Model Generation from Voices

2021-04-21 · Cho-Ying Wu, Ke Xu, Chin-Cheng Hsu, Ulrich Neumann

This work focuses on the analysis that whether 3D face models can be learned from only the speech inputs of speakers. Previous works for cross-modal face synthesis study image generation from voices. However, image synthesis includes variations such as hairstyles, backgrounds, and facial textures, that are arguably irrelevant to voice or without direct studies to show correlations. We instead investigate the ability to reconstruct 3D faces to concentrate on only geometry, which is more physiologically grounded. We propose both the supervised learning and unsupervised learning frameworks. Especially we demonstrate how unsupervised learning is possible in the absence of a direct voice-to-3D-face dataset under limited availability of 3D face scans when the model is equipped with knowledge distillation. To evaluate the performance, we also propose several metrics to measure the geometric fitness of two 3D faces based on points, lines, and regions. We find that 3D face shapes can be reconstructed from voices. Experimental results suggest that 3D faces can be reconstructed from voices, and our method can improve the performance over the baseline. The best performance gains (15% - 20%) on ear-to-ear distance ratio metric (ER) coincides with the intuition that one can roughly envision whether a speaker's face is overall wider or thinner only from a person's voice. See our project page for codes and data.

📄 PDF Abstract BibTeX arXiv:2104.10299

Code (1)

choyingw/Voice2Mesh 공식 구현 pytorch

Tasks

Face GenerationFace ModelImage GenerationKnowledge Distillation

Similar Papers 제목 키워드 기반

Cross-Modal Perceptionist: Can Face Geometry be Gleaned from Voices?

2022-03-18 · CVPR 2022 1 · Cho-Ying Wu, Chin-Cheng Hsu, Ulrich Neumann

This work digs into a root question in human perception: can face geometry be gleaned from one's voices? Previous works that study this question only adopt developments in image synthesis and convert voices into face ima…

3D Face Modelling3D Face ReconstructionImage Generation

Cross-modal Face- and Voice-style Transfer

2023-02-27 · Naoya Takahashi, Mayank K. Singh, Yuki Mitsufuji

Image-to-image translation and voice conversion enable the generation of a new facial image and voice while maintaining some of the semantics such as a pose in an image and linguistic content in audio, respectively. They…

DiversityImage-to-Image TranslationOpen-Ended Question AnsweringStyle Transfer+2

MS2Mesh-XR: Multi-modal Sketch-to-Mesh Generation in XR Environments

2024-12-12 · Yuqi Tong, Yue Qiu, Ruiyang Li, Shi Qiu 외

We present MS2Mesh-XR, a novel multi-modal sketch-to-mesh generation pipeline that enables users to create realistic 3D objects in extended reality (XR) environments using hand-drawn sketches assisted by voice inputs. In…

Integrating Facial Generation into Full-Duplex Spoken Dialogue Systems

2026-06-20 · Jingjing Jiang, Atsumoto Ohashi, Ryuichiro Higashinaka arxiv

Full-duplex spoken dialogue models, such as Moshi, enable natural, low-latency voice conversations. However, they remain limited to the audio modality, lacking the facial expressions that are integral to human communicat…

The Hidden Dance of Phonemes and Visage: Unveiling the Enigmatic Link between Phonemes and Facial Features

2023-07-26 · Liao Qu, Xianwei Zou, Xiang Li, Yandong Wen 외

This work unveils the enigmatic link between phonemes and facial features. Traditional studies on voice-face correlations typically involve using a long period of voice input, including generating face images from voices…