paper-with-me

Papers

IMUSE: IMU-based Facial Expression Capture

2024-02-03 · Youjia Wang, Yiwen Wu, Hengan Zhou, Hongyang Lin, Xingyue Peng, Yingwenqi Jiang, Yingsheng Zhu, Guanpeng Long, Yatu Zhang, Jingya Wang, Lan Xu, Jingyi Yu

For facial motion capture and analysis, the dominated solutions are generally based on visual cues, which cannot protect privacy and are vulnerable to occlusions. Inertial measurement units (IMUs) serve as potential rescues yet are mainly adopted for full-body motion capture. In this paper, we propose IMUSE to fill the gap, a novel path for facial expression capture using purely IMU signals, significantly distant from previous visual solutions.The key design in our IMUSE is a trilogy. We first design micro-IMUs to suit facial capture, companion with an anatomy-driven IMU placement scheme. Then, we contribute a novel IMU-ARKit dataset, which provides rich paired IMU/visual signals for diverse facial expressions and performances. Such unique multi-modality brings huge potential for future directions like IMU-based facial behavior analysis. Moreover, utilizing IMU-ARKit, we introduce a strong baseline approach to accurately predict facial blendshape parameters from purely IMU signals. The IMUSE framework empowers us to perform accurate facial capture in scenarios where visual methods falter and simultaneously safeguard user privacy. We conduct extensive experiments about both the IMU configuration and technical components to validate the effectiveness of our IMUSE approach. Notably, IMUSE enables various potential and novel applications, i.e., facial capture against occlusions or in a moving performance. We will release our dataset and implementations to enrich more possibilities of facial capture and analysis in our community.

📄 PDF Abstract BibTeX arXiv:2402.03944

Code (0)

등록된 구현이 없습니다.

Tasks

Anatomy

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Residual Connection 설명 없음
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Position-Wise Feed-Forward Layer 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…

Similar Papers 제목 키워드 기반

LiMuSE: Lightweight Multi-modal Speaker Extraction

2021-11-07 · Qinghua Liu, Yating Huang, Yunzhe Hao, Jiaming Xu 외

Multi-modal cues, including spatial information, facial expression and voiceprint, are introduced to the speech separation and speaker extraction tasks to serve as complementary information to achieve better performance.…

Model CompressionQuantizationSpeech Separation

SEREP: Semantic Facial Expression Representation for Robust In-the-Wild Capture and Retargeting

2024-12-18 · Arthur Josi, Luiz Gustavo Hafemann, Abdallah Dib, Emeline Got 외

Monocular facial performance capture in-the-wild is challenging due to varied capture conditions, face shapes, and expressions. Most current methods rely on linear 3D Morphable Models, which represent facial expressions …

Domain Adaptation

Capturing Complex Spatio-temporal Relations among Facial Muscles for Facial Expression Recognition

2013-06-01 · CVPR 2013 6 · Ziheng Wang, Shangfei Wang, Qiang Ji

Spatial-temporal relations among facial muscles carry crucial information about facial expressions yet have not been thoroughly exploited. One contributing factor for this is the limited ability of the current dynamic mo…

Facial Expression RecognitionFacial Expression Recognition (FER)

LEED: Label-Free Expression Editing via Disentanglement

2020-07-17 · ECCV 2020 8 · Rongliang Wu, Shijian Lu

Recent studies on facial expression editing have obtained very promising progress. On the other hand, existing methods face the constraint of requiring a large amount of expression labels which are often expensive and ti…

AttributeDisentanglement

High-Quality Real Time Facial Capture Based on Single Camera

2021-11-15 · Hongwei Xu, Leijia Dai, Jianxing Fu, Xiangyuan Wang 외

We propose a real time deep learning framework for video-based facial expression capture. Our process uses a high-end facial capture pipeline based on FACEGOOD to capture facial expression. We train a convolutional neura…

Vocal Bursts Intensity Prediction