Expression-aware video inpainting for HMD removal in XR applications
Head-mounted displays (HMDs) serve as indispensable devices for observing extended reality (XR) environments and virtual content. However, HMDs present an obstacle to external recording techniques as they block the upper face of the user. This limitation significantly affects social XR applications, specifically teleconferencing, where facial features and eye gaze information play a vital role in creating an immersive user experience. In this study, we propose a new network for expression-aware video inpainting for HMD removal (EVI-HRnet) based on generative adversarial networks (GANs). Our model effectively fills in missing information with regard to facial landmarks and a single occlusion-free reference image of the user. The framework and its components ensure the preservation of the user's identity across frames using the reference frame. To further improve the level of realism of the inpainted output, we introduce a novel facial expression recognition (FER) loss function for emotion preservation. Our results demonstrate the remarkable capability of the proposed framework to remove HMDs from facial videos while maintaining the subject's facial expression and identity. Moreover, the outputs exhibit temporal consistency along the inpainted frames. This lightweight framework presents a practical approach for HMD occlusion removal, with the potential to enhance various collaborative XR applications without the need for additional hardware.
Code (1)
Tasks
Facial Expression RecognitionFacial Expression Recognition (FER)Video InpaintingMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Towards Realistic Landmark-Guided Facial Video Inpainting Based on GANs
Facial video inpainting plays a crucial role in a wide range of applications, including but not limited to the removal of obstructions in video conferencing and telemedicine, enhancement of facial expression analysis, pr…
Facial Expression RecognitionFacial Expression Recognition (FER)Video InpaintingRaformer: Redundancy-Aware Transformer for Video Wire Inpainting
Video Wire Inpainting (VWI) is a prominent application in video inpainting, aimed at flawlessly removing wires in films or TV series, offering significant time and labor savings compared to manual frame-by-frame removal.…
Video InpaintingDAOVI: Distortion-Aware Omnidirectional Video Inpainting
Omnidirectional videos that capture the entire surroundings are employed in a variety of fields such as VR applications and remote sensing. However, their wide field of view often causes unwanted objects to appear in the…
Video InpaintingGeometry-Aware Video Inpainting for Joint Headset Occlusion Removal and Face Reconstruction in Social XR
Head-mounted displays (HMDs) are essential for experiencing extended reality (XR) environments and observing virtual content. However, they obscure the upper part of the user's face, complicating external video recording…
3D Face ReconstructionVideo InpaintingShort-Term and Long-Term Context Aggregation Network for Video Inpainting
Video inpainting aims to restore missing regions of a video and has many applications such as video editing and object removal. However, existing methods either suffer from inaccurate short-term context aggregation or ra…
Video EditingVideo Inpainting