paper-with-me

홈 › Papers

Latents2Segments: Disentangling the Latent Space of Generative Models for Semantic Segmentation of Face Images

2022-07-05 · Snehal Singh Tomar, A. N. Rajagopalan

With the advent of an increasing number of Augmented and Virtual Reality applications that aim to perform meaningful and controlled style edits on images of human faces, the impetus for the task of parsing face images to produce accurate and fine-grained semantic segmentation maps is more than ever before. Few State of the Art (SOTA) methods which solve this problem, do so by incorporating priors with respect to facial structure or other face attributes such as expression and pose in their deep classifier architecture. Our endeavour in this work is to do away with the priors and complex pre-processing operations required by SOTA multi-class face segmentation models by reframing this operation as a downstream task post infusion of disentanglement with respect to facial semantic regions of interest (ROIs) in the latent space of a Generative Autoencoder model. We present results for our model's performance on the CelebAMask-HQ and HELEN datasets. The encoded latent space of our model achieves significantly higher disentanglement with respect to semantic ROIs than that of other SOTA works. Moreover, it achieves a 13% faster inference rate and comparable accuracy with respect to the publicly available SOTA for the downstream task of semantic segmentation of face images.

📄 PDF Abstract BibTeX arXiv:2207.01871

Code (0)

등록된 구현이 없습니다.

Tasks

DisentanglementSegmentationSemantic Segmentation

Similar Papers 제목 키워드 기반

Disentangling Shared and Private Neural Dynamics with SPIRE: A Latent Modeling Framework for Deep Brain Stimulation

2025-10-28 · Rahil Soroushmojdehi, Sina Javadzadeh, Mehrnaz Asadi, Terence D. Sanger arxiv

Disentangling shared network-level dynamics from region-specific activity is a central challenge in modeling multi-region neural data. We introduce SPIRE (Shared-Private Inter-Regional Encoder), a deep multi-encoder auto…

latentSplat: Autoencoding Variational Gaussians for Fast Generalizable 3D Reconstruction

2024-03-24 · Christopher Wewer, Kevin Raj, Eddy Ilg, Bernt Schiele 외

We present latentSplat, a method to predict semantic Gaussians in a 3D latent space that can be splatted and decoded by a light-weight generative 2D architecture. Existing methods for generalizable 3D reconstruction eith…

3D ReconstructionDecoder

Exploring and Exploiting Hubness Priors for High-Quality GAN Latent Sampling

2022-06-13 · Yuanbang Liang, Jing Wu, Yu-Kun Lai, Yipeng Qin

Despite the extensive studies on Generative Adversarial Networks (GANs), how to reliably sample high-quality images from their latent spaces remains an under-explored topic. In this paper, we propose a novel GAN latent s…

Vocal Bursts Intensity Prediction

LatentSpeech: Latent Diffusion for Text-To-Speech Generation

2024-12-11 · Haowei Lou, Helen Paik, Pari Delir Haghighi, Wen Hu 외

Diffusion-based Generative AI gains significant attention for its superior performance over other generative techniques like Generative Adversarial Networks and Variational Autoencoders. While it has achieved notable adv…

text-to-speechText to Speech

PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion

2026-05-22 · Yifan Lu, Qi Wu, Jay Zhangjie Wu, Zian Wang 외 arxiv

Most practical high-resolution text-to-image systems, including latent diffusion and autoregressive models, perform generation in a compact latent space, and a decoder maps the generated latents back to pixels. Yet the l…