paper-with-me

Papers

Cross-modal Latent Space Alignment for Image to Avatar Translation

2023-01-01 · ICCV 2023 1 · Manuel Ladron De Guevara, Jose Echevarria, Yijun Li, Yannick Hold-Geoffroy, Cameron Smith, Daichi Ito

We present a novel method for automatic vectorized avatar generation from a single portrait image. Most existing approaches that create avatars rely on image-to-image translation methods, which present some limitations when applied to 3D rendering, animation, or video. Instead, we leverage modality-specific autoencoders trained on large-scale unpaired portraits and parametric avatars, and then learn a mapping between both modalities via an alignment module trained on a significantly smaller amount of data. The resulting cross-modal latent space preserves facial identity, producing more visually appealing and higher fidelity avatars than previous methods, as supported by our quantitative and qualitative evaluations. Moreover, our method's virtue of being resolution-independent makes it highly versatile and applicable in a wide range of settings.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Image-to-Image TranslationTranslation

Similar Papers 제목 키워드 기반

Paired Cross-Modal Data Augmentation for Fine-Grained Image-to-Text Retrieval

2022-07-29 · Hao Wang, Guosheng Lin, Steven C. H. Hoi, Chunyan Miao

This paper investigates an open research problem of generating text-image pairs to improve the training of fine-grained image-to-text cross-modal retrieval task, and proposes a novel framework for paired data augmentatio…

Cross-Modal RetrievalData AugmentationImage to textImage-to-Text Retrieval+2

LatentUMM: Dual Latent Alignment for Unified Multimodal Models

2026-05-18 · Yinyi Luo, Wenwen Wang, Hayes Bai, Marios Savvides 외 arxiv

Unified multimodal models (UMMs) achieve strong performance in both understanding and generation by learning a shared latent space, yet they often exhibit functional inconsistency between these two capabilities. We obser…

Closing the gap in multimodal medical representation alignment

2026-02-23 · Eleonora Grassucci, Giordano Cicchetti, Danilo Comminiello arxiv

In multimodal learning, CLIP has emerged as the de-facto approach for mapping different modalities into a shared latent space by bringing semantically similar representations closer while pushing apart dissimilar ones. H…

Cross-Modal RetrievalImage Captioning

Learning Aligned Cross-Modal Representation for Generalized Zero-Shot Classification

2021-12-24 · Zhiyu Fang, Xiaobin Zhu, Chun Yang, Zheng Han 외

Learning a common latent embedding by aligning the latent spaces of cross-modal autoencoders is an effective strategy for Generalized Zero-Shot Classification (GZSC). However, due to the lack of fine-grained instance-wis…

Classificationzero-shot-classificationZero-Shot Learning

Latent Space Consistency for Sparse-View CT Reconstruction

2025-07-15 · Duoyou Chen, Yunqing Chen, Can Zhang, Zhou Wang 외

Computed Tomography (CT) is a widely utilized imaging modality in clinical settings. Using densely acquired rotational X-ray arrays, CT can capture 3D spatial features. However, it is confronted with challenged such as s…

Computed Tomography (CT)Contrastive LearningCT ReconstructionImage Generation+1