paper-with-me

홈 › Papers

Latent Translation: Crossing Modalities by Bridging Generative Models

2019-02-21 · Yingtao Tian, Jesse Engel

End-to-end optimization has achieved state-of-the-art performance on many specific problems, but there is no straight-forward way to combine pretrained models for new problems. Here, we explore improving modularity by learning a post-hoc interface between two existing models to solve a new task. Specifically, we take inspiration from neural machine translation, and cast the challenging problem of cross-modal domain transfer as unsupervised translation between the latent spaces of pretrained deep generative models. By abstracting away the data representation, we demonstrate that it is possible to transfer across different modalities (e.g., image-to-audio) and even different types of generative models (e.g., VAE-to-GAN). We compare to state-of-the-art techniques and find that a straight-forward variational autoencoder is able to best bridge the two generative models through learning a shared latent space. We can further impose supervised alignment of attributes in both domains with a classifier in the shared latent space. Through qualitative and quantitative evaluations, we demonstrate that locality and semantic alignment are preserved through the transfer process, as indicated by high transfer accuracies and smooth interpolations within a class. Finally, we show this modular structure speeds up training of new interface models by several orders of magnitude by decoupling it from expensive retraining of base generative models.

📄 PDF Abstract BibTeX arXiv:1902.08261

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTranslation

Methods 이 논문이 사용한 방법론

Solana Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

Latent Domain Transfer: Crossing modalities with Bridging Autoencoders

2019-05-01 · ICLR 2019 5 · Yingtao Tian, Jesse Engel

Domain transfer is a exciting and challenging branch of machine learning because models must learn to smoothly transfer between domains, preserving local variations and capturing many aspects of variation without labels.…

Generative Adversarial Network

CyIN: Cyclic Informative Latent Space for Bridging Complete and Incomplete Multimodal Learning

2026-02-04 · Ronghao Lin, Qiaolin He, Sijie Mai, Ying Zeng 외 arxiv

Multimodal machine learning, mimicking the human brain's ability to integrate various modalities has seen rapid growth. Most previous multimodal models are trained on perfectly paired multimodal input to reach optimal pe…

FlowBind: Efficient Any-to-Any Generation with Bidirectional Flows

2025-12-17 · Yeonwoo Cha, Semin Kim, Jinhyeon Kwon, Seunghoon Hong arxiv

Any-to-any generation seeks to translate between arbitrary subsets of modalities, enabling flexible cross-modal synthesis. Despite recent success, existing flow-based approaches are challenged by their inefficiency, as t…

Supervised Image Translation from Visible to Infrared Domain for Object Detection

2024-08-03 · Prahlad Anand, Qiranul Saadiyean, Aniruddh Sikdar, Nalini N 외

This study aims to learn a translation from visible to infrared imagery, bridging the domain gap between the two modalities so as to improve accuracy on downstream tasks including object detection. Previous approaches at…

Generative Adversarial NetworkObjectobject-detectionObject Detection+2

Bridging Modalities and Tasks: A Unified Hierarchical ViT for SAR-to-Optical Translation and Semantic Segmentation

2026-09-04 · Siyuan Liu, Xuze Zhang, Yongshun Wang, Licong Pan 외 arxiv

Synthetic Aperture Radar (SAR) images have all-weather, day-and-night observation capabilities. However, compared with optical images, their speckle noise and non-intuitive scattering mechanism limit the interpretability…

Semantic Segmentation