paper-with-me

Papers

Unified Cross-Modal Image Synthesis with Hierarchical Mixture of Product-of-Experts

2024-10-25 · Reuben Dorent, Nazim Haouchine, Alexandra Golby, Sarah Frisken, Tina Kapur, William Wells

We propose a deep mixture of multimodal hierarchical variational auto-encoders called MMHVAE that synthesizes missing images from observed images in different modalities. MMHVAE's design focuses on tackling four challenges: (i) creating a complex latent representation of multimodal data to generate high-resolution images; (ii) encouraging the variational distributions to estimate the missing information needed for cross-modal image synthesis; (iii) learning to fuse multimodal information in the context of missing data; (iv) leveraging dataset-level information to handle incomplete data sets at training time. Extensive experiments are performed on the challenging problem of pre-operative brain multi-parametric magnetic resonance and intra-operative ultrasound imaging.

📄 PDF Abstract BibTeX arXiv:2410.19378

Code (1)

reubendo/mhvae 공식 구현 pytorch

Tasks

Image Generation

Similar Papers 제목 키워드 기반

Unified Brain MR-Ultrasound Synthesis using Multi-Modal Hierarchical Representations

2023-09-15 · Reuben Dorent, Nazim Haouchine, Fryderyk Kögl, Samuel Joutard 외

We introduce MHVAE, a deep hierarchical variational auto-encoder (VAE) that synthesizes missing images from various modalities. Extending multi-modal VAEs with a hierarchical latent structure, we introduce a probabilisti…

Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis

2026-03-31 · Shuang Chen, Quanxin Shou, Hangting Chen, Yucheng Zhou 외 arxiv

Unified multimodal models provide a natural and promising architecture for understanding diverse and complex real-world knowledge while generating high-quality images. However, they still rely primarily on frozen paramet…

Image Generation

ARMANI: Part-level Garment-Text Alignment for Unified Cross-Modal Fashion Design

2022-08-11 · Xujie Zhang, Yu Sha, Michael C. Kampffmeyer, Zhenyu Xie 외

Cross-modal fashion image synthesis has emerged as one of the most promising directions in the generation domain due to the vast untapped potential of incorporating multiple modalities and the wide range of fashion image…

Image Generation

M3D-GAN: Multi-Modal Multi-Domain Translation with Universal Attention

2019-07-09 · Shuang Ma, Daniel McDuff, Yale Song

Generative adversarial networks have led to significant advances in cross-modal/domain translation. However, typically these networks are designed for a specific task (e.g., dialogue generation or image synthesis, but no…

Dialogue GenerationImage CaptioningImage GenerationMachine Translation+5

Any-to-All MRI Synthesis: A Unified Foundation Model for Nasopharyngeal Carcinoma and Its Downstream Applications

2026-02-09 · Yao Pu, Yiming Shi, Zhenxi Zhang, Peixin Yu 외 arxiv

Magnetic resonance imaging (MRI) is essential for nasopharyngeal carcinoma (NPC) radiotherapy (RT), but practical constraints, such as patient discomfort, long scan times, and high costs often lead to incomplete modaliti…

Representation Learning