paper-with-me

홈 › Papers

From Transthoracic to Transesophageal: Cross-Modality Generation using LoRA Diffusion

2025-08-18 · Emmanuel Oladokun, Yuxuan Ou, Anna Novikova, Daria Kulikova, Sarina Thomas, Jurica Šprem, Vicente Grau arxiv

Deep diffusion models excel at realistic image synthesis but demand large training sets-an obstacle in data-scarce domains like transesophageal echocardiography (TEE). While synthetic augmentation has boosted performance in transthoracic echo (TTE), TEE remains critically underrepresented, limiting the reach of deep learning in this high-impact modality. We address this gap by adapting a TTE-trained, mask-conditioned diffusion backbone to TEE with only a limited number of new cases and adapters as small as $10^5$ parameters. Our pipeline combines Low-Rank Adaptation with MaskR$^2$, a lightweight remapping layer that aligns novel mask formats with the pretrained model's conditioning channels. This design lets users adapt models to new datasets with a different set of anatomical structures to the base model's original set. Through a targeted adaptation strategy, we find that adapting only MLP layers suffices for high-fidelity TEE synthesis. Finally, mixing less than 200 real TEE frames with our synthetic echoes improves the dice score on a multiclass segmentation task, particularly boosting performance on underrepresented right-heart structures. Our results demonstrate that (1) semantically controlled TEE images can be generated with low overhead, (2) MaskR$^2$ effectively transforms unseen mask formats into compatible formats without damaging downstream task performance, and (3) our method generates images that are effective for improving performance on a downstream task of multiclass segmentation.

📄 PDF Abstract BibTeX arXiv:2508.13077

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Transesophageal Echocardiography Generation using Anatomical Models

2024-10-09 · Emmanuel Oladokun, Musa Abdulkareem, Jurica Šprem, Vicente Grau

Through automation, deep learning (DL) can enhance the analysis of transesophageal echocardiography (TEE) images. However, DL methods require large amounts of high-quality data to produce accurate results, which is diffi…

Data AugmentationSemantic SegmentationTranslation

Pareto LoRA: Mitigating Modality Imbalance in Unified Multimodal Models via Pareto-Optimal Gradient Integration

2026-06-15 · Xiwen Wei, Mark Nutter, Madhusudhanan Srinivasan, Radu Marculescu arxiv

Unified multimodal models (UMMs) have recently emerged as a promising paradigm for integrating multimodal understanding and generation within a single autoregressive transformer. However, during multimodal instruction tu…

parameter-efficient fine-tuningmultimodal generationImage Generation

Lateralization LoRA: Interleaved Instruction Tuning with Modality-Specialized Adaptations

2024-07-04 · Zhiyang Xu, Minqian Liu, Ying Shen, Joy Rimchala 외

Recent advancements in Vision-Language Models (VLMs) have led to the development of Vision-Language Generalists (VLGs) capable of understanding and generating interleaved images and text. Despite these advances, VLGs sti…

AttributeImage Generation

An algorithm for Left Atrial Thrombi detection using Transesophageal Echocardiography

2015-08-24 · Jianrui Ding, Min Xian, H. D. Cheng, Yang Li 외

Transesophageal echocardiography (TEE) is widely used to detect left atrium (LA)/left atrial appendage (LAA) thrombi. In this paper, the local binary pattern variance (LBPV) features are extracted from region of interest…

Multiple Instance Learning

Harmonizing Visual Text Comprehension and Generation

2024-07-23 · Zhen Zhao, Jingqun Tang, Binghong Wu, Chunhui Lin 외

In this work, we present TextHarmony, a unified and versatile multimodal generative model proficient in comprehending and generating visual text. Simultaneously generating images and texts typically results in performanc…

multimodal generationReading ComprehensionText Generation