paper-with-me

홈 › Papers

Spatio-Temporal Mixture-of-Modality-Experts Diffusion for Quantitative DCE-MRI Synthesis from Incomplete MR Sequences

2026-06-24 · Junhyeok Lee, Kyu Sung Choi arxiv

Quantitative maps from dynamic contrast-enhanced MRI (DCE-MRI) are essential for tumor assessment but are often unavailable due to contrast-agent risks and protocol variability. Prior methods predict these maps from other MRI modalities, yet most assume fixed, fully observed inputs and fail under realistic missingness. We present Spatio-Temporal Mixture-of-Modality-Experts (ST-MoME), a conditional diffusion framework that synthesizes 3D DCE parameter maps from diverse subsets of multimodal MRI. ST-MoME fuses modality-specific expert features through a spatio-temporal gating network that produces voxel-wise, timestep-dependent weights, forming a conditioning tensor that guides denoising. To preserve quantitative fidelity, ST-MoME performs diffusion directly in image space with 3D patch-based training and a Swin-based backbone. On a clinical brain-tumor cohort of 386 patients, we evaluate ST-MoME across 16 controlled modality-availability scenarios. It achieves the lowest mean Normalized Mean Square Error (NMSE) aggregated across all three DCE parameters, with leading performance on $v_p$ and $v_e$, competitive results on $K^{\mathrm{trans}}$, and the lowest reconstruction error within the clinically critical tumor region. A post-hoc analysis of the learned gating dynamics shows a structural-early, physiological-late fusion schedule consistent with clinical intuition.

📄 PDF Abstract BibTeX arXiv:2606.25535

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MoCA: Identity-Preserving Text-to-Video Generation via Mixture of Cross Attention

2025-08-05 · Qi Xie, Yongjia Ma, Donglin Di, Xuehao Gao 외 arxiv

Achieving ID-preserving text-to-video (T2V) generation remains challenging despite recent advances in diffusion-based models. Existing approaches often fail to capture fine-grained facial dynamics or maintain temporal id…

Text-to-Video Generation

Spatiotemporal-Untrammelled Mixture of Experts for Multi-Person Motion Prediction

2025-12-25 · Zheng Yin, Chengjian Li, Xiangbo Shu, Meiqi Cao 외 arxiv

Comprehensively and flexibly capturing the complex spatio-temporal dependencies of human motion is critical for multi-person motion prediction. Existing methods grapple with two primary limitations: i) Inflexible spatiot…

FineMoGen: Fine-Grained Spatio-Temporal Motion Generation and Editing

2023-12-22 · NeurIPS 2023 11 · Mingyuan Zhang, Huirong Li, Zhongang Cai, Jiawei Ren 외

Text-driven motion generation has achieved substantial progress with the emergence of diffusion models. However, existing methods still struggle to generate complex motion sequences that correspond to fine-grained descri…

Mixture-of-ExpertsMotion GenerationMotion Synthesis

Unified Multimodal Visual Tracking with Dual Mixture-of-Experts

2026-05-05 · Lingyi Hong, Jinglun Li, Xinyu Zhou, Kaixun Jiang 외 arxiv

Multimodal visual object tracking can be divided into to several kinds of tasks (e.g. RGB and RGB+X tracking), based on the input modality. Existing methods often train separate models for each modality or rely on pretra…

Visual Object TrackingModel CompressionVisual Tracking

SpaR3D-MoE: Adaptive 3D Spatial Reasoning from Sparse Views Meets Geometry-Inductive Mixture-of-Experts

2026-07-07 · Haida Feng, Hao Wei, Haolin Wang, Shiwei Li 외 arxiv

Recent Multimodal Large Language Models (MLLMs) struggle to bridge the representational gap between 2D semantic understanding and 3D spatial geometry. Existing 3D-aware models either rely on costly 3D-specific data or ut…

Spatial Reasoning