paper-with-me

홈 › Papers

Structure-Aware Fusion with Progressive Injection for Multimodal Molecular Representation Learning

2025-10-24 · Zihao Jing, Yan Sun, Yan Yi Li, Sugitha Janarthanan, Alana Deng, Pingzhao Hu arxiv

Multimodal molecular models often suffer from 3D conformer unreliability and modality collapse, limiting their robustness and generalization. We propose MuMo, a structured multimodal fusion framework that addresses these challenges in molecular representation through two key strategies. To reduce the instability of conformer-dependent fusion, we design a Structured Fusion Pipeline (SFP) that combines 2D topology and 3D geometry into a unified and stable structural prior. To mitigate modality collapse caused by naive fusion, we introduce a Progressive Injection (PI) mechanism that asymmetrically integrates this prior into the sequence stream, preserving modality-specific modeling while enabling cross-modal enrichment. Built on a state space backbone, MuMo supports long-range dependency modeling and robust information propagation. Across 29 benchmark tasks from Therapeutics Data Commons (TDC) and MoleculeNet, MuMo achieves an average improvement of 2.7% over the best-performing baseline on each task, ranking first on 22 of them, including a 27% improvement on the LD50 task. These results validate its robustness to 3D conformer noise and the effectiveness of multimodal fusion in molecular representation. The code is available at: github.com/selmiss/MuMo.

📄 PDF Abstract BibTeX arXiv:2510.23640

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning

Similar Papers 제목 키워드 기반

Let Geometry GUIDE: Layer-wise Unrolling of Geometric Priors in Multimodal LLMs

2026-04-07 · Chongyu Wang, Ting Huang, Chunyu Sun, Xinyu Ning 외 arxiv

Multimodal Large Language Models (MLLMs) have achieved remarkable progress in 2D visual tasks but still exhibit limited physical spatial awareness when processing real-world visual streams. Recently, feed-forward geometr…

Spatial Reasoning

Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers

2026-02-06 · Yuxuan Yao, Yuxuan Chen, Hui Li, Kaihui Cheng 외 arxiv

Multimodal Diffusion Transformers (MMDiTs) for text-to-image generation maintain separate text and image branches, with bidirectional information flow between text tokens and visual latents throughout denoising. In this …

Text-to-Image Generation

Towards Reliable Image Outpainting: Learning Structure-Aware Multimodal Fusion with Depth Guidance

2022-04-12 · Lei Zhang, Kang Liao, Chunyu Lin, Yao Zhao

Image outpainting technology generates visually plausible content regardless of authenticity, making it unreliable to be applied in practice. Thus, we propose a reliable image outpainting task, introducing the sparse dep…

Image Outpainting

Navigating the Emotion Tree: Hierarchical Hyperbolic RAG for Multimodal Emotion Recognition

2026-05-16 · Zeheng Wang, Bo Zhao, Yijie Zhu, Zhishu Liu 외 arxiv

Multimodal emotion recognition aims to integrate text, audio, and video sources to understand human affective states. Although multimodal large language models excel at multimodal reasoning, they typically treat emotion …

Multimodal Emotion RecognitionEmotion ClassificationMultimodal Reasoning

Modality-Guided Mixture of Graph Experts with Entropy-Triggered Routing for Multimodal Recommendation

2026-02-24 · Ji Dai, Quan Fang, Dengsheng Cai arxiv

Multimodal recommendation enhances ranking by integrating user-item interactions with item content, which is particularly effective under sparse feedback and long-tail distributions. However, multimodal signals are inher…

Multimodal RecommendationGraph Learning