paper-with-me

홈 › Papers

Spatial Knowledge Graph-Guided Multimodal Synthesis

2025-05-28 · Yida Xue, Zhen Bi, Jinnan Yang, Jungang Lou, Huajun Chen, Ningyu Zhang

Recent advances in multimodal large language models (MLLMs) have significantly enhanced their capabilities; however, their spatial perception abilities remain a notable limitation. To address this challenge, multimodal data synthesis offers a promising solution. Yet, ensuring that synthesized data adhere to spatial common sense is a non-trivial task. In this work, we introduce SKG2Data, a novel multimodal synthesis approach guided by spatial knowledge graphs, grounded in the concept of knowledge-to-data generation. SKG2Data automatically constructs a Spatial Knowledge Graph (SKG) to emulate human-like perception of spatial directions and distances, which is subsequently utilized to guide multimodal data synthesis. Extensive experiments demonstrate that data synthesized from diverse types of spatial knowledge, including direction and distance, not only enhance the spatial perception and reasoning abilities of MLLMs but also exhibit strong generalization capabilities. We hope that the idea of knowledge-based data synthesis can advance the development of spatial intelligence.

📄 PDF Abstract BibTeX arXiv:2505.22633

Code (0)

등록된 구현이 없습니다.

Tasks

Common Sense ReasoningKnowledge Graphs

Similar Papers 제목 키워드 기반

CalliMaster: Mastering Page-level Chinese Calligraphy via Layout-guided Spatial Planning

2026-03-12 · Tianshuo Xu, Tiantian Hong, Zhifei Chen, Fei Chao 외 arxiv

Page-level calligraphy synthesis requires balancing glyph precision with layout composition. Existing character models lack spatial context, while page-level methods often compromise brushwork detail. In this paper, we p…

Spatial-Temporal Graph Mamba for Music-Guided Dance Video Synthesis

2025-07-09 · Hao Tang, Ling Shao, Zhenyu Zhang, Luc Van Gool 외 arxiv

We propose a novel spatial-temporal graph Mamba (STG-Mamba) for the music-guided dance video synthesis task, i.e., to translate the input music to a dance video. STG-Mamba consists of two translation mappings: music-to-s…

MIRAGE: Knowledge Graph-Guided Cross-Cohort MRI Synthesis for Alzheimer's Disease Prediction

2026-03-02 · Guanchen Wu, Zhe Huang, Yuzhang Xie, Runze Yan 외 arxiv

Reliable Alzheimer's disease (AD) diagnosis increasingly relies on multimodal assessments combining structural Magnetic Resonance Imaging (MRI) and Electronic Health Records (EHR). However, deploying these models is bott…

MultiMat: Multimodal Program Synthesis for Procedural Materials using Large Multimodal Models

2025-09-26 · Jonas Belouadi, Tamy Boubekeur, Adrien Kaiser arxiv

Material node graphs are programs that generate the 2D channels of procedural materials, including geometry such as roughness and displacement maps, and reflectance such as albedo and conductivity maps. They are essentia…

Program Synthesis

HeartBeat: Towards Controllable Echocardiography Video Synthesis with Multimodal Conditions-Guided Diffusion Models

2024-06-20 · Xinrui Zhou, Yuhao Huang, Wufeng Xue, Haoran Dou 외

Echocardiography (ECHO) video is widely used for cardiac examination. In clinical, this procedure heavily relies on operator experience, which needs years of training and maybe the assistance of deep learning-based syste…