paper-with-me

홈 › Papers

Mosaic of Modalities: A Comprehensive Benchmark for Multimodal Graph Learning

2024-06-24 · CVPR 2025 1 · Jing Zhu, YuHang Zhou, Shengyi Qian, Zhongmou He, Tong Zhao, Neil Shah, Danai Koutra

Graph machine learning has made significant strides in recent years, yet the integration of visual information with graph structure and its potential for improving performance in downstream tasks remains an underexplored area. To address this critical gap, we introduce the Multimodal Graph Benchmark (MM-GRAPH), a pioneering benchmark that incorporates both visual and textual information into graph learning tasks. MM-GRAPH extends beyond existing text-attributed graph benchmarks, offering a more comprehensive evaluation framework for multimodal graph learning Our benchmark comprises seven diverse datasets of varying scales (ranging from thousands to millions of edges), designed to assess algorithms across different tasks in real-world scenarios. These datasets feature rich multimodal node attributes, including visual data, which enables a more holistic evaluation of various graph learning frameworks in complex, multimodal environments. To support advancements in this emerging field, we provide an extensive empirical study on various graph learning frameworks when presented with features from multiple modalities, particularly emphasizing the impact of visual information. This study offers valuable insights into the challenges and opportunities of integrating visual data into graph learning.

📄 PDF Abstract BibTeX arXiv:2406.16321

Code (0)

등록된 구현이 없습니다.

Tasks

Graph Learning

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

TerraMesh: A Planetary Mosaic of Multimodal Earth Observation Data

2025-04-15 · Benedikt Blumenstiel, Paolo Fraccaro, Valerio Marsocci, Johannes Jakubik 외

Large-scale foundation models in Earth Observation can learn versatile, label-efficient representations by leveraging massive amounts of unlabeled data. However, existing public datasets are often limited in scale, geogr…

Earth Observation

MOSAIC: Modality-agnostic Spectral Alignment for Federated Image-level Weakly Supervised Tumor Segmentation under Client-specific Missing Modalities

2026-08-20 · Tarun Kumar Garg, Vaanathi Sundaresan arxiv

Trustworthy multimodal fusion in clinical settings requires handling incomplete and heterogeneous modality subsets across institutions, where privacy constraints prohibit centralized data sharing. Federated learning (FL)…

Federated LearningTumor Segmentation

MM-OpenFGL: A Comprehensive Benchmark for Multimodal Federated Graph Learning

2026-01-29 · Xunkai Li, Yuming Ai, Yinlin Zhu, Haodong Lu 외 arxiv

Multimodal-attributed graphs (MMAGs) provide a unified framework for modeling complex relational data by integrating heterogeneous modalities with graph structures. While centralized learning has shown promising performa…

Graph Learning

SpatialMosaic: A Multiview VLM Dataset for Partial Visibility

2025-12-29 · Kanghee Lee, Jungi Hong, Sion Lee, Injae Lee 외 arxiv

Recent progress in Multimodal Large Language Models (MLLMs) has enabled 3D scene understanding and spatial reasoning directly from multi-view images, without requiring explicit 3D reconstructions. Nevertheless, key chall…

Scene UnderstandingSpatial Reasoning

CLIMB: Data Foundations for Large Scale Multimodal Clinical Foundation Models

2025-03-09 · Wei Dai, Peilin Chen, Malinda Lu, Daniel Li 외

Recent advances in clinical AI have enabled remarkable progress across many clinical domains. However, existing benchmarks and models are primarily limited to a small set of modalities and tasks, which hinders the develo…