paper-with-me

Papers

QuantumCanvas: A Multimodal Benchmark for Visual Learning of Atomic Interactions

2025-12-01 · Can Polat, Erchin Serpedin, Mustafa Kurban, Hasan Kurban arxiv

Despite rapid advances in molecular and materials machine learning, most models still lack physical transferability: they fit correlations across whole molecules or crystals rather than learning the quantum interactions between atomic pairs. Yet bonding, charge redistribution, orbital hybridization, and electronic coupling all emerge from these two-body interactions that define local quantum fields in many-body systems. We introduce QuantumCanvas, a large-scale multimodal benchmark that treats two-body quantum systems as foundational units of matter. The dataset spans 2,850 element-element pairs, each annotated with 18 electronic, thermodynamic, and geometric properties and paired with ten-channel image representations derived from l- and m-resolved orbital densities, angular field transforms, co-occupancy maps, and charge-density projections. These physically grounded images encode spatial, angular, and electrostatic symmetries without explicit coordinates, providing an interpretable visual modality for quantum learning. Benchmarking eight architectures across 18 targets, we report mean absolute errors of 0.201 eV on energy gap using GATv2, 0.265 eV on HOMO and 0.274 eV on LUMO using EGNN. For energy-related quantities, DimeNet attains 2.27 eV total-energy MAE and 0.132 eV repulsive-energy MAE, while a multimodal fusion model achieves a 2.15 eV Mermin free-energy MAE. Pretraining on QuantumCanvas further improves convergence stability and generalization when fine-tuned on larger datasets such as QM9, MD17, and CrysMTM. By unifying orbital physics with vision-based representation learning, QuantumCanvas provides a principled and interpretable basis for learning transferable quantum interactions through coupled visual and numerical modalities. Dataset and model implementations are available at https://github.com/KurbanIntelligenceLab/QuantumCanvas.

📄 PDF Abstract BibTeX arXiv:2512.01519

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning

Similar Papers 제목 키워드 기반

WorldVQA: Measuring Atomic World Knowledge in Multimodal Large Language Models

2026-01-28 · Runjie Zhou, Youbo Shao, Haoyu Lu, Bowei Xing 외 arxiv

We introduce WorldVQA, a benchmark designed to evaluate the atomic visual world knowledge of Multimodal Large Language Models (MLLMs). Unlike current evaluations, which often conflate visual knowledge retrieval with reas…

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models

2025-05-26 · Hyunsik Chae, Seungwoo Yoon, Jaden Park, Chloe Yewon Chun 외

Recent Vision-Language Models (VLMs) have demonstrated impressive multimodal comprehension and reasoning capabilities, yet they often struggle with trivially simple visual tasks. In this work, we focus on the domain of b…

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

2026-07-27 · Zichao Lin, Yifeng Xie, Bowen Qu, Haiming Wang 외 hf

We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: hol…

MUREL: Multimodal Relational Reasoning for Visual Question Answering

2019-02-25 · CVPR 2019 6 · Remi Cadene, Hedi Ben-Younes, Matthieu Cord, Nicolas Thome

Multimodal attentional networks are currently state-of-the-art models for Visual Question Answering (VQA) tasks involving real images. Although attention allows to focus on the visual content relevant to the question, th…

Relational ReasoningVisual Question AnsweringVisual Question Answering (VQA)

AnatomiX, an Anatomy-Aware Grounded Multimodal Large Language Model for Chest X-Ray Interpretation

2026-01-06 · Anees Ur Rehman Hashmi, Numan Saeed, Christoph Lippert arxiv

Multimodal medical large language models have shown substantial progress in chest X-ray interpretation but continue to face challenges in spatial reasoning and anatomical understanding. Although existing grounding techni…

Visual Question AnsweringSpatial ReasoningPhrase Grounding