paper-with-me

Papers

MatBind: A Shared Embedding Space for Multimodal Materials Characterization

2026-07-09 · Le Yang, Anoop K. Chandran, Jona Östreicher, Evgenii Sovetkin, Adrian Mirza, Sebastien Bompas, Bashir Kazimi, Pascal Friederich, Stefan Kesselheim, Kevin Maik Jablonka, Stefan Sandfeld arxiv

Fully characterizing a crystalline material requires integrating heterogeneous data sources -- atomic structures, diffraction patterns, electronic density of states, and natural language -- each of which captures a different facet of the same physical object. In practice, however, these modalities are stored and analyzed in isolation, making it difficult to relate or query materials across representational boundaries. We present MatBind, a contrastive learning framework that aligns four materials modalities -- crystal structure, powder X-ray diffraction (pXRD) simulated from structures, density of states (DOS), and text -- into a unified embedding space using crystal structure as the central physical anchor. The framework induces alignment between modalities never explicitly paired during training, enabling emergent zero-shot cross-modal retrieval as a direct consequence of the shared representation. The learned embedding space organizes materials according to physically meaningful properties without explicit supervision, and retrieval performance improves systematically when modalities are combined at query time. These results demonstrate that treating heterogeneous materials data as complementary projections of a single physical reality, rather than as isolated data sources, is not a practical choice but is consistent with the underlying physics.

📄 PDF Abstract BibTeX arXiv:2607.08470

Code (0)

등록된 구현이 없습니다.

Tasks

Zero-Shot Cross-Modal RetrievalContrastive Learning

Similar Papers 제목 키워드 기반

Unsupervised physics-informed disentanglement of multimodal data for high-throughput scientific discovery

2022-02-07 · Nathaniel Trask, Carianne Martinez, Kookjin Lee, Brad Boyce

We introduce physics-informed multimodal autoencoders (PIMA) - a variational inference framework for discovering shared information in multimodal scientific datasets representative of high-throughput testing. Individual …

DecoderDisentanglementscientific discoveryVariational Inference

Unifying Scientific Communication: Fine-Grained Correspondence Across Scientific Media

2026-05-07 · Megha Mariam K. M, Vineeth N. Balasubramanian, C. V. Jawahar arxiv

The communication of scientific knowledge has become increasingly multimodal, spanning text, visuals, and speech through materials such as research papers, slides, and recorded presentations. These different representati…

Approximate Fiber Product: A Preliminary Algebraic-Geometric Perspective on Multimodal Embedding Alignment

2024-11-30 · Dongfang Zhao

Multimodal tasks, such as image-text retrieval and generation, require embedding data from diverse modalities into a shared representation space. Aligning embeddings from heterogeneous sources while preserving shared and…

Image-text RetrievalRepresentation LearningText Retrieval

SPANER: Shared Prompt Aligner for Multimodal Semantic Representation

2025-08-18 · Thye Shan Ng, Caren Soyeon Han, Eun-Jung Holden arxiv

Recent advances in multimodal Parameter-Efficient Fine-Tuning (PEFT) have significantly improved performance on downstream tasks such as few-shot retrieval. However, most existing approaches focus on task-specific gains …

parameter-efficient fine-tuning

Unaligning Everything: Or Aligning Any Text to Any Image in Multimodal Models

2024-07-01 · Shaeke Salman, Md Montasir Bin Shams, Xiuwen Liu

Utilizing a shared embedding space, emerging multimodal models exhibit unprecedented zero-shot capabilities. However, the shared embedding space could lead to new vulnerabilities if different modalities can be misaligned…