paper-with-me

홈 › Papers

Multimodal Crystal Flow: Any-to-Any Modality Generation for Unified Crystal Modeling

2026-02-23 · Kiyoung Seong, Sungsoo Ahn, Sehui Han, Changyoung Park arxiv

Crystal modeling spans a family of conditional and unconditional generation tasks, including crystal structure prediction (CSP) and de novo generation (DNG). While recent deep generative models have shown promising performance, they remain largely task-specific, lacking a unified framework that shares crystal representations across tasks. To address this limitation, we propose Multimodal Crystal Flow (MCFlow), a unified multimodal flow model that realizes multiple crystal generation tasks as distinct inference trajectories via independent time variables for atom types and crystal structures. To enable multimodal flow in a standard transformer model, we introduce a composition- and symmetry-aware atom ordering with hierarchical permutation augmentation, injecting compositional and crystallographic priors without explicit structural templates. Experiments on the MP-20 and MPTS-52 benchmarks show that a single MCFlow model is competitive with task-specific baselines across CSP, DNG, and structure-conditioned atom type generation.

📄 PDF Abstract BibTeX arXiv:2602.20210

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DMFlow: Disordered Materials Generation by Flow Matching

2026-02-04 · Liming Wu, Rui Jiao, Qi Li, Mingze Li 외 arxiv

The design of materials with tailored properties is crucial for technological progress. However, most deep generative models focus exclusively on perfectly ordered crystals, neglecting the important class of disordered m…

Graph Neural Network

Uni-ViGU: Towards Unified Video Generation and Understanding via A Diffusion-Based Video Generator

2026-04-09 · Luozheng Qin, Jia Gong, Qian Qiao, Tianjiao Li 외 arxiv

Unified multimodal models integrating visual understanding and generation face a fundamental challenge: visual generation incurs substantially higher computational costs than understanding, particularly for video. This i…

multimodal generationVideo GenerationText Generation

Unispeaker: A Unified Approach for Multimodality-driven Speaker Generation

2025-01-11 · Zhengyan Sheng, Zhihao Du, Heng Lu, Shiliang Zhang 외

Recent advancements in personalized speech generation have brought synthetic speech increasingly close to the realism of target speakers' recordings, yet multimodal speaker generation remains on the rise. This paper intr…

Diversity

UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation

2025-06-20 · Teng Li, Quanfeng Lu, Lirui Zhao, Hao Li 외

Unified image understanding and generation has emerged as a promising paradigm in multimodal artificial intelligence. Despite recent progress, the optimal architectural design for such unified models remains an open chal…

Representation Learning

CrystalDiT: A Diffusion Transformer for Crystal Generation

2025-08-13 · Xiaohan Yi, Guikun Xu, Xi Xiao, Zhong Zhang 외 arxiv

We present CrystalDiT, a diffusion transformer for crystal structure generation that achieves state-of-the-art performance by challenging the trend of architectural complexity. Instead of intricate, multi-stream designs,…