paper-with-me

Papers

Stress-Testing Multimodal Foundation Models for Crystallographic Reasoning

2025-06-16 · Can Polat, Hasan Kurban, Erchin Serpedin, Mustafa Kurban

Evaluating foundation models for crystallographic reasoning requires benchmarks that isolate generalization behavior while enforcing physical constraints. This work introduces a multiscale multicrystal dataset with two physically grounded evaluation protocols to stress-test multimodal generative models. The Spatial-Exclusion benchmark withholds all supercells of a given radius from a diverse dataset, enabling controlled assessments of spatial interpolation and extrapolation. The Compositional-Exclusion benchmark omits all samples of a specific chemical composition, probing generalization across stoichiometries. Nine vision--language foundation models are prompted with crystallographic images and textual context to generate structural annotations. Responses are evaluated via (i) relative errors in lattice parameters and density, (ii) a physics-consistency index penalizing volumetric violations, and (iii) a hallucination score capturing geometric outliers and invalid space-group predictions. These benchmarks establish a reproducible, physically informed framework for assessing generalization, consistency, and reliability in large-scale multimodal models. Dataset and code are available at https://github.com/KurbanIntelligenceLab/StressTestingMMFMinCR.

📄 PDF Abstract BibTeX arXiv:2506.13051

Code (1)

kurbanintelligencelab/stresstestingmmfmincr 공식 구현

Tasks

HallucinationSpatial Interpolation

Similar Papers 제목 키워드 기반

CrystalXRD-Bench: Benchmarking Vision-Language Models for XRD Peak Indexing Across Diverse Crystalline Materials

2026-05-28 · Chengliang Xu, Xiaogang Li, Peiyao Xiao, Beng Wang 외 arxiv

Miller-index identification from powder XRD patterns requires capabilities untested by existing multimodal benchmarks: the model must read a narrow peak location from a rendered scientific curve and then connect that obs…

MORSE-500: A Programmatically Controllable Video Benchmark to Stress-Test Multimodal Reasoning

2025-06-05 · Zikui Cai, Andrew Wang, Anirudh Satheesh, Ankit Nakhawa 외

Despite rapid advances in vision-language models (VLMs), current benchmarks for multimodal reasoning fall short in three key dimensions. First, they overwhelmingly rely on static images, failing to capture the temporal c…

Dataset GenerationMathematical Problem-SolvingMultimodal Reasoning

Miller-Index-Based Latent Crystallographic Fracture Plane Reasoning and generation with Vision-Language Models

2026-05-19 · Qinwu Xu, Xiaofu Ma, Yifan Jiang arxiv

We study whether multimodal large language models (MLLMs) can leverage crystallographic plane indices (Miller indices) as a structured latent representation for reasoning about fracture geometry. We formulate Miller indi…

FCMBench: The First Large-scale Financial Credit Multimodal Benchmark for Real-world Applications

2026-01-01 · Yehui Yang, Dalu Yang, Fangxin Shang, Wenshuo Zhou 외 arxiv

FCMBench is the first large-scale and privacy-compliant multimodal benchmark for real-world financial credit applications, covering tasks and robustness challenges from domain specific workflows and constraints. The curr…

Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models

2025-03-04 · Zhifei Xie, Mingbao Lin, Zihang Liu, Pengcheng Wu 외

Recent advancements in multimodal reasoning have largely overlooked the audio modality. We introduce Audio-Reasoner, a large-scale audio language model for deep reasoning in audio tasks. We meticulously curated a large-s…

Language ModelingLanguage ModellingMultimodal Reasoning