paper-with-me

홈 › Papers

AtomWorld: A Benchmark for Evaluating Spatial Reasoning in Large Language Models on Crystalline Materials

2025-10-06 · Taoyuze Lv, Alexander Chen, Fengyu Xie, Chu Wu, Jeffrey Meng, Dongzhan Zhou, Yingheng Wang, Bram Hoex, Zhicheng Zhong, Tong Xie arxiv

Large language models (LLMs) have shown promising potential in scientific research, enabling tasks ranging from knowledge retrieval to property prediction. Existing science benchmarks mainly focus on perceptual or knowledge-based tasks, largely ignoring the modelling tasks, a fundamental starting point for any real scientific research. For materials science, constructing and manipulating atomic structures is one of the most creative and least automated steps. In this work, we introduce AtomWorld, a benchmark designed to evaluate the abilities of LLMs on structure modifications. The benchmark includes ten fundamental actions under four widely used modelling categories, enabling verifiable evaluation metrics. We find that Claude Opus 4.6 generally performs the best. While the success rate decreases markedly with increasing modelling complexity, with particularly low success rates (below 12\% for rotation) for operations involving complex spatial relations. Our results suggest that contemporary LLMs are better suited as copilots for materials structure modelling rather than fully unsupervised autonomous scientific agents. Beyond evaluation, AtomWorld also serves as a testbed and playground for developing future structure-aware models, including reinforcement learning and agentic approaches.

📄 PDF Abstract BibTeX arXiv:2510.04704

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningSpatial Reasoning

Similar Papers 제목 키워드 기반

MapVerse: A Benchmark for Geospatial Question Answering on Diverse Real-World Maps

2026-02-11 · Sharat Bhat, Harshita Khandelwal, Tushar Kataria, Vivek Gupta arxiv

Maps are powerful carriers of structured and contextual knowledge, encompassing geography, demographics, infrastructure, and environmental patterns. Reasoning over such knowledge requires models to integrate spatial rela…

Multimodal ReasoningQuestion AnsweringSpatial Reasoning

Spatial4D-Bench: A Versatile 4D Spatial Intelligence Benchmark

2025-12-31 · Pan Wang, Yang Liu, Guile Wu, Eduardo R. Corral-Soto 외 arxiv

4D spatial intelligence involves perceiving and processing how objects move or change over time. Humans naturally possess 4D spatial intelligence, supporting a broad spectrum of spatial reasoning abilities. To what exten…

Scene UnderstandingAction RecognitionSpatial Reasoning

Spatial-DISE: A Unified Benchmark for Evaluating Spatial Reasoning in Vision-Language Models

2025-10-15 · Xinmiao Huang, Qisong He, Zhenglin Huang, Boxuan Wang 외 arxiv

Spatial reasoning ability is crucial for Vision Language Models (VLMs) to support real-world applications in diverse domains including robotics, augmented reality, and autonomous navigation. Unfortunately, existing bench…

Spatial Reasoning

FloorplanQA: A Benchmark for Spatial Reasoning in LLMs using Structured Representations

2025-07-10 · Fedor Rodionov, Abdelrahman Eldesokey, Michael Birsak, John Femiani 외 arxiv

We introduce FloorplanQA, a diagnostic benchmark for evaluating spatial reasoning in large language models (LLMs). FloorplanQA is grounded in structured representations of indoor scenes, such as (e.g., kitchens, living r…

Spatial Reasoning

Thinking in Structures: Evaluating Spatial Intelligence in Constraint-Governed Spaces

2026-02-08 · Chen Yang, Guanxin Lin, Youquan He, Peiyao Chen 외 arxiv

Spatial intelligence is crucial for vision--language models (VLMs), yet many scene-centric benchmarks evaluate unconstrained environments where a single image may admit multiple plausible 3D interpretations. We introduce…

Spatial Reasoning