paper-with-me

홈 › Papers

NUMINA: A Natural Understanding Benchmark for Multi-dimensional Intelligence and Numerical Reasoning Abilities

2025-09-20 · Changyu Zeng, Yifan Wang, Zimu Wang, Wei Wang, Zhengni Yang, Muyi Bao, Jiming Xiao, Anh Nguyen, Yutao Yue arxiv

Recent advancements in 2D multimodal large language models (MLLMs) have significantly improved performance in vision-language tasks. However, extending these capabilities to 3D environments remains a distinct challenge due to the complexity of spatial reasoning. Nevertheless, existing 3D benchmarks often lack fine-grained numerical reasoning task annotations, limiting MLLMs' ability to perform precise spatial measurements and complex numerical reasoning. To address this gap, we introduce NUMINA, the first Natural Understanding benchmark for Multi-dimensional Intelligence and Numerical reasoning Abilities to enhance multimodal indoor perceptual understanding. NUMINA features multi-scale annotations and various question-answer pairs, generated using NUMINA-Flow, an automated annotation pipeline that integrates LLM rewriting and rule-based self-verification. We evaluate the performance of various state-of-the-art LLMs on NUMINA following the Chat-Scene framework, demonstrating that current LLMs struggle with multimodal numerical reasoning, particularly in performing precise computations such as distance and volume estimation, highlighting the need for further advancements in 3D models. The dataset and source codes can be obtained from https://github.com/fengshun124/NUMINA.

📄 PDF Abstract BibTeX arXiv:2509.16656

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

Numina-Lean-Agent: An Open and General Agentic Reasoning System for Formal Mathematics

2026-01-20 · Junqi Liu, Zihao Zhou, Zekai Zhu, Marco Dos Santos 외 arxiv

Agentic systems have recently become the dominant paradigm for formal theorem proving, achieving strong performance by coordinating multiple models and tools. However, existing approaches often rely on task-specific pipe…

When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models

2026-04-09 · Zhengyang Sun, Yu Chen, Xin Zhou, Xiaofan Li 외 arxiv

Text-to-video diffusion models have enabled open-ended video synthesis, but often struggle with generating the correct number of objects specified in a prompt. We introduce NUMINA , a training-free identify-then-guide fr…

Goedel-Prover: A Frontier Model for Open-Source Automated Theorem Proving

2025-02-11 · Yong Lin, Shange Tang, Bohan Lyu, Jiayun Wu 외

We introduce Goedel-Prover, an open-source large language model (LLM) that achieves the state-of-the-art (SOTA) performance in automated formal proof generation for mathematical problems. The key challenge in this field …

Automated Theorem ProvingLarge Language ModelMath

Emotion Entanglement and Bayesian Inference for Multi-Dimensional Emotion Understanding

2026-04-01 · Hemanth Kotaprolu, Kishan Maharaj, Raey Zhao, Abhijit Mishra 외 arxiv

Understanding emotions in natural language is inherently a multi-dimensional reasoning problem, where multiple affective signals interact through context, interpersonal relations, and situational cues. However, most exis…

Bayesian Inference

Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate

2025-01-29 · YuBo Wang, Xiang Yue, Wenhu Chen

Supervised Fine-Tuning (SFT) is commonly used to train language models to imitate annotated responses for given instructions. In this paper, we propose Critique Fine-Tuning (CFT), a method more effective than SFT for rea…

Instruction FollowingMathMathematical Reasoning