paper-with-me

홈 › Papers

3DBench: A Scalable 3D Benchmark and Instruction-Tuning Dataset

2024-04-23 · Junjie Zhang, Tianci Hu, Xiaoshui Huang, Yongshun Gong, Dan Zeng

Evaluating the performance of Multi-modal Large Language Models (MLLMs), integrating both point cloud and language, presents significant challenges. The lack of a comprehensive assessment hampers determining whether these models truly represent advancements, thereby impeding further progress in the field. Current evaluations heavily rely on classification and caption tasks, falling short in providing a thorough assessment of MLLMs. A pressing need exists for a more sophisticated evaluation method capable of thoroughly analyzing the spatial understanding and expressive capabilities of these models. To address these issues, we introduce a scalable 3D benchmark, accompanied by a large-scale instruction-tuning dataset known as 3DBench, providing an extensible platform for a comprehensive evaluation of MLLMs. Specifically, we establish the benchmark that spans a wide range of spatial and semantic scales, from object-level to scene-level, addressing both perception and planning tasks. Furthermore, we present a rigorous pipeline for automatically constructing scalable 3D instruction-tuning datasets, covering 10 diverse multi-modal tasks with more than 0.23 million QA pairs generated in total. Thorough experiments evaluating trending MLLMs, comparisons against existing datasets, and variations of training protocols demonstrate the superiority of 3DBench, offering valuable insights into current limitations and potential research directions.

📄 PDF Abstract BibTeX arXiv:2404.14678

Code (1)

Inshsang/3DBench 공식 구현

Similar Papers 제목 키워드 기반

IF-RewardBench: Benchmarking Judge Models for Instruction-Following Evaluation

2026-03-05 · Bosi Wen, Yilin Niu, Cunxiang Wang, Xiaoying Ling 외 arxiv

Instruction-following is a foundational capability of large language models (LLMs), with its improvement hinging on scalable and accurate feedback from judge models. However, the reliability of current judge models in in…

SimdBench: Benchmarking Large Language Models for SIMD-Intrinsic Code Generation

2025-07-21 · Yibo He, Shuoran Zhao, Jiaming Huang, Yingjie Fu 외 arxiv

SIMD (Single Instruction Multiple Data) instructions and their compiler intrinsics are widely supported by modern processors to accelerate performance-critical tasks. SIMD intrinsic programming, a trade-off between codin…

Code Generation

MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding

2025-12-06 · Yuhao Su, Anwesa Choudhuri, Zhongpai Gao, Benjamin Planche 외 arxiv

Large vision-language models struggle with medical video understanding, where spatial precision, temporal reasoning, and clinical semantics are critical. To address this, we first introduce \textbf{MedVidBench}, a large-…

Reinforcement Learning

PodBench: A Comprehensive Benchmark for Instruction-Aware Audio-Oriented Podcast Script Generation

2026-01-21 · Chenning Xu, Mao Zheng, Mingyu Zheng, Mingyang Song arxiv

Podcast script generation requires LLMs to synthesize structured, context-grounded dialogue from diverse inputs, yet systematic evaluation resources for this task remain limited. To bridge this gap, we introduce PodBench…

Instruction Following

Embodied3DBench: Benchmarking Low-Level Embodied Spatial Intelligence of Vision Language Models

2026-05-27 · Jiyao Zhang, Mingxu Zhang, Yitong Peng, Haoxuan Liu 외 arxiv

Are current Vision Language Models (VLMs) ready to comprehend and reason about complex embodied interactions in 3D environments? We introduce Embodied3DBench, a robot-centric benchmark targeting low-level spatial intelli…

Trajectory PredictionSpatial Reasoning