paper-with-me

홈 › Papers

LLM-Cave: A benchmark and light environment for large language models reasoning and decision-making system

2025-11-27 · Huanyu Li, Zongyuan Li, Wei Huang, Xian Guo arxiv

Large language models (LLMs) such as ChatGPT o1, ChatGPT o3, and DeepSeek R1 have shown great potential in solving difficult problems. However, current LLM evaluation benchmarks are limited to one-step interactions. Some of the existing sequence decision-making environments, such as TextStarCraftII and LLM-PySC2, are too complicated and require hours of interaction to complete a game. In this paper, we introduce LLM-Cave, a benchmark and light environment for LLM reasoning and decision-making systems. This environment is a classic instance in the era of Symbolism. Artificial intelligence enables the agent to explore the environment and avoid potential losses by reasoning about nearby dangers using partial observable state information. In the experiment, we evaluated the sequential reasoning ability, decision-making performance and computational efficiency of mainstream large language models (LLMs) such as GPT-4o-mini, o1-mini, and DeepSeek-R1. Experiments show that while Deepseek-R1 achieved the highest success rate on complex reasoning tasks, smaller models like 4o-mini significantly narrowed the performance gap on challenges by employing Chain of Speculation and Planner-Critic strategies, at the expense of reduced computational efficiency. This indicates that structured, multi-step reasoning combined with an LLM-based feedback mechanism can substantially enhance an LLM's decision-making capabilities, providing a promising direction for improving reasoning in weaker models and suggesting a new reasoning-centered benchmark for LLM assessment. Our code is open-sourced in https://github.com/puleya1277/CaveEnv.

📄 PDF Abstract BibTeX arXiv:2511.22598

Code (0)

등록된 구현이 없습니다.

Tasks

Computational Efficiency

Similar Papers 제목 키워드 기반

CAVE-NAV: VLM-Based Autonomous 3D Navigation in Underwater Cave Environments

2026-08-28 · Zhenqi Wu, Yuanjie Lu, Yisheng Zhang, Miao Yu 외 arxiv

Autonomous navigation in underwater cave environments is essential for search-and-rescue operations, scientific exploration, and emergency egress. Traditional navigation systems commonly depend on dense visual features f…

CaveSeg: Deep Semantic Segmentation and Scene Parsing for Autonomous Underwater Cave Exploration

2023-09-20 · A. Abdullah, T. Barua, R. Tibbetts, Z. Chen 외

In this paper, we present CaveSeg - the first visual learning pipeline for semantic segmentation and scene parsing for AUV navigation inside underwater caves. We address the problem of scarce annotated training data by p…

Scene ParsingSegmentationSemantic Segmentation

CAVERS: Multimodal SLAM Data from a Natural Karstic Cave with Ground Truth Motion Capture

2026-04-16 · Giacomo Franchini, David Rodríguez-Martínez, Alfonso Martínez-Petersen, C. J. Pérez-del-Pulgar 외 arxiv

Autonomous robots operating in natural karstic caves face perception and navigation challenges that are qualitatively distinct from those encountered in mines or tunnels: irregular geometry, reflective wet surfaces, near…

3D Reconstruction

Weakly Supervised Caveline Detection For AUV Navigation Inside Underwater Caves

2023-03-07 · Boxiao Yu, Reagan Tibbetts, Titon Barua, Ailani Morales 외

Underwater caves are challenging environments that are crucial for water resource management, and for our understanding of hydro-geology and history. Mapping underwater caves is a time-consuming, labor-intensive, and haz…

Management

CAVE: Detecting and Explaining Commonsense Anomalies in Visual Environments

2025-10-29 · Rishika Bhagwatkar, Syrielle Montariol, Angelika Romanou, Beatriz Borges 외 arxiv

Humans can naturally identify, reason about, and explain anomalies in their environment. In computer vision, this long-standing challenge remains limited to industrial defects or unrealistic, synthetically generated anom…

Anomaly DetectionVisual Grounding