paper-with-me

Papers

JRDB-Reasoning: A Difficulty-Graded Benchmark for Visual Reasoning in Robotics

2025-08-14 · Simindokht Jahangard, Mehrzad Mohammadi, Yi Shen, Zhixi Cai, Hamid Rezatofighi arxiv

Recent advances in Vision-Language Models (VLMs) and large language models (LLMs) have greatly enhanced visual reasoning, a key capability for embodied AI agents like robots. However, existing visual reasoning benchmarks often suffer from several limitations: they lack a clear definition of reasoning complexity, offer have no control to generate questions over varying difficulty and task customization, and fail to provide structured, step-by-step reasoning annotations (workflows). To bridge these gaps, we formalize reasoning complexity, introduce an adaptive query engine that generates customizable questions of varying complexity with detailed intermediate annotations, and extend the JRDB dataset with human-object interaction and geometric relationship annotations to create JRDB-Reasoning, a benchmark tailored for visual reasoning in human-crowded environments. Our engine and benchmark enable fine-grained evaluation of visual reasoning frameworks and dynamic assessment of visual-language models across reasoning levels.

📄 PDF Abstract BibTeX arXiv:2508.10287

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Reasoning

Similar Papers 제목 키워드 기반

JRDB-Pose: A Large-scale Dataset for Multi-Person Pose Estimation and Tracking

2022-10-20 · CVPR 2023 1 · Edward Vendrow, Duy Tho Le, Jianfei Cai, Hamid Rezatofighi

Autonomous robotic systems operating in human environments must understand their surroundings to make accurate and safe decisions. In crowded human scenes with close-up human-robot interaction and robot navigation, a dee…

DiversityMulti-Person Pose EstimationMulti-Person Pose Estimation and TrackingPose Estimation+2

JRDB-PanoTrack: An Open-world Panoptic Segmentation and Tracking Robotic Dataset in Crowded Human Environments

2024-04-02 · CVPR 2024 1 · Duy-Tho Le, Chenhui Gou, Stavya Datta, Hengcan Shi 외

Autonomous robot systems have attracted increasing research attention in recent years, where environment understanding is a crucial step for robot navigation, human-robot interaction, and decision. Real-world robot syste…

Decision MakingPanoptic SegmentationRobot Navigation

DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training

2025-04-24 · Xiaoyu Tian, Sitong Zhao, Haotian Wang, Shuaiting Chen 외

Although large language models (LLMs) have recently achieved remarkable performance on various complex reasoning benchmarks, the academic community still lacks an in-depth understanding of base model training processes a…

Mathematical Reasoning

SuperCLUE-Math6: Graded Multi-Step Math Reasoning Benchmark for LLMs in Chinese

2024-01-22 · Liang Xu, Hang Xue, Lei Zhu, Kangkang Zhao

We introduce SuperCLUE-Math6(SC-Math6), a new benchmark dataset to evaluate the mathematical reasoning abilities of Chinese language models. SC-Math6 is designed as an upgraded Chinese version of the GSM8K dataset with e…

DiversityGSM8KMathMathematical Reasoning

MathMixup: Boosting LLM Mathematical Reasoning with Difficulty-Controllable Data Synthesis and Curriculum Learning

2026-01-14 · Xuchen Li, Jing Chen, Xuzhao Li, Hao Liang 외 arxiv

In mathematical reasoning tasks, the advancement of Large Language Models (LLMs) relies heavily on high-quality training data with clearly defined and well-graded difficulty levels. However, existing data synthesis metho…

Mathematical Reasoning