paper-with-me

홈 › Papers

Revisiting Reliability in the Reasoning-based Pose Estimation Benchmark

2025-07-17 · Junsu Kim, Naeun Kim, Jaeho Lee, Incheol Park, Dongyoon Han, Seungryul Baek

The reasoning-based pose estimation (RPE) benchmark has emerged as a widely adopted evaluation standard for pose-aware multimodal large language models (MLLMs). Despite its significance, we identified critical reproducibility and benchmark-quality issues that hinder fair and consistent quantitative evaluations. Most notably, the benchmark utilizes different image indices from those of the original 3DPW dataset, forcing researchers into tedious and error-prone manual matching processes to obtain accurate ground-truth (GT) annotations for quantitative metrics (\eg, MPJPE, PA-MPJPE). Furthermore, our analysis reveals several inherent benchmark-quality limitations, including significant image redundancy, scenario imbalance, overly simplistic poses, and ambiguous textual descriptions, collectively undermining reliable evaluations across diverse scenarios. To alleviate manual effort and enhance reproducibility, we carefully refined the GT annotations through meticulous visual matching and publicly release these refined annotations as an open-source resource, thereby promoting consistent quantitative evaluations and facilitating future advancements in human pose-aware multimodal reasoning.

📄 PDF Abstract BibTeX arXiv:2507.13314

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal ReasoningPose Estimation

Similar Papers 제목 키워드 기반

Revisiting Uncertainty Estimation and Calibration of Large Language Models

2025-05-29 · Linwei Tao, Yi-Fan Yeh, Minjing Dong, Tao Huang 외

As large language models (LLMs) are increasingly deployed in high-stakes applications, robust uncertainty estimation is essential for ensuring the safe and trustworthy deployment of LLMs. We present the most comprehensiv…

Mixture-of-ExpertsMMLUQuantization

Calibration vs Decision Making: Revisiting the Reliability Paradox in Unlearned Language Models

2026-05-20 · Divyaksh Shukla, Ashutosh Modi arxiv

Machine unlearning aims to remove the influence of specific training data from a model while preserving reliable behavior on the remaining data, making reliable prediction and uncertainty estimation essential for evaluat…

Decision Making

Position Rebinding Cache Reuse: Replay-Free Visual Revisiting for Interleaved Multimodal Reasoning

2026-06-25 · Mengzhao Wang, Yanli Ji, Wangmeng Zuo, Peng Ye 외 arxiv

Interleaved multimodal reasoning improves visual grounding by revisiting visual evidence during multi-step generation, yet existing methods typically rely on token replay, repeatedly forwarding selected visual tokens. A …

Multimodal ReasoningVisual Grounding

When Does Intrinsic Self-Correction Help? A Task-Sensitive Analysis

2026-06-22 · Elroy Stav, Dvir Berlowitz, Maayan Orner, Sarit Kraus arxiv

Intrinsic self-correction (SC) aims to improve large language model outputs by prompting a model to revisit its own initial answer without external feedback. Recent studies have questioned the reliability of this approac…

Revisiting the Reliability of Language Models in Instruction-Following

2025-12-15 · Jianshuo Dong, Yutong Zhang, Yan Liu, Zhenyu Zhong 외 arxiv

Advanced LLMs have achieved near-ceiling instruction-following accuracy on benchmarks such as IFEval. However, these impressive scores do not necessarily translate to reliable services in real-world use, where users ofte…

Data Augmentation