paper-with-me

홈 › Papers

Metrics vs Surveys: An Analysis for Human-Aligned Benchmarking in Social Robot Navigation

2025-10-03 · Stefano Trepella, Mauro Martini, Noé Pérez-Higueras, Andrea Ostuni, Fernando Caballero, Luis Merino, Marcello Chiaberge arxiv

Social, also called human-aware, navigation is a key challenge for integrating mobile robots into human environments. The evaluation of such systems is complex, as factors such as comfort, safety, and legibility must be considered. Human-centered assessments, typically conducted through surveys, provide reliable insights but are costly, resource-intensive, and difficult to reproduce or compare across systems. Alternatively, numerical social navigation metrics are easy to compute and facilitate comparisons, yet the community lacks consensus on a standard set of metrics. This work explores the relationship between numerical metrics and human-centered evaluations to identify potential correlations. If specific quantitative measures align with human perceptions, they could serve as preliminary benchmarking tools, providing a human-aligned assessment when large-scale surveys are not feasible. Our results indicate that while current metrics capture some aspects of robot navigation behavior, important subjective factors remain insufficiently represented, necessitating new metrics.

📄 PDF Abstract BibTeX arXiv:2510.02941

Code (0)

등록된 구현이 없습니다.

Tasks

Robot Navigation

Similar Papers 제목 키워드 기반

A short methodological review on social robot navigation benchmarking

2025-10-25 · Pranup Chhetri, Alejandro Torrejon, Sergio Eslava, Luis J. Manso arxiv

Social Robot Navigation is the skill that allows robots to move efficiently in human-populated environments while ensuring safety, comfort, and trust. Unlike other areas of research, the scientific community has not yet …

Robot Navigation

Benchmarking Contemporary Deep Learning Hardware and Frameworks:A Survey of Qualitative Metrics

2019-07-05 · Wei Dai, Daniel Berleant

This paper surveys benchmarking principles, machine learning devices including GPUs, FPGAs, and ASICs, and deep learning software frameworks. It also reviews these technologies with respect to benchmarking from the persp…

BenchmarkingBIG-bench Machine LearningDeep Learning

VizSeq: A Visual Analysis Toolkit for Text Generation Tasks

2019-09-12 · IJCNLP 2019 11 · Changhan Wang, Anirudh Jain, Danlu Chen, Jiatao Gu

Automatic evaluation of text generation tasks (e.g. machine translation, text summarization, image captioning and video description) usually relies heavily on task-specific metrics, such as BLEU and ROUGE. They, however,…

BenchmarkingImage CaptioningMachine TranslationText Generation+3

Enhancing Trust in LLMs: Algorithms for Comparing and Interpreting LLMs

2024-06-04 · Nik Bear Brown

This paper surveys evaluation techniques to enhance the trustworthiness and understanding of Large Language Models (LLMs). As reliance on LLMs grows, ensuring their reliability, fairness, and transparency is crucial. We …

BenchmarkingFairnessFew-Shot LearningHallucination+1

SVGauge: Towards Human-Aligned Evaluation for SVG Generation

2025-09-08 · Leonardo Zini, Elia Frigieri, Sebastiano Aloscari, Marcello Generali 외 arxiv

Generated Scalable Vector Graphics (SVG) images demand evaluation criteria tuned to their symbolic and vectorial nature: criteria that existing metrics such as FID, LPIPS, or CLIPScore fail to satisfy. In this paper, we …