paper-with-me

홈 › Papers

EgoNormia: Benchmarking Physical Social Norm Understanding

2025-02-27 · MohammadHossein Rezaei, Yicheng Fu, Phil Cuvin, Caleb Ziems, Yanzhe Zhang, Hao Zhu, Diyi Yang

Human activity is moderated by norms. However, machines are often trained without explicit supervision on norm understanding and reasoning, particularly when norms are physically- or socially-grounded. To improve and evaluate the normative reasoning capability of vision-language models (VLMs), we present \dataset{} $\|\epsilon\|$, consisting of 1,853 challenging, multi-stage MCQ questions based on ego-centric videos of human interactions, evaluating both the prediction and justification of normative actions. The normative actions encompass seven categories: safety, privacy, proxemics, politeness, cooperation, coordination/proactivity, and communication/legibility. To compile this dataset at scale, we propose a novel pipeline leveraging video sampling, automatic answer generation, filtering, and human validation. Our work demonstrates that current state-of-the-art vision-language models lack robust norm understanding, scoring a maximum of 54\% on \dataset{} (versus a human bench of 92\%). Our analysis of performance in each dimension highlights the significant risks of safety, privacy, and the lack of collaboration and communication capability when applied to real-world agents. We additionally show that through a retrieval-based generation (RAG) method, it is possible to use \dataset{} to enhance normative reasoning in VLMs.

📄 PDF Abstract BibTeX arXiv:2502.20490

Code (1)

open-social-world/egonormia 공식 구현

Tasks

Answer GenerationBenchmarkingRAG

Similar Papers 제목 키워드 기반

"A Woman is More Culturally Knowledgeable than A Man?": The Effect of Personas on Cultural Norm Interpretation in LLMs

2024-09-18 · Mahammed Kamruzzaman, Hieu Nguyen, Nazmul Hassan, Gene Louis Kim

As the deployment of large language models (LLMs) expands, there is an increasing demand for personalized LLMs. One method to personalize and guide the outputs of these models is by assigning a persona -- a role that des…

Norms, Institutions, and Robots

2018-07-30 · Stevan Tomic, Federico Pecora, Alessandro Saffiotti

Interactions within human societies are usually regulated by social norms. If robots are to be accepted into human society, it is essential that they are aware of and capable of reasoning about social norms. In this pape…

Where Norms and References Collide: Evaluating LLMs on Normative Reasoning

2026-02-03 · Mitchell Abrams, Kaveh Eskandari Miandoab, Felix Gervits, Vasanth Sarathy 외 arxiv

Embodied agents, such as robots, will need to interact in situated environments where successful communication often depends on reasoning over social norms: shared expectations that constrain what actions are appropriate…

Eyes on VLM: Benchmarking Gaze Following and Social Gaze Prediction in Vision Language Models

2026-05-19 · Hengfei Wang, Anshul Gupta, Pierre Vuillecard, Jean-Marc Odobez arxiv

Vision-language models (VLMs) have rapidly evolved into general-purpose multimodal reasoners with strong zero-shot generalization. In this context, VLMs could greatly benefit the analysis of human gaze and attention, a c…

Zero-shot GeneralizationRelational Reasoning

Measuring Physical-World Privacy Awareness of Large Language Models: An Evaluation Benchmark

2025-09-27 · Xinjie Shen, Mufei Li, Pan Li arxiv

The deployment of Large Language Models (LLMs) in embodied agents creates an urgent need to measure their privacy awareness in the physical world. Existing evaluation methods, however, are confined to natural language ba…