paper-with-me

Papers

HumaniBench: A Human-Centric Framework for Large Multimodal Models Evaluation

2025-05-16 · Shaina Raza, Aravind Narayanan, Vahid Reza Khazaie, Ashmal Vayani, Mukund S. Chettiar, Amandeep Singh, Mubarak Shah, Deval Pandya

Large multimodal models (LMMs) now excel on many vision language benchmarks, however, they still struggle with human centered criteria such as fairness, ethics, empathy, and inclusivity, key to aligning with human values. We introduce HumaniBench, a holistic benchmark of 32K real-world image question pairs, annotated via a scalable GPT4o assisted pipeline and exhaustively verified by domain experts. HumaniBench evaluates seven Human Centered AI (HCAI) principles: fairness, ethics, understanding, reasoning, language inclusivity, empathy, and robustness, across seven diverse tasks, including open and closed ended visual question answering (VQA), multilingual QA, visual grounding, empathetic captioning, and robustness tests. Benchmarking 15 state of the art LMMs (open and closed source) reveals that proprietary models generally lead, though robustness and visual grounding remain weak points. Some open-source models also struggle to balance accuracy with adherence to human-aligned principles. HumaniBench is the first benchmark purpose built around HCAI principles. It provides a rigorous testbed for diagnosing alignment gaps and guiding LMMs toward behavior that is both accurate and socially responsible. Dataset, annotation prompts, and evaluation code are available at: https://vectorinstitute.github.io/HumaniBench

📄 PDF Abstract BibTeX arXiv:2505.11454

Code (1)

vectorinstitute/humanibench 공식 구현 pytorch

Tasks

BenchmarkingEthicsFairnessQuestion AnsweringVisual GroundingVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

EgoM2P: Egocentric Multimodal Multitask Pretraining

2025-06-09 · Gen Li, Yutong Chen, Yiqian Wu, Kaifeng Zhao 외

Understanding multimodal signals in egocentric vision, such as RGB video, depth, camera poses, and gaze, is essential for applications in augmented reality, robotics, and human-computer interaction. These capabilities en…

Depth EstimationGaze PredictionMonocular Depth Estimation

MIND: Multimodal Shopping Intention Distillation from Large Vision-language Models for E-commerce Purchase Understanding

2024-06-15 · Baixuan Xu, Weiqi Wang, Haochen Shi, Wenxuan Ding 외

Improving user experience and providing personalized search results in E-commerce platforms heavily rely on understanding purchase intention. However, existing methods for acquiring large-scale intentions bank on distill…

HERM: Benchmarking and Enhancing Multimodal LLMs for Human-Centric Understanding

2024-10-09 · Keliang Li, Zaifei Yang, Jiahe Zhao, Hongze Shen 외

The significant advancements in visual understanding and instruction following from Multimodal Large Language Models (MLLMs) have opened up more possibilities for broader applications in diverse and universal human-centr…

BenchmarkingInstruction Following

Nymeria: A Massive Collection of Multimodal Egocentric Daily Motion in the Wild

2024-06-14 · Lingni Ma, Yuting Ye, Fangzhou Hong, Vladimir Guzov 외

We introduce Nymeria - a large-scale, diverse, richly annotated human motion dataset collected in the wild with multiple multimodal egocentric devices. The dataset comes with a) full-body ground-truth motion; b) multiple…

Action RecognitionGaze EstimationMotion Synthesis

When does CLIP generalize better than unimodal models? When judging human-centric concepts

2022-05-01 · RepL4NLP (ACL) 2022 5 · Romain Bielawski, Benjamin Devillers, Tim Van De Cruys, Rufin VanRullen

CLIP, a vision-language network trained with a multimodal contrastive learning objective on a large dataset of images and captions, has demonstrated impressive zero-shot ability in various tasks. However, recent work sho…

ClassificationContrastive LearningGenre classificationSentiment Analysis+1