paper-with-me

홈 › Papers

AGIBench: A Multi-granularity, Multimodal, Human-referenced, Auto-scoring Benchmark for Large Language Models

2023-09-05 · Fei Tang, Wanling Gao, Luzhou Peng, Jianfeng Zhan

Large language models (LLMs) like ChatGPT have revealed amazing intelligence. How to evaluate the question-solving abilities of LLMs and their degrees of intelligence is a hot-spot but challenging issue. First, the question-solving abilities are interlaced with different ability branches like understanding and massive knowledge categories like mathematics. Second, the inputs of questions are multimodal that may involve text and images. Third, the response format of LLMs is diverse and thus poses great challenges for result extraction and evaluation. In this paper, we propose AGIBench -- a multi-granularity, multimodal, human-referenced, and auto-scoring benchmarking methodology for LLMs. Instead of a collection of blended questions, AGIBench focuses on three typical ability branches and adopts a four-tuple <ability branch, knowledge, difficulty, modal> to label the attributes of each question. First, it supports multi-granularity benchmarking, e.g., per-question, per-ability branch, per-knowledge, per-modal, per-dataset, and per-difficulty level granularities. Second, it contains multimodal input, including text and images. Third, it classifies all the questions into five degrees of difficulty according to the average accuracy rate of abundant educated humans (human-referenced). Fourth, it adopts zero-shot learning to avoid introducing additional unpredictability and provides an auto-scoring method to extract and judge the result. Finally, it defines multi-dimensional metrics, including accuracy under the average, worst, best, and majority voting cases, and repeatability. AGIBench is publically available from \url{https://www.benchcouncil.org/agibench}.

📄 PDF Abstract BibTeX arXiv:2309.06495

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingZero-Shot Learning

Similar Papers 제목 키워드 기반

Benchmarking Egocentric Multimodal Goal Inference for Assistive Wearable Agents

2025-10-25 · Vijay Veerabadran, Fanyi Xiao, Nitin Kamra, Pedro Matias 외 arxiv

There has been a surge of interest in assistive wearable agents: agents embodied in wearable form factors (e.g., smart glasses) who take assistive actions toward a user's goal/query (e.g. "Where did I leave my keys?"). I…

ISO-Bench: Benchmarking Multimodal Causal Reasoning in Visual-Language Models through Procedural Plans

2025-07-30 · Ananya Sadana, Yash Kumar Lal, Jiawei Zhou arxiv

Understanding causal relationships across modalities is a core challenge for multimodal models operating in real-world environments. We introduce ISO-Bench, a benchmark for evaluating whether models can infer causal depe…

Enhancing Multimodal Emotion Recognition through Multi-Granularity Cross-Modal Alignment

2024-12-30 · Xuechen Wang, Shiwan Zhao, Haoqin Sun, Hui Wang 외

Multimodal emotion recognition (MER), leveraging speech and text, has emerged as a pivotal domain within human-computer interaction, demanding sophisticated methods for effective multimodal integration. The challenge of …

cross-modal alignmentEmotion RecognitionMultimodal Emotion Recognition

Each Fake News is Fake in its Own Way: An Attribution Multi-Granularity Benchmark for Multimodal Fake News Detection

2024-12-19 · Hao Guo, Zihan Ma, Zhi Zeng, Minnan Luo 외

Social platforms, while facilitating access to information, have also become saturated with a plethora of fake news, resulting in negative consequences. Automatic multimodal fake news detection is a worthwhile pursuit. E…

Fake News Detection

Target-Oriented Object Grasping via Multimodal Human Guidance

2024-08-20 · Pengwei Xie, Siang Chen, Dingchang Hu, Yixiang Dai 외

In the context of human-robot interaction and collaboration scenarios, robotic grasping still encounters numerous challenges. Traditional grasp detection methods generally analyze the entire scene to predict grasps, lead…

Motion PlanningObjectRobotic Grasping