paper-with-me

홈 › Papers

FB-Bench: A Fine-Grained Multi-Task Benchmark for Evaluating LLMs' Responsiveness to Human Feedback

2024-10-12 · Youquan Li, Miao Zheng, Fan Yang, Guosheng Dong, Bin Cui, WeiPeng Chen, Zenan Zhou, Wentao Zhang

Human feedback is crucial in the interactions between humans and Large Language Models (LLMs). However, existing research primarily focuses on benchmarking LLMs in single-turn dialogues. Even in benchmarks designed for multi-turn dialogues, the user inputs are often independent, neglecting the nuanced and complex nature of human feedback within real-world usage scenarios. To fill this research gap, we introduce FB-Bench, a fine-grained, multi-task benchmark designed to evaluate LLMs' responsiveness to human feedback in real-world usage scenarios. Drawing from the two main interaction scenarios, FB-Bench comprises 734 meticulously curated samples, encompassing eight task types, five deficiency types of response, and nine feedback types. We extensively evaluate a broad array of popular LLMs, revealing significant variations in their performance across different interaction scenarios. Further analysis indicates that task, human feedback, and deficiencies of previous responses can also significantly impact LLMs' responsiveness. Our findings underscore both the strengths and limitations of current models, providing valuable insights and directions for future research. Both the toolkits and the dataset of FB-Bench are available at https://github.com/PKU-Baichuan-MLSystemLab/FB-Bench.

📄 PDF Abstract BibTeX arXiv:2410.09412

Code (1)

pku-baichuan-mlsystemlab/fb-bench 공식 구현

Tasks

Benchmarking

Similar Papers 제목 키워드 기반

MMDocBench: Benchmarking Large Vision-Language Models for Fine-Grained Visual Document Understanding

2024-10-25 · Fengbin Zhu, Ziyang Liu, Xiang Yao Ng, Haohui Wu 외

Large Vision-Language Models (LVLMs) have achieved remarkable performance in many vision-language tasks, yet their capabilities in fine-grained visual understanding remain insufficiently evaluated. Existing benchmarks ei…

Benchmarkingdocument understandingOptical Character Recognition (OCR)

JourneyBench: A Challenging One-Stop Vision-Language Understanding Benchmark of Generated Images

2024-09-19 · Zhecan Wang, Junzhang Liu, Chia-Wei Tang, Hani AlOmari 외

Existing vision-language understanding benchmarks largely consist of images of objects in their usual contexts. As a consequence, recent multimodal large language models can perform well with only a shallow visual unders…

HallucinationImage CaptioningMultimodal ReasoningVisual Question Answering (VQA)+1

African or European Swallow? Benchmarking Large Vision-Language Models for Fine-Grained Object Classification

2024-06-20 · Gregor Geigle, Radu Timofte, Goran Glavaš

Recent Large Vision-Language Models (LVLMs) demonstrate impressive abilities on numerous image understanding and reasoning tasks. The task of fine-grained object classification (e.g., distinction between \textit{animal s…

BenchmarkingClassificationMultiple-choiceObject+1

MARINER: A 3E-Driven Benchmark for Fine-Grained Perception and Complex Reasoning in Open-Water Environments

2026-04-09 · Xingming Liao, Ning Chen, Muying Shu, Yunpeng Yin 외 arxiv

Fine-grained visual understanding and high-level reasoning in real-world open-water environments remain under-explored due to the lack of dedicated benchmarks. We introduce MARINER, a comprehensive benchmark built under …

Visual Question AnsweringObject Detection

FINER: MLLMs Hallucinate under Fine-grained Negative Queries

2026-03-18 · Rui Xiao, Sanghwan Kim, Yongqin Xian, Zeynep Akata 외 arxiv

Multimodal large language models (MLLMs) struggle with hallucinations, particularly with fine-grained queries, a challenge underrepresented by existing benchmarks that focus on coarse image-related questions. We introduc…