paper-with-me

Papers

D2VBench: Benchmarking Large Language Models with Value Dilemmas in Daily Scenarios

2026-07-22 · Siyi Hao, Yidi Cao, Linhao Yu, Yuqi Ren, Deyi Xiong arxiv

With the wide application of large language models (LLMs) in real-world scenarios, the value implication of their outputs is crucial. However, existing evaluation benchmarks suffer from insufficient coverage of value dilemmas in daily scenarios involving multiple value conflicts and simplistic evaluation formalisms that fail to assess LLMs' value alignment. To address these issues, we propose D2VBench, a value alignment benchmark comprising 10,000 instances of real daily dilemma scenarios constructed through a multi-stage collaboration between LLMs and humans, grounded in 158 manually annotated fine-grained value concepts. For evaluation on the benchmark, we present a hybrid evaluation paradigm that integrates multiple-choice questions with open-ended questions. We conduct comprehensive evaluations on eight mainstream LLMs. Experimental results demonstrate that D2VBench exhibits high reliability and robustness, effectively reflecting the LLMs' alignment across different value categories and dimensions, and providing a more realistic and fine-grained tool for research on value alignment. The dataset is available at https://github.com/tjunlp-lab/D2VBench.

📄 PDF Abstract BibTeX arXiv:2607.19834

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CLASH: Evaluating Language Models on Judging High-Stakes Dilemmas from Multiple Perspectives

2025-04-15 · Ayoung Lee, Ryan Sungmo Kwon, Peter Railton, Lu Wang

Navigating high-stakes dilemmas involving conflicting values is challenging even for humans, let alone for AI. Yet prior work in evaluating the reasoning capabilities of large language models (LLMs) in such situations ha…

Benchmarking

UAVBench and UAVIT-1M: Benchmarking and Enhancing MLLMs for Low-Altitude UAV Vision-Language Understanding

2026-03-15 · Yang Zhan, Yuan Yuan arxiv

Multimodal Large Language Models (MLLMs) have made significant strides in natural images and satellite remote sensing images. However, understanding low-altitude drone scenarios remains a challenge. Existing datasets pri…

A Comparative Analysis on Ethical Benchmarking in Large Language Models

2024-10-11 · Kira Sam, Raja Vavekanand

This work contributes to the field of Machine Ethics (ME) benchmarking, which develops tests to assess whether intelligent systems accurately represent human values and act accordingly. We identify three major issues wit…

BenchmarkingDecision MakingEthicsQuestion Generation+1

Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities

2026-07-07 · So Hasegawa, Shailaja Keyur Sampat, Lei Liu, Wei-Peng Chen arxiv

Current benchmarks for evaluating Large Language Models (LLMs) in data analysis often fail to reflect real-world settings. They typically focus on fact retrieval from small tables and overlook the challenges of large mul…

Can LLMs Imagine Moral Alternatives Beyond Binary Dilemmas?

2026-06-30 · Jongchan Choi, Nari Yang, Sung Soo Park, Jaemin Cho 외 arxiv

As large language models (LLMs) are increasingly deployed as moral advisors and agents, they need to address dilemmas between two competing values. However, existing research on LLMs with moral dilemmas overlooks a centr…