paper-with-me

홈 › Papers

Logical Judges Challenge Human Judges on the Strange Case of B.C.-Valjean

2020-09-22 · Viviana Mascardi, Domenico Pellegrini

On May 12th, 2020, during the course entitled Artificial Intelligence and Jurisdiction Practice organized by the Italian School of Magistracy, more than 70 magistrates followed our demonstration of a Prolog logical judge reasoning on an armed robbery case. Although the implemented logical judge is just an exercise of knowledge representation and simple deductive reasoning, a practical demonstration of an automated reasoning tool to such a large audience of potential end-users represents a first and unique attempt in Italy and, to the best of our knowledge, in the international panorama. In this paper we present the case addressed by the logical judge - a real case already addressed by a human judge in 2015 - and the feedback on the demonstration collected from the attendees.

📄 PDF Abstract BibTeX arXiv:2010.05694

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

JudgeBench: A Benchmark for Evaluating LLM-based Judges

2024-10-16 · Sijun Tan, Siyuan Zhuang, Kyle Montgomery, William Y. Tang 외

LLM-based judges have emerged as a scalable alternative to human evaluation and are increasingly used to assess, compare, and improve models. However, the reliability of LLM-based judges themselves is rarely scrutinized.…

Math

Beyond Surface Judgments: Human-Grounded Risk Evaluation of LLM-Generated Disinformation

2026-04-08 · Zonghuan Xu, Xiang Zheng, Yutao Wu, Xingjun Ma arxiv

Large language models (LLMs) can generate persuasive narratives at scale, raising concerns about their potential use in disinformation campaigns. Assessing this risk ultimately requires understanding how readers receive …

Comparing Developer and LLM Biases in Code Evaluation

2026-03-25 · Aditya Mittal, Ryan Shar, Zichu Wu, Shyam Agarwal 외 arxiv

As LLMs are increasingly used as judges in code applications, they should be evaluated in realistic interactive settings that capture partial context and ambiguous intent. We present TRACE (Tool for Rubric Analysis in Co…

Can Machines Imitate Humans? Integrative Turing Tests for Vision and Language Demonstrate a Narrowing Gap

2022-11-23 · Mengmi Zhang, Giorgia Dellaferrera, Ankur Sikarwar, Caishun Chen 외

As AI algorithms increasingly participate in daily activities, it becomes critical to ascertain whether the agents we interact with are human or not. To address this question, we turn to the Turing test and systematicall…

Image Captioningobject-detectionObject Detection

Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

2024-06-18 · Aman Singh Thakur, Kartik Choudhary, Venkat Srinik Ramayapally, Sankaran Vaidyanathan 외

Offering a promising solution to the scalability challenges associated with human evaluation, the LLM-as-a-judge paradigm is rapidly gaining traction as an approach to evaluating large language models (LLMs). However, th…

TriviaQA