paper-with-me

Papers

Parameterized Argumentation-based Reasoning Tasks for Benchmarking Generative Language Models

2025-05-02 · Cor Steging, Silja Renooij, Bart Verheij

Generative large language models as tools in the legal domain have the potential to improve the justice system. However, the reasoning behavior of current generative models is brittle and poorly understood, hence cannot be responsibly applied in the domains of law and evidence. In this paper, we introduce an approach for creating benchmarks that can be used to evaluate the reasoning capabilities of generative language models. These benchmarks are dynamically varied, scalable in their complexity, and have formally unambiguous interpretations. In this study, we illustrate the approach on the basis of witness testimony, focusing on the underlying argument attack structure. We dynamically generate both linear and non-linear argument attack graphs of varying complexity and translate these into reasoning puzzles about witness testimony expressed in natural language. We show that state-of-the-art large language models often fail in these reasoning puzzles, already at low complexity. Obvious mistakes are made by the models, and their inconsistent performance indicates that their reasoning capabilities are brittle. Furthermore, at higher complexity, even state-of-the-art models specifically presented for reasoning capabilities make mistakes. We show the viability of using a parametrized benchmark with varying complexity to evaluate the reasoning capabilities of generative language models. As such, the findings contribute to a better understanding of the limitations of the reasoning capabilities of generative models, which is essential when designing responsible AI systems in the legal domain.

📄 PDF Abstract BibTeX arXiv:2505.01539

Code (1)

corsteging/parameterizedargumentationbasedreasoningtasks 공식 구현

Tasks

Benchmarking

Similar Papers 제목 키워드 기반

ArgBench: Benchmarking LLMs on Computational Argumentation Tasks

2026-04-19 · Yamen Ajjour, Carlotta Quensel, Nedim Lipka, Henning Wachsmuth arxiv

Argumentation skills are an essential toolkit for large language models (LLMs). These skills are crucial in various use cases, including self-reflection, debating collaboratively for diverse answers, and countering hate …

Counting Complexity for Reasoning in Abstract Argumentation

2018-11-28 · Johannes K. Fichte, Markus Hecher, Arne Meier

In this paper, we consider counting and projected model counting of extensions in abstract argumentation for various semantics. When asking for projected counts we are interested in counting the number of extensions of a…

Abstract Argumentation

Parameterized Complexity of Logic-Based Argumentation in Schaefer's Framework

2021-02-23 · Yasir Mahmood, Arne Meier, Johannes Schmidt

Logic-based argumentation is a well-established formalism modelling nonmonotonic reasoning. It has been playing a major role in AI for decades, now. Informally, a set of formulas is the support for a given claim if it is…

Harnessing Incremental Answer Set Solving for Reasoning in Assumption-Based Argumentation

2021-08-09 · Tuomo Lehtonen, Johannes P. Wallner, Matti Järvisalo

Assumption-based argumentation (ABA) is a central structured argumentation formalism. As shown recently, answer set programming (ASP) enables efficiently solving NP-hard reasoning tasks of ABA in practice, in particular …

An Argumentation-Based Legal Reasoning Approach for DL-Ontology

2022-09-07 · Zhe Yu, Yiwei Lu

Ontology is a popular method for knowledge representation in different domains, including the legal domain, and description logics (DL) is commonly used as its description language. To handle reasoning based on inconsist…

Autonomous VehiclesLegal Reasoning