paper-with-me

홈 › Papers

Discovering the Rationale of Decisions: Experiments on Aligning Learning and Reasoning

2021-05-14 · Cor Steging, Silja Renooij, Bart Verheij

In AI and law, systems that are designed for decision support should be explainable when pursuing justice. In order for these systems to be fair and responsible, they should make correct decisions and make them using a sound and transparent rationale. In this paper, we introduce a knowledge-driven method for model-agnostic rationale evaluation using dedicated test cases, similar to unit-testing in professional software development. We apply this new method in a set of machine learning experiments aimed at extracting known knowledge structures from artificial datasets from fictional and non-fictional legal settings. We show that our method allows us to analyze the rationale of black-box machine learning systems by assessing which rationale elements are learned or not. Furthermore, we show that the rationale can be adjusted using tailor-made training data based on the results of the rationale evaluation.

📄 PDF Abstract BibTeX arXiv:2105.06758

Code (1)

CorSteging/DiscoveringTheRationaleOfDecisions 공식 구현

Similar Papers 제목 키워드 기반

ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails

2026-05-29 · Yan Wang, Zhixuan Chu, Zihao Xue, Zhen Bi 외 arxiv

Reasoning-based LLM guardrails improve safety moderation by generating explicit rationales before issuing final decisions. However, their rationales do not always lead to faithful enforcement: a model may recognize a har…

R-Align: Enhancing Generative Reward Models through Rationale-Centric Meta-Judging

2026-02-06 · Yanlin Lai, Mitt Huang, Hangyu Guo, Xiangfeng Wang 외 arxiv

Reinforcement Learning from Human Feedback (RLHF) remains indispensable for aligning large language models (LLMs) in subjective domains. To enhance robustness, recent work shifts toward Generative Reward Models (GenRMs) …

Reinforcement LearningInstruction Following

Aligning Deep Implicit Preferences by Learning to Reason Defensively

2025-10-13 · Peiming Li, Zhiyuan Hu, Yang Tang, Shiyu Li 외 arxiv

Personalized alignment is crucial for enabling Large Language Models (LLMs) to engage effectively in user-centric interactions. However, current methods face a dual challenge: they fail to infer users' deep implicit pref…

Reinforcement Learning

From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought

2025-07-01 · Wentao Tan, Qiong Cao, Yibing Zhan, Chao Xue 외 arxiv

Achieving human-like reasoning capabilities in Multimodal Large Language Models (MLLMs) has long been a goal. Current methods primarily focus on synthesizing positive rationales, typically relying on manual annotations o…

Multimodal Reasoning

How Ambiguous are the Rationales for Natural Language Reasoning? A Simple Approach to Handling Rationale Uncertainty

2024-02-22 · Hazel Kim

Rationales behind answers not only explain model decisions but boost language models to reason well on complex reasoning tasks. However, obtaining impeccable rationales is often impossible. Besides, it is non-trivial to …

Informativeness