paper-with-me

Papers

Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models

2026-08-18 · Sahab Zandi, Noah Kostesku, Christophe Mues, María Óskarsdóttir, Cristián Bravo arxiv

Credit decisioning is a high-stakes task in which model outputs must be accurate and explainable to support compliant decisions. Although modern credit risk models such as eXtreme Gradient Boosting (XGBoost) and Graph Neural Networks (GNNs) improve predictive performance, their explanations are often too technical for stakeholders creating communication gaps that can shape approvals, denials, and fairness judgments. We examine whether Large Language Models (LLMs) can serve as explanation layers that translate post-hoc explanation artefacts into stakeholder-appropriate risk narratives. Using Freddie Mac single-family loan-level data, we develop three pipelines: standard tabular (XGBoost + SHAP), and two with alternative data, a pure network-based (GNN + GNNExplainer), and a bimodal one (combining tabular and network data). We generate narratives with three LLM configurations: a small fine-tuned LLM (Gemma 3 4B), a large fine-tuned LLM (DeepSeek R1 70B), and a zero-shot commercial LLM (Gemini 2.5). Explanation quality is evaluated through automated checks across all pipelines and a human study of bimodal explanations comparing credit risk professionals and non-professionals on eight decision-relevant dimensions. We have three main findings. First, the pipeline accounts for higher variance in evidence-grounding scores than the language model, meaning that the binding constraint on explanation quality is the evidence representation, not the model used. Second, the explanation narratives reliably name the influential factors but are less reliable when stating the direction of influence, which may be consequential for adverse-action communication. Finally, professionals apply stricter evidentiary standards than non-professionals. We discuss implications for the governance of risk models, including deployment considerations and the value of domain-aligned LLMs in regulated credit settings.

📄 PDF Abstract BibTeX arXiv:2608.17715

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Empowering Many, Biasing a Few: Generalist Credit Scoring through Large Language Models

2023-10-01 · Duanyu Feng, Yongfu Dai, Jimin Huang, Yifang Zhang 외

In the financial industry, credit scoring is a fundamental element, shaping access to credit and determining the terms of loans for individuals and businesses alike. Traditional credit scoring methods, however, often gra…

Decision MakingLanguage ModellingLarge Language Model

LendNova: Towards Automated Credit Risk Assessment with Language Models

2026-01-05 · Kiarash Shamsi, Danijel Novokmet, Joshua Peters, Mao Lin Liu 외 arxiv

Credit risk assessment is essential in the financial sector, but has traditionally depended on costly feature-based models that often fail to utilize all available information in raw credit records. This paper introduces…

Feature Engineering

Interpretable LLMs for Credit Risk: A Systematic Review and Taxonomy

2025-06-04 · Muhammed Golec, Maha AlabdulJalil

Large Language Models (LLM), which have developed in recent years, enable credit risk assessment through the analysis of financial texts such as analyst reports and corporate disclosures. This paper presents the first sy…

LLMs Should Not Yet Be Credited with Decision Explanation

2026-05-01 · Wenshuo Wang arxiv

This position paper argues that LLMs should not yet be credited with decision explanation. This matters because recent work increasingly treats accurate behavioral prediction, plausible rationales, and outcome-conditione…

A Meta Path Based Evaluation Method for Enterprise Credit Risk

2021-10-22 · Marui Du, Yue Ma, Zuoquan Zhang

Nowadays small and medium-sized enterprises have become an essential part of the national economy. With the increasing number of such enterprises, how to evaluate their credit risk becomes a hot issue. Unlike big enterpr…