SearcharxivSearch

arXiv · 2608.17715

Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models

Abstract

Credit decisioning is a high-stakes task in which model outputs must be accurate and explainable to support compliant decisions. Although modern credit risk models such as eXtreme Gradient Boosting (XGBoost) and Graph Neural Networks (GNNs) improve predictive performance, their explanations are often too technical for stakeholders creating communication gaps that can shape approvals, denials, and fairness judgments. We examine whether Large Language Models (LLMs) can serve as explanation layers that translate post-hoc explanation artefacts into stakeholder-appropriate risk narratives. Using Freddie Mac single-family loan-level data, we develop three pipelines: standard tabular (XGBoost + SHAP), and two with alternative data, a pure network-based (GNN + GNNExplainer), and a bimodal one (combining tabular and network data). We generate narratives with three LLM configurations: a small fine-tuned LLM (Gemma 3 4B), a large fine-tuned LLM (DeepSeek R1 70B), and a zero-shot commercial LLM (Gemini 2.5). Explanation quality is evaluated through automated checks across all pipelines and a human study of bimodal explanations comparing credit risk professionals and non-professionals on eight decision-relevant dimensions. We have three main findings. First, the pipeline accounts for higher variance in evidence-grounding scores than the language model, meaning that the binding constraint on explanation quality is the evidence representation, not the model used. Second, the explanation narratives reliably name the influential factors but are less reliable when stating the direction of influence, which may be consequential for adverse-action communication. Finally, professionals apply stricter evidentiary standards than non-professionals. We discuss implications for the governance of risk models, including deployment considerations and the value of domain-aligned LLMs in regulated credit settings.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Sahab Zandi, Noah Kostesku, Christophe Mues, María Óskarsdóttir, Cristián Bravo. 2026-08-18. Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models. https://arxiv.org/abs/2608.17715

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Simplifying Cyber Cat(astrophe)s with Cyber Kittens: Power Law Plausibility for Cyber Insurance Risks

Cyber insurance requires accurate modeling of worst-case catastrophic (cat) events, but the field lacks robust quantitative approaches for estimating upper-bound losses. Building on a recent dataset of 24 cyber cat events over 30 years, this work tests whether cyber economic losses follow a power law distribution. We analyze "cyber kittens" - sub-1B USD events distinguished from cat events (1B+ USD) only by magnitude - extracted via LLM from cyber insurance claims data (2020-2024). Using victim count (weighted by claim year) as a proxy for economic loss, we link kitten-sized events to known cat events to estimate losses. The kitten distribution proved consistent with the cat dataset, and power laws were statistically plausible: each order-of-magnitude increase in event size corresponds to a 5-7x drop in probability. Extrapolating, an event 100x the largest 2020-2024 cat event is expected roughly every 206 years, translating to 100-250B USD in losses - catastrophic, but not extraordinary relative to other insurance lines.

q-fin.RM

Pricing the DeFi Tail: Do Protocols or Depositors Price Operational Risk?

Similar to banks, DeFi protocols expose depositors to operational risk (USD 9.45 billion across 1,075 events since 2020). Unlike banks, they are not required to hold capital against it. A protocol may maintain a buffer voluntarily. Absent one, the risk falls on the depositor, who should then demand a risk premium in the supply yield. I quantify the underlying tail on one benchmark, a per-sector Basel loss-distribution approach fitted to a new operational risk event dataset, and test both margins against it. Tails in the four core sectors are no heavier than the Moscadelli banking band $[0.85, 1.39]$. Bridge, Derivatives, and the residual Other sector exhibit cyber-loss-level tails ($\hat\xi \approx 1.6$), with point estimates past the infinite-mean boundary. The Lending tail implies a $\mathrm{VaR}_{99.9}$ capital buffer of 18% of TVL and of the ten largest Lending venues, the four holding a buffer cover on average 5% of it. Under market discipline, depositors should demand a higher yield in compensation where a venue does not maintain a buffer. I find that venues without a buffer pay a higher premium than those with (a 125-bps gap in medians): evidence the market discriminates in the right direction. However, the premium falls far short of an adequately priced tail. This unpriced tail falls disproportionately on the retail depositor, who sees only the posted rate but lacks the information and skills to price it. Because these products are not bank-regulated, I recommend disclosure over capital mandates: protocols, and any service providers that front access to it, should publish standardized losses, existing capital buffers and tail coverage.

q-fin.RM

Illiquidity at Risk

Market efficiency relies fundamentally on stable liquidity. Consequently, forecasting liquidity dynamics is a priority for both investors and regulators. We introduce a new tail-risk metric, Illiquidity-at-Risk (IlliQaR), designed to quantify the magnitude of extreme liquidity dry-ups. Relying upon the realized Amihud (a precise illiquidity measurement derived from high-frequency data as the ratio of realized volatility to trading volume) we assess the predictive power of various linear and non-linear econometric models, with a specific focus on the impact of discontinuous jump components. Accounting for these jumps is essential for achieving accurate probability coverage and better IlliQaR predictions during periods of systemic stress, where standard continuous models systematically underestimate the severity of liquidity evaporation. Our empirical analysis, encompassing the S&P 500 index and a cross-section of 25 large U.S. equities, demonstrates that incorporating jumps significantly improves forecasts of illiquidity. Our results suggest that individual stock IlliQaR violations often cluster during periods of S&P 500 liquidity stress. This indicates that Illiquidity at Risk is not just a localized concern but a systemic one, where the main index acts as a leading indicator for extreme dry-ups in individual stock liquidity.

q-fin.RM