Searcharxiv⌕ Search

arXiv · 2610.07888

A Functional Representation of Credit Behavior for Probability of Default Modeling

Abstract

This paper proposes a framework for modeling probability of default via functional data analysis. By representing a series of credit variables as functions, we investigate whether intra-monthly information improves default predictability in linear models. We further show how a range of widely used variables, among them available funds, utilization rate, and overdraft, can all be derived from three quantities observed over time. Namely, 1) the type of each account, 2) the balance of the account, and 3) the size of the credit limit on the account. Retaining these quantities as continuous-time processes, rather than reducing them to monthly aggregated values, yields a continuous faithful representation of the borrower. When analyzing these three processes, we discovered that recurring events associated with the ordinal position among banking days created strong cyclical patterns. We therefore develop a relative time framework that aligns the recurring events across borrowers, ensuring that borrowers possess the same cyclical pattern, regardless of real time. We assess the framework using functional logistic regression. This approach accommodates the continuous representation while remaining closely related to a logistic regression model commonly used in credit risk practice, due to strict regulatory constraints. We show that, when equipped with an effective functional representation of transactional trajectories, the proposed model attains predictive performance on par with XGBoost while consistently outperforming logistic regression. Importantly, the model balances predictive performance of a machine learning model with the interpretability of linear default models, potentially enabling financial institutions to use the model, even under strict regulation.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jonas Brunholm, Bjarne Højgaard, Thomas D. Nielsen, Orimar Sauri. 2026-10-06. A Functional Representation of Credit Behavior for Probability of Default Modeling. https://arxiv.org/abs/2610.07888

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

When Is the Gini Loading More Prudent? Tail Structure and the Ordering of the Standard Deviation and the Gini Mean Difference

The standard deviation (SD) and the Gini mean difference (GMD) are the two canonical measures of variability used to load premiums, set risk margins and allocate capital, yet no universal ordering between them exists. We show that the comparison is \emph{equivalent} to asking whether the coefficient of variation of the spacing $|X-X'|$ generated by two independent copies of the risk exceeds unity, so that the exponential law -- whose spacing is again exponential -- is the universal knife-edge separating the two regimes. Reading the GMD as twice the maxiance, that is, as a second-order \emph{dual} moment in the sense of Yaari's dual theory, the problem becomes an explicit comparison of primal and dual second-order variability. We derive a closed-form representation of the mean excess function of the spacing in terms of the hazard and reverse hazard rates of $X$, and use it to prove that heavy-tailed behavior -- a decreasing hazard rate or an increasing reverse hazard rate -- yields SD dominance, whereas two-sided light tails yield GMD dominance; when both the hazard and the reverse hazard rate are monotone, equality characterizes the exponential law. The sufficient conditions for SD dominance are stable under mixing, those for GMD dominance under convolution, and both under tail truncation, which makes them operational in frailty and collective risk models. We classify the severity, lifetime and frequency distributions of actuarial practice accordingly, quantify the consequences for SD- and Gini-loaded premium principles and for Gini-type tail risk measures, and show that the sign of $\SD-\GMD$ across thresholds furnishes a simple diagnostic for tail aging.

q-fin.RM↗

Comparing two approaches for modelling the loss given default of credit cards: Run-off triangles vs regression

The use of run-off triangles (ROTs) is a common industry practice in estimating the loss given default (LGD) risk parameter when predicting credit losses in banking. We benchmark this industry practice using credit card data against a more sophisticated (though classical) regression-based approach, which is able to leverage various types of input variables in producing loan-level LGD-estimates. This regression-based approach can demonstrably recover the typical characteristics of the 'U-shaped' empirical LGD-distribution, which the ROT-based approach cannot do. First, we critically review the ROT-based approach and identify multiple demerits using data-driven diagnostics. We then estimate a two-stage regression-based LGD-model and favourably assess the model performance of each component (or 'stage'). Finally, we aggregate the LGD-estimates produced by each approach over time, and compare each time series to the mean empirical loss rate over time. The ROT-based aggregates diverge substantially from the empirical rate over most time periods, whilst the regression-based aggregates follow the empirical trends much closer. These results underscore the greater prediction accuracy of the regression-based LGD-model, relative to the ROT-based one. By implication, the former approach is probably better than the latter ROT-based approach when estimating the LGD under the IFRS 9 accounting framework, which prioritises accuracy.

q-fin.RM↗

Distributionally Robust Insurance under Bregman-Wasserstein Divergence

This paper investigates two optimal insurance contracting problems under distributional uncertainty from the perspective of a potential policyholder, utilizing a Bregman-Wasserstein (BW) ball to characterize the ambiguity set of loss distributions. The first problem examines an insurance demand model where the policyholder adopts an $α$-maxmin preference with Value-at-Risk (VaR). We derive the optimal indemnity function in closed form and study, both analytically and numerically, how the asymmetry inherent in BW divergence influences the optimal indemnity structure. The second problem employs a robust optimization framework, where the policyholder aims to secure robust insurance indemnity by minimizing the worst-case convex distortion risk measure while adhering to a guaranteed VaR constraint. In this context, we provide explicit characterizations of both the optimal indemnity and the worst-case distribution in closed form through a combined approach using the Lagrangian and relaxation methods. To illustrate the practical implications of our theoretical findings, we include a concrete example based on Tail Value-at-Risk (TVaR).

q-fin.RM↗