SearcharxivSearch

arXiv subjects

X. Sheldon Lin

Publications and source records attributed to X. Sheldon Lin.

9 recordsLinked to original sources

A Portfolio-Anchored Frequency-Severity Risk Index for Trip and Driver Assessment Using Telematics Signals

In this paper, we propose a novel frequency-severity joint trip-level risk index that combines the frequency of abnormal driving patterns with a severity component reflecting how extreme such behavior is relative to a portfolio-level baseline. Severity is quantified through an inverse-probability penalty that increases with the rarity of observed tail extremes, rather than being interpreted as a claim size. Based on high-frequency telematics data, we construct a multi-scale representation of longitudinal acceleration using the maximal overlap discrete wavelet transform (MODWT), which preserves localized driving patterns across multiple time scales. To capture severity as tail rarity, we model the portfolio distribution using a Gaussian-Uniform mixture with a layered tail structure, where Gaussian components describe typical driving behavior and the tail is partitioned into ordered severity layers that reflect increasing extremeness. We develop a likelihood-based estimation procedure that makes inference feasible for this mixture model. The resulting severity layers are then used to construct multi-layer tail counts (MLTC) at the trip level, which are modeled within a Poisson-Gamma framework to yield a closed-form posterior risk index that jointly reflects frequency and severity. This conjugate structure naturally supports sequential updating, enabling the construction of dynamically evolving driver-level risk profiles. Using the UAH-DriveSet controlled dataset, we demonstrate that the proposed index enables reliable discrimination across behavioral driving states, identification of high-risk trips, and coherent ranking of drivers, yielding a purely behavior-driven risk measure suitable for actuarial ratemaking and potentially mitigating fairness concerns associated with traditional covariates.

stat.AP

Static marginal expected shortfall: Systemic risk measurement under dependence uncertainty

Measuring the contribution of a bank or an insurance company to overall systemic risk is a key concern, particularly in the aftermath of the 2007--2009 financial crisis and the 2020 downturn. In this paper, we derive worst-case and best-case bounds for the marginal expected shortfall (MES) -- a key measure of systemic risk contribution -- under the assumption that individual firms' risk distributions are known but their dependence structure is not. We further derive tighter MES bounds when partial information on companies' risk exposures, and thus their dependence, is available. To represent this partial information, we employ three standard factor models: additive, minimum-based, and multiplicative background risk models. Additionally, we propose an alternative set of improved MES bounds based on a linear relationship between firm-specific and market-wide risks, consistent with the Capital Asset Pricing Model in finance and the Weighted Insurance Pricing Model in insurance. Finally, empirical analyses demonstrate the practical relevance of the theoretical bounds for industry practitioners and policymakers.

q-fin.RM

A Population Sampling Framework for Claim Reserving in General Insurance

Claim reserving in insurance has been studied through two primary frameworks: the macro-level approach, which estimates reserves at an aggregate level (e.g., Chain-Ladder), and the micro-level approach, which estimates reserves at the individual claim level Antonio and Plat (2014). These frameworks are based on fundamentally different theoretical foundations, creating a degree of incompatibility that limits the adoption of more flexible models. This paper introduces a unified statistical framework for claim reserving, grounded in population sampling theory. We show that macro- and micro-level models represent extreme yet natural cases of an augmented inverse probability weighting (AIPW) estimator. This formulation allows for a seamless integration of principles from both aggregate and individual models, enabling more accurate and flexible estimations. Moreover, this paper also addresses critical issues of sampling bias arising from partially observed claims data-an often overlooked challenge in insurance. By adapting advanced statistical methods from the sampling literature, such as double-robust estimators, weighted estimating equations, and synthetic data generation, we improve predictive accuracy and expand the tools available for actuaries. The framework is illustrated using Canadian auto insurance data, highlighting the practical benefits of the sampling-based methods.

stat.ME

Assessing Driving Risk Through Unsupervised Detection of Anomalies in Telematics Time Series Data

Vehicle telematics provides granular data for dynamic driving risk assessment, but current methods often rely on aggregated metrics (e.g., harsh braking counts) and do not fully exploit the rich time-series structure of telematics data. In this paper, we introduce a flexible framework using continuous-time hidden Markov model (CTHMM) to model and analyze trip-level telematics data. Unlike existing methods, the CTHMM models raw time-series data without predefined thresholds on harsh driving events or assumptions about accident probabilities. Moreover, our analysis is based solely on telematics data, requiring no traditional covariates such as driver or vehicle characteristics. Through unsupervised anomaly detection based on pseudo-residuals, we identify deviations from normal driving patterns -- defined as the prevalent behaviour observed in a driver's history or across the population -- which are linked to accident risk. Validated on both controlled and real-world datasets, the CTHMM effectively detects abnormal driving behaviour and trips with increased accident likelihood. In real data analysis, higher anomaly levels in longitudinal and lateral accelerations consistently correlate with greater accident risk, with classification models using this information achieving ROC-AUC values as high as 0.86 for trip-level analysis and 0.78 for distinguishing drivers with claims. Furthermore, the methodology reveals significant behavioural differences between drivers with and without claims, offering valuable insights for insurance applications, accident analysis, and prevention.

stat.AP

Claim Reserving via Inverse Probability Weighting: A Micro-Level Chain-Ladder Method

Claim reserving primarily relies on macro-level models, with the Chain-Ladder method being the most widely adopted. These methods were heuristically developed without minimal statistical foundations, relying on oversimplified data assumptions and neglecting policyholder heterogeneity, often resulting in conservative reserve predictions. Micro-level reserving, utilizing stochastic modeling with granular information, can improve predictions but tends to involve less attractive and complex models for practitioners. This paper aims to strike a practical balance between aggregate and individual models by introducing a methodology that enables the Chain-Ladder method to incorporate individual information. We achieve this by proposing a novel framework, formulating the claim reserving problem within a population sampling context. We introduce a reserve estimator in a frequency and severity distribution-free manner that utilizes inverse probability weights (IPW) driven by individual information, akin to propensity scores. We demonstrate that the Chain-Ladder method emerges as a particular case of such an IPW estimator, thereby inheriting a statistically sound foundation based on population sampling theory that enables the use of granular information, and other extensions.

econ.EM

Data Mining of Telematics Data: Unveiling the Hidden Patterns in Driving Behaviour

With the advancement in technology, telematics data which capture vehicle movements information are becoming available to more insurers. As these data capture the actual driving behaviour, they are expected to improve our understanding of driving risk and facilitate more accurate auto-insurance ratemaking. In this paper, we analyze an auto-insurance dataset with telematics data collected from a major European insurer. Through a detailed discussion of the telematics data structure and related data quality issues, we elaborate on practical challenges in processing and incorporating telematics information in loss modelling and ratemaking. Then, with an exploratory data analysis, we demonstrate the existence of heterogeneity in individual driving behaviour, even within the groups of policyholders with and without claims, which supports the study of telematics data. Our regression analysis reiterates the importance of telematics data in claims modelling; in particular, we propose a speed transition matrix that describes discretely recorded speed time series and produces statistically significant predictors for claim counts. We conclude that large speed transitions, together with higher maximum speed attained, nighttime driving and increased harsh braking, are associated with increased claim counts. Moreover, we empirically illustrate the learning effects in driving behaviour: we show that both severe harsh events detected at a high threshold and expected claim counts are not directly proportional with driving time or distance, but they increase at a decreasing rate.

stat.AP

Effective experience rating for large insurance portfolios via surrogate modeling

Experience rating in insurance uses a Bayesian credibility model to upgrade the current premiums of a contract by taking into account policyholders' attributes and their claim history. Most data-driven models used for this task are mathematically intractable, and premiums must be obtained through numerical methods such as simulation via MCMC. However, these methods can be computationally expensive and even prohibitive for large portfolios when applied at the policyholder level. Additionally, these computations become ``black-box" procedures as there is no analytical expression showing how the claim history of policyholders is used to upgrade their premiums. To address these challenges, this paper proposes a surrogate modeling approach to inexpensively derive an analytical expression for computing the Bayesian premiums for any given model, approximately. As a part of the methodology, the paper introduces a \emph{likelihood-based summary statistic} of the policyholder's claim history that serves as the main input of the surrogate model and that is sufficient for certain families of distribution, including the exponential dispersion family. As a result, the computational burden of experience rating for large portfolios is reduced through the direct evaluation of such analytical expression, which can provide a transparent and interpretable way of computing Bayesian premiums.

stat.ME

A Posteriori Risk Classification and Ratemaking with Random Effects in the Mixture-of-Experts Model

A well-designed framework for risk classification and ratemaking in automobile insurance is key to insurers' profitability and risk management, while also ensuring that policyholders are charged a fair premium according to their risk profile. In this paper, we propose to adapt a flexible regression model, called the Mixed LRMoE, to the problem of a posteriori risk classification and ratemaking, where policyholder-level random effects are incorporated to better infer their risk profile reflected by the claim history. We also develop a stochastic variational Expectation-Conditional-Maximization algorithm for estimating model parameters and inferring the posterior distribution of random effects, which is numerically efficient and scalable to large insurance portfolios. We then apply the Mixed LRMoE model to a real, multiyear automobile insurance dataset, where the proposed framework is shown to offer better fit to data and produce posterior premium which accurately reflects policyholders' claim history.

stat.AP

A Marked Cox Model for IBNR Claims: Model and Theory

Incurred but not reported (IBNR) loss reserving is an important issue for Property & Casualty (P&C) insurers. The modeling of the claim arrival process, especially its temporal dependence, has not been closely examined in many of the current loss reserving models. In this paper, we propose modeling the claim arrival process together with its reporting delays as a marked Cox process. Our model is versatile in modeling temporal dependence, allowing also for natural interpretations. This paper focuses mainly on the theoretical aspects of the proposed model. We show that the associated reported claim process and IBNR claim process are both marked Cox processes with easily convertible intensity functions and marking distributions. The proposed model can also account for fluctuations in the exposure. By an order statistics property, we show that the corresponding discretely observed process preserves all the information about the claim arrival epochs. Finally, we derive closed-form expressions for both the autocorrelation function (ACF) and the distributions of the numbers of reported claims and IBNR claims. Model estimation and its applications are considered in a subsequent paper, Badescu et al.(2015b)

stat.AP