SearcharxivSearch

arXiv subjects

Pratheepa Jeganathan

Publications and source records attributed to Pratheepa Jeganathan.

8 recordsLinked to original sources

Penalized Copula Mixed Models for Intercompany Loss Reserving and Risk Capital

Intercompany loss reserving provides an opportunity to improve reserve estimation by borrowing information across insurers while accounting for company-level heterogeneity. We propose a penalized generalized copula mixed model for multivariate loss reserving and risk capital analysis using multiple companies' loss triangles. The framework combines mixed-effects marginal models with company-specific copula dependence parameters, allowing residual dependence between lines of business to vary across insurers. Penalization is introduced through an $L_1$ penalty on the fixed effects to stabilize estimation in the tail of the loss triangles, where observations are limited. Estimation is carried out using an iterative two-stage procedure that combines likelihood-based estimation of the marginal mixed models with rank-based copula estimation using residual pseudo-observations. To obtain predictive reserve distributions, we develop a modified bootstrap procedure that accommodates penalized estimation while preserving the dependence structure. Using Schedule P data from the National Association of Insurance Commissioners, we show that the proposed model provides a more stable decomposition of unpaid losses across lines of business, reduces predictive variability relative to silo and fixed-effect copula benchmarks, and leads to lower risk capital after accounting for diversification. A simulation study further evaluates parameter recovery, sparsity selection, reserve accuracy, and robustness to random-effect misspecification. Overall, the proposed model offers an interpretable and flexible framework for intercompany multivariate reserving and capital assessment.

stat.ME

A Robust Nonparametric Framework for Detecting Repeated Spatial Patterns

Identifying spatially contiguous clusters and repeated spatial patterns (RSP) characterized by similar underlying distributions that are spatially apart is a key challenge in modern spatial statistics. Existing constrained clustering methods enforce spatial contiguity but are limited in their ability to identify RSP. We propose a novel nonparametric framework that addresses this limitation by combining constrained clustering with a post-clustering reassigment step based on the maximum mean discrepancy (MMD) statistic. We employ a block permutation strategy within each cluster that preserves local attribute structure when approximating the null distribution of the MMD. We also show that the MMD$^2$ statistic is asymptotically consistent under second-order stationarity and spatial mixing conditions. This two-stage approach enables the detection of clusters that are both spatially distant and similar in distribution. Through simulation studies that vary spatial dependence, cluster sizes, shapes, and multivariate dimensionality, we demonstrate the robustness of our proposed framework in detecting RSP. We further illustrate its applicability through an analysis of spatial proteomics data from patients with triple-negative breast cancer. Overall, our framework presents a methodological advancement in spatial clustering, offering a flexible and robust solution for spatial datasets that exhibit repeated patterns.

stat.ME

A Bayesian Framework for Post-disruption Travel Time Prediction in Metro Networks

Disruptions are an inherent feature of transportation systems, occurring unpredictably and with varying durations. Even after an incident is reported as resolved, disruptions can induce irregular train operations that generate substantial uncertainty in passenger waiting and travel times. Accurately forecasting post-disruption travel times therefore remains a critical challenge for transit operators and passenger information systems. This paper develops a Bayesian spatiotemporal modeling framework for post-disruption train travel times that explicitly captures train interactions, headway imbalance, and non-Gaussian distributional characteristics observed during recovery periods. The proposed model decomposes travel times into delay and journey components and incorporates a moving-average error structure to represent dependence between consecutive trains. Skew-normal and skew-$t$ distributions are employed to flexibly accommodate heteroskedasticity, skewness, and heavy-tailed behavior in post-disruption travel times. The framework is evaluated using high-resolution track-occupancy and disruption log data from the Montréal metro system, covering two lines in both travel directions. Empirical results indicate that post-disruption travel times exhibit pronounced distributional asymmetries that vary with traveled distance, as well as significant error dependence across trains. The proposed models consistently outperform baseline specifications in both point prediction accuracy and uncertainty quantification, with the skew-$t$ model demonstrating the most robust performance for longer journeys. These findings underscore the importance of incorporating both distributional flexibility and error dependence when forecasting post-disruption travel times in urban rail systems.

stat.AP

Scalable Spatiotemporal Modeling for Bicycle Count Prediction

We propose a novel sparse spatiotemporal dynamic generalized linear model for efficient inference and prediction of bicycle count data. Assuming Poisson distributed counts with spacetime-varying rates, we model the log-rate using spatiotemporal intercepts, dynamic temporal covariates, and site-specific effects additively. Spatiotemporal dependence is modeled using a spacetime-varying intercept that evolves smoothly over time with spatially correlated errors, and coefficients of some temporal covariates including seasonal harmonics also evolve dynamically over time. Inference is performed following the Bayesian paradigm, and uncertainty quantification is naturally accounted for when predicting bicycle counts for unobserved locations and future times of interest. To address the challenges of high-dimensional inference of spatiotemporal data in a Bayesian setting, we develop a customized hybrid Markov Chain Monte Carlo (MCMC) algorithm. To address the computational burden of dense covariance matrices, we extend our framework to high-dimensional spatial settings using the sparse SPDE approach of Lindgren et al. (2011), demonstrating its accuracy and scalability on both synthetic data and Montreal Island bicycle datasets. The proposed approach naturally provides missing value imputations, kriging, future forecasting, spatiotemporal predictions, and inference of model components. Moreover, it provides ways to predict average annual daily bicycles (AADB), a key metric often sought when designing bicycle networks.

stat.ME

Recurrent Neural Networks for Multivariate Loss Reserving and Risk Capital Analysis

In the property and casualty (P&C) insurance industry, reserves comprise most of a company's liabilities. These reserves are the best estimates made by actuaries for future unpaid claims. Notably, reserves for different lines of business (LOBs) are related due to dependent events or claims. While the actuarial industry has developed both parametric and non-parametric methods for loss reserving, only a few tools have been developed to capture dependence between loss reserves. This paper introduces the use of the Deep Triangle (DT), a recurrent neural network, for multivariate loss reserving, incorporating an asymmetric loss function to combine incremental paid losses of multiple LOBs. The input and output to the DT are the vectors of sequences of incremental paid losses that account for the pairwise and time dependence between and within LOBs. In addition, we extend generative adversarial networks (GANs) by transforming the two loss triangles into a tabular format and generating synthetic loss triangles to obtain the predictive distribution for reserves. We call the combination of DT for multivariate loss reserving and GAN for risk capital analysis the extended Deep Triangle (EDT). To illustrate EDT, we apply and calibrate these methods using data from multiple companies from the National Association of Insurance Commissioners database. For validation, we compare EDT to the copula regression models and find that the EDT outperforms the copula regression models in predicting total loss reserve. Furthermore, with the obtained predictive distribution for reserves, we show that risk capitals calculated from EDT are smaller than that of the copula regression models, suggesting a more considerable diversification benefit. Finally, these findings are also confirmed in a simulation study.

stat.AP

Microbiome Intervention Analysis with Transfer Functions and Mirror Statistics

Microbiome interventions provide valuable data about microbial ecosystem structure and dynamics. Despite their ubiquity in microbiome research, few rigorous data analysis approaches are available. In this study, we extend transfer function-based intervention analysis to the microbiome setting, drawing from advances in statistical learning and selective inference. Our proposal supports the simulation of hypothetical intervention trajectories and False Discovery Rate-guaranteed selection of significantly perturbed taxa. We explore the properties of our approach through simulation and re-analyze three contrasting microbiome studies. An R package, mbtransfer, is available at https://go.wisc.edu/crj6k6. Notebooks to reproduce the simulation and case studies can be found at https://go.wisc.edu/dxuibh and https://go.wisc.edu/emxv33.

stat.AP

A Statistical Perspective on the Challenges in Molecular Microbial Biology

High throughput sequencing (HTS)-based technology enables identifying and quantifying non-culturable microbial organisms in all environments. Microbial sequences have enhanced our understanding of the human microbiome, the soil and plant environment, and the marine environment. All molecular microbial data pose statistical challenges due to contamination sequences from reagents, batch effects, unequal sampling, and undetected taxa. Technical biases and heteroscedasticity have the strongest effects, but different strains across subjects and environments also make direct differential abundance testing unwieldy. We provide an introduction to a few statistical tools that can overcome some of these difficulties and demonstrate those tools on an example. We show how standard statistical methods, such as simple hierarchical mixture and topic models, can facilitate inferences on latent microbial communities. We also review some nonparametric Bayesian approaches that combine visualization and uncertainty quantification. The intersection of molecular microbial biology and statistics is an exciting new venue. Finally, we list some of the important open problems that would benefit from more careful statistical method development.

stat.AP

The Block Bootstrap Method for Longitudinal Microbiome Data

Microbial ecology serves as a foundation for a wide range of scientific and biomedical studies. Rapidly-evolving high-throughput sequencing technology enables the comprehensive search for microbial biomarkers using longitudinal experiments. Such experiments consist of repeated biological observations from each subject over time and are essential in accounting for the high between-subject and within-subject variability. Unfortunately, many of the statistical tests based on parametric models rely on correctly specifying temporal dependence structure which is unavailable in most microbiome data. In this paper, we propose an extension of the nonparametric bootstrap method that enables inference on these types longitudinal data. The proposed moving block bootstrap (MBB) method accounts for within-subject dependency by using overlapping blocks of repeated observations within each subject to draw valid inferences based on approximately pivotal statistics. Our simulation studies show an increase in power compared to merge-by-subject (MBS) strategies. We also show that compared to tests that presume independent samples (PIS), our proposed method reduces false microbial biomarker discovery rates. In this paper, we illustrated the MBB method using three different pregnancy data and an oral microbiome data. We provide an open-source R package https://github.com/PratheepaJ/bootLong to make our method accessible and the study in this paper reproducible.

stat.ME