SearcharxivSearch

arXiv subjects

Jiakun Jiang

Publications and source records attributed to Jiakun Jiang.

2 recordsLinked to original sources

Sparse Convex Biclustering

Biclustering is an essential unsupervised machine learning technique for simultaneously clustering rows and columns of a data matrix, with widespread applications in genomics, transcriptomics, and other high-dimensional omics data. Despite its importance, existing biclustering methods struggle to meet the demands of modern large-scale datasets. The challenges stem from the accumulation of noise in high-dimensional features, the limitations of non-convex optimization formulations, and the computational complexity of identifying meaningful biclusters. These issues often result in reduced accuracy and stability as the size of the dataset increases. To overcome these challenges, we propose Sparse Convex Biclustering (SpaCoBi), a novel method that penalizes noise during the biclustering process to improve both accuracy and robustness. By adopting a convex optimization framework and introducing a stability-based tuning criterion, SpaCoBi achieves an optimal balance between cluster fidelity and sparsity. Comprehensive numerical studies, including simulations and an application to mouse olfactory bulb data, demonstrate that SpaCoBi significantly outperforms state-of-the-art methods in accuracy. These results highlight SpaCoBi as a robust and efficient solution for biclustering in high-dimensional and large-scale datasets.

stat.ML

Dynamic online prediction model and its application to automobile claim frequency data

Prediction modelling of claim frequency is an important task for pricing and risk management in non-life insurance and needed to be updated frequently with the changes in the insured population, regulatory legislation and technology. Existing methods are either done in an ad hoc fashion, such as parametric model calibration, or less so for the purpose of prediction. In this paper, we develop a Dynamic Poisson state space (DPSS) model which can continuously update the parameters whenever new claim information becomes available. DPSS model allows for both time-varying and time-invariant coefficients. To account for smoothness trends of time-varying coefficients over time, smoothing splines are used to model time-varying coefficients. The smoothing parameters are objectively chosen by maximum likelihood. The model is updated using batch data accumulated at pre-specified time intervals, which allows for a better approximation of the underlying Poisson density function. The proposed method can be also extended to the distributional assumption of zero-inflated Poisson and negative binomial. In the simulation, we show that the new model has significantly higher prediction accuracy compared to existing methods. We apply this methodology to a real-world automobile insurance claim data set in China over a period of six years and demonstrate its superiority by comparing it with the results of competing models from the literature.

stat.AP