SearcharxivSearch

arXiv subjects

Martin Bladt

Publications and source records attributed to Martin Bladt.

At least 19 recordsLinked to original sources

Censored Heteroscedastic Extremes

We study estimation of tail heterogeneity for non-identically distributed extreme observations subject to random right-censoring. In the uncensored setting, such heterogeneity is described by the event scedasis function, which measures the relative contribution of different design points to the upper tail. Under censoring, however, the observed tail heterogeneity is contaminated by the censoring scedasis functions, and applying uncensored techniques targets the wrong object. We propose a Beran-type estimator of the relative event scedasis, which is consistent under mild conditions. To obtain these results, survival analysis representations at an upper order statistics are extended to the non-identically distributed case; specifically, we develop conditional Nelson--Aalen and Beran theory on increasing intervals whose random endpoint is dominated, with probability tending to one, by a deterministic high local quantile. In particular, we derive a martingale array representation of the conditional Nelson--Aalen estimator with explicit error bounds depending only on the sample fraction and the bandwidth. Simulations demonstrate the finite-sample performance of the method, and an application to French property-casualty insurance claims illustrates how heterogeneous censoring can distort naive scedasis estimates.

math.ST

Survival Isotonic Distributional Regression

We introduce Survival-IDR (S-IDR), a nonparametric estimator of conditional survival distributions under order restrictions, extending Isotonic Distributional Regression (IDR; Henzi et al., 2021) to right-censored outcomes. S-IDR has no tuning parameters and accommodates continuous, discrete, and partially ordered covariates. We first study the direct Kaplan-Meier adaptation of IDR: it is uniformly consistent at the minimax rate, but only when the conditional outcomes are hazard-rate ordered. We trace this restriction to the Kaplan-Meier estimator's failure to satisfy the Cauchy mean value property on non-i.i.d. samples, and use the diagnosis to construct S-IDR. The S-IDR estimator is uniformly consistent under only stochastic dominance of the conditional outcomes, attains the minimax rate when the smoothness of the conditional CDFs is known, and admits a known cross-threshold PAVA acceleration. We further embed S-IDR in a distributional single-index framework on a benchmark suite, and apply it in a case study that validates the MELD score used for liver-transplant wait list management. Accompanying R, Python and Rust packages are available at https://github.com/AlexanderHenzi/isodistrreg.

math.ST

Payment Process Estimation in Aggregated Insurance Models

Insurance payments may depend on latent micro states although only macro states and realized payments are observed. We study a sojourn-payment model for such aggregated multi-state systems under left-truncation and right-censoring. Starting from a micro-to-macro projection, we establish strong consistency and weak convergence for inverse-probability-weighted estimators of state-specific cumulative payment processes.

stat.ME

Local estimation of transition rates of jump processes through discretization

We investigate the Poisson regression method for Markov and semi-Markov jump processes from a nonparametric angle, allowing the lengths of the time and duration intervals in the partition to vary with the number of observations. Imposing no structural assumptions on the true intensities, we obtain asymptotic normality of the occurence/exposure rates under appropriate shrinking conditions on the partition lengths. We derive asymptotic normality results for both Markov and semi-Markov models using only classical central limit theorems and elementary results for counting processes. All results are illustrated on both simulated and real data.

math.ST

Scoring Rules with Normalized Upper Order Statistics for Tail Inference

This paper proposes a scoring-rule-based method for ranking predictive distributions in the Fr\'echet domain that is able to distinguish between different tail indices. The approach is built on normalized order statistics and exploits proper scoring rules to compare tail limit distributions in a distributional framework, with direct relevance for insurance claim-severity tails. On the theoretical side, consistency and asymptotic normality for empirical tail scores based on normalized upper order statistics are obtained through residual estimation theory. Simulation results demonstrate that the scoring-rule-based approach is capable of discriminating between different tail behaviors in finite samples and that systematic scale variation has only a minor impact on stability. We further show that optimizing scoring rules yields consistent tail-index estimators and that the classical Hill estimator arises as a special case. The performance of the Energy Score tail-index estimator is investigated and compared with the Hill estimator across a range of tail indices. Lastly, we analyze an automobile claim-severity data set to demonstrate how scoring rules can be used to rank predictive models based on tail predictions in actuarial settings.

stat.ME

Cure models: from mixture to matrix distributions

Cure rate models address survival data in which a proportion of individuals will never experience the event of interest. Existing parametric approaches are predominantly based on finite mixtures, which impose restrictive assumptions on both the cure mechanism and the distribution of susceptible event times. A cure model based on phase-type distributions is introduced, leveraging their latent Markov jump process representation to allow immunity to occur either at baseline or dynamically during follow-up. This structure yields a flexible and interpretable formulation of long-term survival while encompassing classical mixture cure models as special cases. A unified regression framework is developed for covariate effects on both the cure rate and the susceptible survival distribution, and the proposed model class is dense, reducing the impact of parametric misspecification. Estimation is performed via expectation-maximization algorithms, accompanied by an automatic model selection strategy. Simulation studies and a real-data example demonstrate the practical advantages of the approach.

stat.ME

Consistency of Honest Decision Trees and Random Forests

We study various types of consistency of honest decision trees and random forests in the regression setting. In contrast to related literature, our proofs are elementary and follow the classical arguments used for smoothing methods. Under mild regularity conditions on the regression function and data distribution, we establish weak and almost sure convergence of honest trees and honest forest averages to the true regression function, and moreover we obtain uniform convergence over compact covariate domains. The framework naturally accommodates ensemble variants based on subsampling and also a two-stage bootstrap sampling scheme. Our treatment synthesizes and simplifies existing analyses, in particular recovering several results as special cases. The elementary nature of the arguments clarifies the close relationship between data-adaptive partitioning and kernel-type methods, providing an accessible approach to understanding the asymptotic behavior of tree-based methods.

stat.ME

Nonparametric Survival Estimation with Contaminated and Adjudicated Events

We study the conditional expert Kaplan-Meier estimator, an extension of the classical Kaplan--Meier estimator designed for time-to-event data subject to both right-censoring and contamination. Such contamination, where observed events may not reflect true outcomes, is common in applied settings, including insurance and credit risk, where expert opinion is often used to adjudicate uncertain events. Building on previous work, we develop a comprehensive asymptotic theory for the conditional version incorporating covariates through kernel smoothing. We establish functional consistency and weak convergence under suitable regularity conditions and quantify the bias induced by imperfect expert information. The results show that unbiased expert judgments ensure consistency, while systematic deviations lead to a deterministic asymptotic bias that can be explicitly characterized. We examine finite-sample properties through simulation studies and illustrate the practical use of the estimator with an application to loan default data.

stat.ME

Assessing continuous common-shock risk through matrix distributions

We introduce a class of continuous-time bivariate phase-type distributions for modeling dependencies from common shocks. The construction uses continuous-time Markov processes that evolve identically until an internal common-shock event, after which they diverge into independent processes. We derive and analyze key risk measures for this new class, including joint cumulative distribution functions, dependence measures, and conditional risk measures. Theoretical results establish analytically tractable properties of the model. For parameter estimation, we employ efficient gradient-based methods. Applications to both simulated and real-world data illustrate the ability to capture common-shock dependencies effectively. Our analysis also demonstrates that common-shock continuous phase-type distributions may capture dependencies that extend beyond those explicitly triggered by common shocks.

math.ST

Bayesian non-parametric survival estimation: stochastic hyperparameter sequences and distribution splicing

A Bayesian non-parametric framework for studying time-to-event data is proposed, where the prior distribution is allowed to depend on an additional random source, and may update with the sample size. Such scenarios are natural, for instance, when considering empirical Bayes techniques or dynamic expert information. In this context, a natural stochastic class for studying the cumulative hazard function are conditionally inhomogeneous independent increment processes with non-decreasing sample paths, also known as mixed time-inhomogeneous subordinators or mixed non-decreasing additive processes. The asymptotic behaviour is studied by showing that Bayesian consistency and Bernstein--von~Mises theorems may be recovered under suitable conditions on the asymptotic negligibility of the stochastic prior sequences. The non-asymptotic behaviour of the posterior is also considered. Namely, upon conditioning, an efficient and exact simulation algorithm for the paths of the Beta L\'evy process is provided. As a natural application, it is shown how the model can provide an appropriate definition of non-parametric spliced models. Spliced models target data where an accurate global description of both the body and tail of the distribution is desirable. The Bayesian non-parametric nature of the proposed estimators can offer conceptual and numerical alternatives to their parametric counterparts.

stat.ME

Non-parametric cure models through extreme-value tail estimation

In survival analysis, the estimation of the proportion of subjects who will never experience the event of interest, termed the cure rate, has received considerable attention recently. Its estimation can be a particularly difficult task when follow-up is not sufficient, that is when the censoring mechanism has a smaller support than the distribution of the target data. In the latter case, non-parametric estimators were recently proposed using extreme value methodology, assuming that the distribution of the susceptible population is in the Fr\'echet or Gumbel max-domains of attraction. In this paper, we take the extreme value techniques one step further, to jointly estimate the cure rate and the extreme value index, using probability plotting methodology, and in particular using the full information contained in the top order statistics. In other words, under sufficient or insufficient follow-up, we reconstruct the immune proportion. To this end, a Peaks-over-Threshold approach is proposed under the Gumbel max-domain assumption. Next, the approach is also transferred to more specific models such as Pareto, log-normal and Weibull tail models, allowing to recognize the most important tail characteristics of the susceptible population. We establish the asymptotic behavior of our estimators under regularization. Though simulation studies, our estimators are show to rival and often outperform established models, even when purely considering cure rate estimation. Finally, we provide an application of our method to Norwegian birth registry data.

math.ST

Conditional Extreme Value Estimation for Dependent Time Series

We study the consistency and weak convergence of the conditional tail function and conditional Hill estimators under broad dependence assumptions for a heavy-tailed response sequence and a covariate sequence. Consistency is established under $\alpha$-mixing, while asymptotic normality follows from $\beta$-mixing and second-order conditions. A key aspect of our approach is its versatile functional formulation in terms of the conditional tail process. Simulations demonstrate its performance across dependence scenarios. We apply our method to extreme event modelling in the oil industry, revealing distinct tail behaviours under varying conditioning values.

math.ST

Modeling discrete common-shock risks through matrix distributions

We introduce a novel class of bivariate common-shock discrete phase-type (CDPH) distributions to describe dependencies in loss modeling, with an emphasis on those induced by common shocks. By constructing two jointly evolving terminating Markov chains that share a common evolution up to a random time corresponding to the common shock component, and then proceed independently, we capture the essential features of risk events influenced by shared and individual-specific factors. We derive explicit expressions for the joint distribution of the termination times and prove various class and distributional properties, facilitating tractable analysis of the risks. Extending this framework, we model random sums where aggregate claims are sums of continuous phase-type random variables with counts determined by these termination times, and show that their joint distribution belongs to the multivariate phase-type or matrix-exponential class. We develop estimation procedures for the CDPH distributions using the expectation-maximization algorithm and demonstrate the applicability of our models through simulation studies and an application to bivariate insurance claim frequency data.

math.ST

Censored and extreme losses: functional convergence and applications to tail goodness-of-fit

This paper establishes the functional convergence of the Extreme Nelson--Aalen and Extreme Kaplan--Meier estimators, which are designed to capture the heavy-tailed behaviour of censored losses. The resulting limit representations can be used to obtain the distributions of pathwise functionals with respect to the so-called tail process. For instance, we may recover the convergence of a censored Hill estimator, and we further investigate two goodness-of-fit statistics for the tail of the loss distribution. Using the the latter limit theorems, we propose two rules for selecting a suitable number of order statistics, both based on test statistics derived from the functional convergence results. The effectiveness of these selection rules is investigated through simulations and an application to a real dataset comprised of French motor insurance claim sizes.

stat.ME

Uniform Consistency of Generalized Fr\'echet Means

Loss-based notions of centre on nonlinear spaces range from the Fr\'echet mean and power means to the geometric median and, in a limiting sense, the Chebyshev centre. To use such summaries statistically, one first needs a law of large numbers that remains valid beyond smooth manifolds and beyond a fixed choice of loss. We study generalized Fr\'echet means on metric spaces with the Heine--Borel property, obtained by replacing squared distance with a convex loss under a mild exponential-growth condition. We prove existence and compactness of the population mean set, establish a sharp diameter bound, obtain almost-sure consistency of empirical $\phi$-means, and derive a uniform strong law over compact classes of losses. The analysis is driven by a deterministic argmin principle together with a Glivenko--Cantelli theorem for monotone classes. For isotropic densities on Riemannian symmetric spaces, we identify the population $\phi$-mean for every strictly increasing loss for which the objective is finite, including bounded robust losses. We also illustrate the framework on spheres and on the polyhedral space of ultrametric phylogenetic trees.

math.ST

Heterogeneous extremes in the presence of random covariates and censoring

The task of analyzing extreme events with censoring effects is considered under a framework allowing for random covariate information. A wide class of estimators that can be cast as product-limit integrals is considered, for when the conditional distributions belong to the Frechet max-domain of attraction. The main mathematical contribution is establishing uniform conditions on the families of the regularly varying tails for which the asymptotic behaviour of the resulting estimators is tractable. In particular, a decomposition of the integral estimators in terms of exchangeable sums is provided, which leads to a law of large numbers and several central limit theorems. Subsequently, the finite-sample behaviour of the estimators is explored through a simulation study, and through the analysis of two real-life datasets. In particular, the inclusion of covariates makes the model significantly versatile and, as a consequence, practically relevant.

math.ST

Individual claims reserving using the Aalen--Johansen estimator

We propose an individual claims reserving model based on the conditional Aalen-Johansen estimator, as developed in Bladt and Furrer (2023b). In our approach, we formulate a multi-state problem, where the underlying variable is the individual claim size, rather than time. The states in this model represent development periods, and we estimate the cumulative density function of individual claim sizes using the conditional Aalen-Johansen method as transition probabilities to an absorbing state. Our methodology reinterprets the concept of multi-state models and offers a strategy for modeling the complete curve of individual claim sizes. To illustrate our approach, we apply our model to both simulated and real datasets. Having access to the entire dataset enables us to support the use of our approach by comparing the predicted total final cost with the actual amount, as well as evaluating it in terms of the continuously ranked probability score.

stat.AP

Extremile scalar-on-function regression

Extremiles provide a generalization of quantiles which are not only robust, but also have an intrinsic link with extreme value theory. This paper introduces an extremile regression model tailored for functional covariate spaces. The estimation procedure turns out to be a weighted version of local linear scalar-on-function regression, where now a double kernel approach plays a crucial role. Asymptotic expressions for the bias and variance are established, applicable to both decreasing bandwidth sequences and automatically selected bandwidths. The methodology is then investigated in detail through a simulation study. Furthermore, we illustrate the method's applicability with an analysis of the Berkeley Growth data, showcasing its performance in a real-world functional data setting.

stat.ME