SearcharxivSearch

arXiv subjects

Simon Boge Brant

Publications and source records attributed to Simon Boge Brant.

3 recordsLinked to original sources

Generalised logistic regression with vine copulas

We propose a generalisation of the logistic regression model, that aims to account for non-linear main effects and complex interactions, while keeping the model inherently explainable. This is obtained by starting with log-odds that are linear in the covariates, and adding non-linear terms that depend on at least two covariates. More specifically, we use a generative specification of the model, consisting of a combination of certain margins on natural exponential form, combined with vine copulas. The estimation of the model is however based on the discriminative likelihood, and dependencies between covariates are included in the model, only if they contribute significantly to the distinction between the two classes. Further, a scheme for model selection and estimation is presented. The methods described in this paper are implemented in the R package LogisticCopula. In order to assess the performance of our model, we ran an extensive simulation study. The results from the study, as well as from a couple of examples on real data, showed that our model performs at least as well as natural competitors, especially in the presence of non-linearities and complex interactions, even when $n$ is not large compared to $p$.

stat.ME

Boosting with copula-based components

The authors propose new additive models for binary outcomes, where the components are copula-based regression models (Noh et al, 2013), and designed such that the model may capture potentially complex interaction effects. The models do not require discretisation of continuous covariates, and are therefore suitable for problems with many such covariates. A fitting algorithm, and efficient procedures for model selection and evaluation of the components are described. Software is provided in the R-package copulaboost. Simulations and illustrations on data sets indicate that the method's predictive performance is either better than or comparable to the other methods.

stat.ME

The fraud loss for selecting the model complexity in fraud detection

In fraud detection applications, the investigator is typically limited to controlling a restricted number k of cases. The most efficient manner of allocating the resources is then to try selecting the k cases with the highest probability of being fraudulent. The prediction model used for this purpose must normally be regularized to avoid overfitting and consequently bad prediction performance. A new loss function, denoted the fraud loss, is proposed for selecting the model complexity via a tuning parameter. A simulation study is performed to find the optimal settings for validation. Further, the performance of the proposed procedure is compared to the most relevant competing procedure, based on the area under the receiver operating characteristic curve (AUC), in a set of simulations, as well as on a VAT fraud dataset. In most cases, choosing the complexity of the model according to the fraud loss, gave a better than, or comparable performance to the AUC in terms of the fraud loss.

stat.AP