SearcharxivSearch

arXiv subjects

Andrew McInerney

Publications and source records attributed to Andrew McInerney.

2 recordsLinked to original sources

Investigating Statistical Inference and Covariate Effects in Shallow Neural Networks

Feedforward neural networks (FNNs) are typically viewed as pure prediction algorithms, and their strong predictive performance has led to their use in many machine-learning applications. However, their flexibility comes with an interpretability trade-off; thus, FNNs have been historically less popular among statisticians. Nevertheless, for suitably parsimonious shallow FNNs, classical statistical theory, such as significance testing and uncertainty quantification, may still provide useful regression-style summaries. Supplementing FNNs with methods of statistical inference, and covariate-effect visualisations, can shift the focus away from black-box prediction and move FNNs towards traditional statistical models. This can allow for more inferential analysis, and, hence, make FNNs more accessible within the statistical-modelling context. We investigate covariate-level Wald testing in the context of penalised FNNs, and also propose covariate-effect plots that emulate regression coefficients. Simulation studies are used to extensively investigate the performance of Wald-based inference, with particular emphasis on when this approach performs well. This statistical-based approach to neural networks is demonstrated through an application to insurance data.

stat.ME

A Statistical-Modelling Approach to Feedforward Neural Network Model Selection

Feedforward neural networks (FNNs) can be viewed as non-linear regression models, where covariates enter the model through a combination of weighted summations and non-linear functions. Although these models have some similarities to the approaches used within statistical modelling, the majority of neural network research has been conducted outside of the field of statistics. This has resulted in a lack of statistically-based methodology, and, in particular, there has been little emphasis on model parsimony. Determining the input layer structure is analogous to variable selection, while the structure for the hidden layer relates to model complexity. In practice, neural network model selection is often carried out by comparing models using out-of-sample performance. However, in contrast, the construction of an associated likelihood function opens the door to information-criteria-based variable and architecture selection. A novel model selection method, which performs both input- and hidden-node selection, is proposed using the Bayesian information criterion (BIC) for FNNs. The choice of BIC over out-of-sample performance as the model selection objective function leads to an increased probability of recovering the true model, while parsimoniously achieving favourable out-of-sample performance. Simulation studies are used to evaluate and justify the proposed method, and applications on real data are investigated.

stat.ME