SearcharxivSearch

arXiv subjects

Heather Battey

Publications and source records attributed to Heather Battey.

13 recordsLinked to original sources

On inferential equivalence classes of causal models

Causal models in the same Markov equivalence class are, in the absence of strong assumptions, statistically indistinguishable at any sample size. For modest sample sizes there is also the possibility that several such classes are compatible with the data, pointing to a confidence set of causal models as the appropriate presentation of evidence. Non-identifiability of Gaussian causal models from the same Markov equivalence class corresponds to a plurality of inverse-covariance representations in the unconstrained parametrisation of Cox and Wermuth (1993). This parametrisation, avoiding conic constraints that would otherwise complicate distributional approximations, facilitates construction of, and theoretical analysis for, a confidence set of causal models based on standard likelihood theory. By drawing on a geometric formulation of Evans (2020), we provide insight into which causal models are most likely to be included in the confidence set when the true parameter values of the generating model are at the borderline of detectability, delineating those models whose inclusion probabilities are stable at the nominal level, slowly decaying with sample size, and quickly decaying with sample size. We also study settings in which two or more causal models are in operation, exploring how the mixture weights, and the geometry of the models in the mixture relative to candidate models, interact with the inclusion probabilities. The purpose of the paper is to probe, from the perspective of structural properties of the true causal mechanism, the limits of what is achievable in causal inferential settings.

math.ST

Induced replication and the assessment of models

We study the assessment of semiparametric and other highly-parametrised models from the perspective of foundational principles of parametric statistical inference. In doing so, we highlight the possibility of avoiding the usual semiparametric considerations, which typically require estimation of nuisance components through kernel smoothing or basis expansion, with the associated difficulties of tuning-parameter choice that blur the distinction between estimation and model assessment. A key aspect is the inducement of replication under the postulated model. This can be cast in terms of some non-standard inferential separations, in the vein of Fisherian ancillarity/co-ancillarity and sufficiency/co-sufficiency separations, allowing the replacement of out-of-sample prediction error as a criterion for semiparametric model assessment by a type of within-sample prediction error. Framed in this light are new methodological contributions in multiple example settings, including model assessment for the proportional hazards model, for a time-dependent Poisson process with semiparametric intensity function, and for matched-pair and two-group examples. Also subsumed within the framework is a post-reduction inference approach to the construction of confidence sets of sparse regression models. Numerical work confirms recovery of nominal error rates under the postulated model and high sensitivity to departures in the direction of semiparametric alternatives. We conclude by emphasising open challenges and unifying perspectives.

stat.ME

Treatment effect: a critique

Two broad positions within statistics define a treatment effect, on the one hand, as a parameter of a statistical model, and on the other, as an appropriate population-level difference in outcomes or counterfactual outcomes under the different treatment regimes. This short expository paper presents some simple but consequential insights on the two formulations, contrasting the answers under the most favourable fictitious idealisation for the counterfactual framework. These observations clarify the relationship between Fisherian model-based inference and modern counterfactual formulations, and emphasise concerns, raised by Cox and others, regarding the suitability of model-free definitions as targets of inference when scientific conclusions are intended to generalise beyond the observed sample. Parts of the paper are necessarily controversial; we follow Cox (1958a) in not putting these forward in any dogmatic spirit.

stat.OT

Post-reduction inference for confidence sets of models

Sparsity in a regression context makes the model itself an object of interest, pointing to a confidence set of models as the appropriate presentation of evidence. A difficulty in areas such as genomics, where the number of candidate variables is vast, arises from the need for preliminary reduction prior to the assessment of models. The present paper considers a resolution using inferential separations fundamental to the Fisherian approach to conditional inference, namely, the sufficiency/co-sufficiency separation, and the ancillary/co-ancillary separation. The advantage of these separations is that no direction for departure from any hypothesised model is needed, avoiding issues that would otherwise arise from using the same data for reduction and for model assessment. In idealised cases with no nuisance parameters, the separations extract all the information in the data solely for the purpose for which it is useful, without loss or redundancy. The extent to which estimation of nuisance parameters affects the idealised information extraction is illustrated in detail for the normal-theory linear regression model, extending immediately to a log-normal accelerated-life model for time-to-event outcomes. This idealised analysis provides insight into when sample-splitting is likely to perform as well as, or better than, the co-sufficient or ancillary tests, and when it may be unreliable. The considerations involved in extending the detailed implementation to canonical exponential-family and more general regression models are briefly discussed. As part of the analysis for the Gaussian model, we introduce a modified version of the refitted cross-validation estimator of Fan et al. (2012), whose distribution theory is tractable in the appropriate conditional sense.

math.ST

Non-standard boundary behaviour in two-component mixture models

Consider a binary mixture model of the form $F_\theta = (1-\theta)F_0 + \theta F_1$, where $F_0$ is standard Gaussian and $F_1$ is a completely specified heavy-tailed distribution with the same support. For a sample of $n$ independent and identically distributed values $X_i \sim F_\theta$, the maximum likelihood estimator $\hat\theta_n$ is asymptotically normal provided that $0 < \theta < 1$ is an interior point. This paper investigates the large-sample behaviour for boundary points, which is entirely different and strikingly asymmetric for $\theta=0$ and $\theta=1$. The reason for the asymmetry has to do with typical choices such that $F_0$ is an extreme boundary point and $F_1$ is usually not extreme. On the right boundary, well known results on boundary parameter problems are recovered, giving $\lim \mathbb{P}_1(\hat\theta_n < 1)=1/2$. On the left boundary, $\lim\mathbb{P}_0(\hat\theta_n > 0)=1-1/\alpha$, where $1\leq \alpha \leq 2$ indexes the domain of attraction of the density ratio $f_1(X)/f_0(X)$ when $X\sim F_0$. For $\alpha=1$, which is the most important case in practice, we show how the tail behaviour of $F_1$ governs the rate at which $\mathbb{P}_0(\hat\theta_n > 0)$ tends to zero. A new limit theorem for the joint distribution of the sample maximum and sample mean conditional on positivity establishes multiple inferential anomalies. Most notably, given $\hat\theta_n > 0$, the likelihood ratio statistic has a conditional null limit distribution $G\neq\chi^2_1$ determined by the joint limit theorem. We show through this route that no advantage is gained by extending the single distribution $F_1$ to the nonparametric composite mixture generated by the same tail-equivalence class.

math.ST

Regression graphs and sparsity-inducing reparametrizations

That parametrization and sparsity are inherently linked raises the possibility that relevant models, not obviously sparse in their natural formulation, exhibit a population-level sparsity after reparametrization. In covariance models, positive-definiteness enforces additional constraints on how sparsity can legitimately manifest. It is therefore natural to consider reparametrization maps in which sparsity respects positive definiteness. The main purpose of this paper is to provide insight into structures on the physically-natural scale that induce and are induced by sparsity after reparametrization. The richest of the four structures initially uncovered can be generated, under a causal ordering, by the joint-response graphs studied by Wermuth & Cox (2004), while the most restrictive is that induced by sparsity on the scale of the matrix logarithm, studied by Battey (2017). The Iwasawa decomposition of the general linear group, combined with the graphical-models interpretation, points to a class of reparametrizations for the chain-graph models (Andersson et al. 2001), with undirected and directed acyclic graphs as special cases. An important insight is the interpretation of approximate zeros, explaining the modelling implications of enforcing sparsity after reparameterization: in effect, the relation between two variables would be declared null if relatively direct regression effects were negligible and others manifested through long paths. The insights have a bearing on methodology, some aspects of which are presented. A detailed simulation uses the theoretical insights to further explore regimes under which reparametrization is beneficial.

math.ST

On the role of parametrization in models with a misspecified nuisance component

The paper is concerned with inference for a parameter of interest in models that share a common interpretation for that parameter but that may differ appreciably in other respects. We study the general structure of models under which the maximum likelihood estimator of the parameter of interest is consistent under arbitrary misspecification of the nuisance part of the model. A specialization of the general results to matched-comparison and two-groups problems gives a more explicit and easily checkable condition in terms of a new notion of symmetric parametrization, leading to a broadening and unification of existing results in those problems. The role of a generalized definition of parameter orthogonality is highlighted, as well as connections to Neyman orthogonality. The issues involved in obtaining inferential guarantees beyond consistency are briefly discussed.

math.ST

Communication-Constrained Distributed Quantile Regression with Optimal Statistical Guarantees

We address the problem of how to achieve optimal inference in distributed quantile regression without stringent scaling conditions. This is challenging due to the non-smooth nature of the quantile regression (QR) loss function, which invalidates the use of existing methodology. The difficulties are resolved through a double-smoothing approach that is applied to the local (at each data source) and global objective functions. Despite the reliance on a delicate combination of local and global smoothing parameters, the quantile regression model is fully parametric, thereby facilitating interpretation. In the low-dimensional regime, we establish a finite-sample theoretical framework for the sequentially defined distributed QR estimators. This reveals a trade-off between the communication cost and statistical error. We further discuss and compare several alternative confidence set constructions, based on inversion of Wald and score-type tests and resampling techniques, detailing an improvement that is effective for more extreme quantile coefficients. In high dimensions, a sparse framework is adopted, where the proposed doubly-smoothed objective function is complemented with an $\ell_1$-penalty. We show that the corresponding distributed penalized QR estimator achieves the global convergence rate after a near-constant number of communication rounds. A thorough simulation study further elucidates our findings.

stat.ME

An Unethical Optimization Principle

If an artificial intelligence aims to maximise risk-adjusted return, then under mild conditions it is disproportionately likely to pick an unethical strategy unless the objective function allows sufficiently for this risk. Even if the proportion ${\eta}$ of available unethical strategies is small, the probability ${p_U}$ of picking an unethical strategy can become large; indeed unless returns are fat-tailed ${p_U}$ tends to unity as the strategy space becomes large. We define an Unethical Odds Ratio Upsilon (${\Upsilon}$) that allows us to calculate ${p_U}$ from ${\eta}$, and we derive a simple formula for the limit of ${\Upsilon}$ as the strategy space becomes large. We give an algorithm for estimating ${\Upsilon}$ and ${p_U}$ in finite cases and discuss how to deal with infinite strategy spaces. We show how this principle can be used to help detect unethical strategies and to estimate ${\eta}$. Finally we sketch some policy implications of this work.

q-fin.RM

HCmodelSets: An R package for specifying sets of well-fitting models in regression with a large number of potential explanatory variables

In the context of regression with a large number of explanatory variables, Cox and Battey (2017) emphasize that if there are alternative reasonable explanations of the data that are statistically indistinguishable, one should aim to specify as many of these explanations as is feasible. The standard practice, by contrast, is to report a single model effective for prediction. The present paper illustrates the R implementation of the new ideas in the package `HCmodelSets', using simple reproducible examples and real data. Results of some simulation experiments are also reported.

stat.CO

Distributed Estimation and Inference with Statistical Guarantees

This paper studies hypothesis testing and parameter estimation in the context of the divide and conquer algorithm. In a unified likelihood based framework, we propose new test statistics and point estimators obtained by aggregating various statistics from $k$ subsamples of size $n/k$, where $n$ is the sample size. In both low dimensional and high dimensional settings, we address the important question of how to choose $k$ as $n$ grows large, providing a theoretical upper bound on $k$ such that the information loss due to the divide and conquer algorithm is negligible. In other words, the resulting estimators have the same inferential efficiencies and estimation rates as a practically infeasible oracle with access to the full sample. Thorough numerical results are provided to back up the theory.

math.ST

A topologically valid definition of depth for functional data

The main focus of this work is on providing a formal definition of statistical depth for functional data on the basis of six properties, recognising topological features such as continuity, smoothness and contiguity. Amongst our depth defining properties is one that addresses the delicate challenge of inherent partial observability of functional data, with fulfilment giving rise to a minimal guarantee on the performance of the empirical depth beyond the idealised and practically infeasible case of full observability. As an incidental product, functional depths satisfying our definition achieve a robustness that is commonly ascribed to depth, despite the absence of a formal guarantee in the multivariate definition of depth. We demonstrate the fulfilment or otherwise of our properties for six widely used functional depth proposals, thereby providing a systematic basis for selection of a depth function.

math.ST

Smooth projected density estimation

We introduce and analyse a new nonparametric estimator of a multi-dimensional density. Our smooth projection estimator (SPE) is defined by a least squares projection of the sample onto an infinite dimensional mixture class via an undersmoothed nonparametric pilot estimate, which acts as a structural filter to regularise the solution. The undersmoothing is required to optimise the convergence rate of the SPE, which is jointly determined by that of the pilot estimator to the true density in squared $\mathbb{L}_{2}$ norm, and by that of the pilot distribution function to the empirical distribution function in uniform norm. Our procedure was conceived with a view to exploiting well known results in convex analysis and their connection to mixture densities. In the context of our work, this translates to the observation that the infinite dimensional minimisation problem, implicit in the construction of the SPE, possesses a solution of dimension at most $n+1$, where $n$ is the sample size. The SPE thus enjoys practical advantages such as computational efficiency, ease of storage and rapid evaluation at a new data point.

stat.ME