SearcharxivSearch

arXiv subjects

Jelena Markovic

Publications and source records attributed to Jelena Markovic.

10 recordsLinked to original sources

Inference after black box selection

We consider the problem of inference for parameters selected to report only after some algorithm, the canonical example being inference for model parameters after a model selection procedure. The conditional correction for selection requires knowledge of how the selection is affected by changes in the underlying data, and current research explicitly describes this selection. In this work, we assume 1) we have in silico access to the selection algorithm and 2) for parameters of interest, the data input into the algorithm satisfies (pre-selection) a central limit theorem jointly with an estimator of our parameter of interest. Under these assumptions, we recast the problem into a statistical learning problem which can be fit with off-the-shelf models for binary regression. The feature points in this problem are set by the user, opening up the possibility of active learning methods for computationally expensive selection algorithms. We consider two examples previously out of reach of this conditional approach: stability selection and multiple cross-validation.

stat.ME

Non-reversible, tuning- and rejection-free Markov chain Monte Carlo via iterated random functions

In this work we present a non-reversible, tuning- and rejection-free Markov chain Monte Carlo which naturally fits in the framework of hit-and-run. The sampler only requires access to the gradient of the log-density function, hence the normalizing constant is not needed. We prove the proposed Markov chain is invariant for the target distribution and illustrate its applicability through a wide range of examples. We show that the sampler introduced in the present paper is intimately related to the continuous sampler of Peters and de With (2012), Bouchard-Cote et al. (2017). In particular, the computation is quite similar in the sense that both are centered around simulating an inhomogenuous Poisson process. The computation can be simplified when the gradient of the log-density admits a computationally efficient directional decomposition into a sum of two monotone functions. We apply our sampler in selective inference, gaining significant improvement over the formerly used sampler (Tian et al. 2016).

stat.CO

Bouncy Hybrid Sampler as a Unifying Device

This work introduces a class of rejection-free Markov chain Monte Carlo (MCMC) samplers, named the Bouncy Hybrid Sampler, which unifies several existing methods from the literature. Examples include the Bouncy Particle Sampler of Peters and de With (2012), Bouchard-Cote et al. (2015) and the Hamiltonian MCMC. Following the introduced general framework, we derive a new sampler called the Quadratic Bouncy Hybrid Sampler. We apply this novel sampler to the problem of sampling from a truncated Gaussian distribution.

stat.CO

Unifying approach to selective inference with applications to cross-validation

We develop tools to do valid post-selective inference for a family of model selection procedures, including choosing a model via cross-validated Lasso. The tools apply universally when the following random vectors are jointly asymptotically multivariate Gaussian: 1. the vector composed of each model's quality value evaluated under certain model selection criteria (e.g. cross-validation errors across folds, AIC, prediction errors etc.) 2. the test statistics from which we make inference on the parameters; it is worth noting that the parameters here are chosen after model selection methods are performed. Under these assumptions, we derive a pivotal quantity that has an asymptotically Unif(0,1) distribution which can be used to perform tests and construct confidence intervals. Both the tests and confidence intervals are selectively valid for the chosen parameter. While the above assumptions may not be satisfied in some applications, we propose a novel variation to these model selection procedures by adding Gaussian randomizations to either one of the two vectors. As a result, the joint distribution of the above random vectors is multivariate Gaussian and our general tools apply. We illustrate our method by applying it to four important procedures for which very few selective inference results have been developed: cross-validated Lasso, cross-validated randomized Lasso, AIC-based model selection among a fixed set of models and inference for a newly introduced novel marginal LOCO parameter, inspired by the LOCO parameter of Rinaldo et al (2016); and we provide complete results for these cases. For randomized model selection procedures, we develop Markov chain Monte Carlo sampling scheme to construct valid post-selective confidence intervals empirically.

stat.ME

More powerful post-selection inference, with application to the Lasso

Investigators often use the data to generate interesting hypotheses and then perform inference for the generated hypotheses. P-values and confidence intervals must account for this explorative data analysis. A fruitful method for doing so is to condition any inferences on the components of the data used to generate the hypotheses, thus preventing information in those components from being used again. Some currently popular methods "over-condition", leading to wide intervals. We show how to perform the minimal conditioning in a computationally tractable way. In high dimensions, even this minimal conditioning can lead to intervals that are too wide to be useful, suggesting that up to now the cost of hypothesis generation has been underestimated. We show how to generate hypotheses in a strategic manner that sharply reduces the cost of data exploration and results in useful confidence intervals. Our discussion focuses on the problem of post-selection inference after fitting a lasso regression model, but we also outline its extension to a much more general setting.

stat.ME

Bootstrap inference after using multiple queries for model selection

In this work, we provide a refinement of the selective CLT result of Tian and Taylor (2015), which allows for selective inference in non-parametric settings by adjusting for the asymptotic Gaussian limit for selection. Under some regularity assumptions on the density of the randomization, including heavier tails than Gaussian satisfied by e.g. logistic distribution, we prove the selective CLT holds without any assumptions on the underlying parameter, allowing for rare selection events. We also show a selective CLT result for Gaussian randomization, though the quantitative results are qualitatively different for the Gaussian randomization as compared to the heavier tailed results. Furthermore, we propose a bootstrap version of this test statistic, which is provably asymptotically pivotal uniformly across a family of non-parametric distributions. This result can be interpreted as resolving the impossibility results of Leeb and Potscher (2006). We describe several sampling methods involving the projected Langevin Monte Carlo to compute the bootstrapped test statistic and the corresponding confidence intervals valid after selection. The applications of our work include valid inferential and sampling tools after running various model selection algorithms including their combinations into multiple views/queries framework. We also present a way to do data carving, providing more powerful tests than classical data splitting by reusing the information in the data from the first stage.

stat.ME

Inferactive data analysis

We describe inferactive data analysis, so-named to denote an interactive approach to data analysis with an emphasis on inference after data analysis. Our approach is a compromise between Tukey's exploratory (roughly speaking "model free") and confirmatory data analysis (roughly speaking classical and "model based"), also allowing for Bayesian data analysis. We view this approach as close in spirit to current practice of applied statisticians and data scientists while allowing frequentist guarantees for results to be reported in the scientific literature, or Bayesian results where the data scientist may choose the statistical model (and hence the prior) after some initial exploratory analysis. While this approach to data analysis does not cover every scenario, and every possible algorithm data scientists may use, we see this as a useful step in concrete providing tools (with frequentist statistical guarantees) for current data scientists. The basis of inference we use is selective inference [Lee et al., 2016, Fithian et al., 2014], in particular its randomized form [Tian and Taylor, 2015a]. The randomized framework, besides providing additional power and shorter confidence intervals, also provides explicit forms for relevant reference distributions (up to normalization) through the {\em selective sampler} of Tian et al. [2016]. The reference distributions are constructed from a particular conditional distribution formed from what we call a DAG-DAG -- a Data Analysis Generative DAG. As sampling conditional distributions in DAGs is generally complex, the selective sampler is crucial to any practical implementation of inferactive data analysis. Our principal goal is in reviewing the recent developments in selective inference as well as describing the general philosophy of selective inference.

math.ST

An MCMC-free approach to post-selective inference

We develop a Monte Carlo-free approach to inference post output from randomized algorithms with a convex loss and a convex penalty. The pivotal statistic based on a truncated law, called the selective pivot, usually lacks closed form expressions. Inference in these settings relies upon standard Monte Carlo sampling techniques at a reference parameter followed by an exponential tilting at the reference. Tilting can however be unstable for parameters that are far off from the reference parameter. We offer in this paper an alternative approach to construction of intervals and point estimates by proposing an approximation to the intractable selective pivot. Such an approximation solves a convex optimization problem in |E| dimensions, where |E| is the size of the active set observed from selection. We empirically show that the confidence intervals obtained by inverting the approximate pivot have valid coverage.

stat.ME

Selective sampling after solving a convex problem

We consider the problem of selective inference after solving a (randomized) convex statistical learning program in the form of a penalized or constrained loss function. Our first main result is a change-of-measure formula that describes many conditional sampling problems of interest in selective inference. Our approach is model-agnostic in the sense that users may provide their own statistical model for inference, we simply provide the modification of each distribution in the model after the selection. Our second main result describes the geometric structure in the Jacobian appearing in the change of measure, drawing connections to curvature measures appearing in Weyl-Steiner volume-of-tubes formulae. This Jacobian is necessary for problems in which the convex penalty is not polyhedral, with the prototypical example being group LASSO or the nuclear norm. We derive explicit formulae for the Jacobian of the group LASSO. To illustrate the generality of our method, we consider many examples throughout, varying both the penalty or constraint in the statistical learning problem as well as the loss function, also considering selective inference after solving multiple statistical learning programs. Penalties considered include LASSO, forward stepwise, stagewise algorithms, marginal screening and generalized LASSO. Loss functions considered include squared-error, logistic, and log-det for covariance matrix estimation. Having described the appropriate distribution we wish to sample from through our first two results, we outline a framework for sampling using a projected Langevin sampler in the (commonly occuring) case that the distribution is log-concave.

math.ST

Estimating the Ratio of Two Functions in a Nonparametric Regression Model

Due to measurement noise, a common problem in in various fields is how to estimate the ratio of two functions. We consider this problem of estimating the ratio of two functions in a nonparametric regression model. Assuming the noise is normally distributed, this is equivalent to estimating the ratio of the means of two normally distributed random variables. We identified a consistent estimator that gives the mean squared loss of order $O(1/n)$ ($n$ is the sample size) when conditioned on a highly probable event. We also present our result applied to both the real data from EAPS and on simulated data, confirming our theoretical results.

stat.ME