SearcharxivSearch

arXiv subjects

Yuta Umezu

Publications and source records attributed to Yuta Umezu.

7 recordsLinked to original sources

Sure independence screening for covariate-dependent extreme value index estimation

One of the main topics in extreme value analysis is the estimation of the extreme value index, which characterizes the tail behavior of a distribution. Although covariate dependent extreme value index estimation has been widely studied, covariate screening for high-dimensional covariates has not been fully investigated. This paper proposes a sure independence screening method for covariate-dependent extreme value index estimation. The proposed method ranks covariates by marginal utilities constructed from a kernel-based conditional Pickands estimator. Unlike ordinary local smoothing, the proposed screening procedure uses a large-bandwidth kernel regime to obtain stable marginal contrasts. We establish the sure screening property under this regime, showing that all truly active covariates are retained with probability tending to one. Simulation studies and a real-data application demonstrate the effectiveness of the proposed method.

stat.ME

Selective Inference in Propensity Score Analysis

Selective inference (post-selection inference) is a methodology that has attracted much attention in recent years in the fields of statistics and machine learning. Naive inference based on data that are also used for model selection tends to show an overestimation, and so the selective inference conditions the event that the model was selected. In this paper, we develop selective inference in propensity score analysis with a semiparametric approach, which has become a standard tool in causal inference. Specifically, for the most basic causal inference model in which the causal effect can be written as a linear sum of confounding variables, we conduct Lasso-type variable selection by adding an $\ell_1$ penalty term to the loss function that gives a semiparametric estimator. Confidence intervals are then given for the coefficients of the selected confounding variables, conditional on the event of variable selection, with asymptotic guarantees. An important property of this method is that it does not require modeling of nonparametric regression functions for the outcome variables, as is usually the case with semiparametric propensity score analysis.

stat.ME

Selective Inference via Marginal Screening for High Dimensional Classification

Post-selection inference is a statistical technique for determining salient variables after model or variable selection. Recently, selective inference, a kind of post-selection inference framework, has garnered the attention in the statistics and machine learning communities. By conditioning on a specific variable selection procedure, selective inference can properly control for so-called selective type I error, which is a type I error conditional on a variable selection procedure, without imposing excessive additional computational costs. While selective inference can provide a valid hypothesis testing procedure, the main focus has hitherto been on Gaussian linear regression models. In this paper, we develop a selective inference framework for binary classification problem. We consider a logistic regression model after variable selection based on marginal screening, and derive the high dimensional statistical behavior of the post-selection estimator. This enables us to asymptotically control for selective type I error for the purposes of hypothesis testing after variable selection. We conduct several simulation studies to confirm the statistical power of the test, and compare our proposed method with data splitting and other methods.

stat.ME

Selective Inference for Change Point Detection in Multi-dimensional Sequences

We study the problem of detecting change points (CPs) that are characterized by a subset of dimensions in a multi-dimensional sequence. A method for detecting those CPs can be formulated as a two-stage method: one for selecting relevant dimensions, and another for selecting CPs. It has been difficult to properly control the false detection probability of these CP detection methods because selection bias in each stage must be properly corrected. Our main contribution in this paper is to formulate a CP detection problem as a selective inference problem, and show that exact (non-asymptotic) inference is possible for a class of CP detection methods. We demonstrate the performances of the proposed selective inference framework through numerical simulations and its application to our motivating medical data analysis problem.

stat.ML

Post Selection Inference with Kernels

We propose a novel kernel based post selection inference (PSI) algorithm, which can not only handle non-linearity in data but also structured output such as multi-dimensional and multi-label outputs. Specifically, we develop a PSI algorithm for independence measures, and propose the Hilbert-Schmidt Independence Criterion (HSIC) based PSI algorithm (hsicInf). The novelty of the proposed algorithm is that it can handle non-linearity and/or structured data through kernels. Namely, the proposed algorithm can be used for wider range of applications including nonlinear multi-class classification and multi-variate regressions, while existing PSI algorithms cannot handle them. Through synthetic experiments, we show that the proposed approach can find a set of statistically significant features for both regression and classification problems. Moreover, we apply the hsicInf algorithm to a real-world data, and show that hsicInf can successfully identify important features.

stat.ML

On the Consistency of the Bias Correction Term of the AIC for the Non-Concave Penalized Likelihood Method

Penalized likelihood methods with an $\ell_γ$-type penalty, such as the Bridge, the SCAD, and the MCP, allow us to estimate a parameter and to do variable selection, simultaneously, if $γ\in (0,1]$. In this method, it is important to choose a tuning parameter which controls the penalty level, since we can select the model as we want when we choose it arbitrarily. Nowadays, several information criteria have been developed to choose the tuning parameter without such an arbitrariness. However the bias correction term of such information criteria depend on the true parameter value in general, then we usually plug-in a consistent estimator of it to compute the information criteria from the data. In this paper, we derive a consistent estimator of the bias correction term of the AIC for the non-concave penalized likelihood method and propose a simple AIC-type information criterion for such models.

stat.ME

AIC for Non-concave Penalized Likelihood Method

Non-concave penalized maximum likelihood methods, such as the Bridge, the SCAD, and the MCP, are widely used because they not only do parameter estimation and variable selection simultaneously but also have a high efficiency as compared to the Lasso. They include a tuning parameter which controls a penalty level, and several information criteria have been developed for selecting it. While these criteria assure the model selection consistency and so have a high value, it is a severe problem that there are no appropriate rules to choose the one from a class of information criteria satisfying such a preferred asymptotic property. In this paper, we derive an information criterion based on the original definition of the AIC by considering the minimization of the prediction error rather than the model selection consistency. Concretely speaking, we derive a function of the score statistic which is asymptotically equivalent to the non-concave penalized maximum likelihood estimator, and then we provide an asymptotically unbiased estimator of the Kullback-Leibler divergence between the true distribution and the estimated distribution based on the function. Furthermore, through simulation studies, we check that the performance of the proposed information criterion gives almost the same as or better than that of the cross-validation.

stat.ME