Searcharxiv⌕ Search

arXiv subjects

Xiongzhi Chen

Publications and source records attributed to Xiongzhi Chen.

At least 19 recordsLinked to original sources

Direct Prediction Set Minimization via Bilevel Conformal Classifier Training

Conformal prediction (CP) is a promising uncertainty quantification framework which works as a wrapper around a black-box classifier to construct prediction sets (i.e., subset of candidate classes) with provable guarantees. However, standard calibration methods for CP tend to produce large prediction sets which makes them less useful in practice. This paper considers the problem of integrating conformal principles into the training process of deep classifiers to directly minimize the size of prediction sets. We formulate conformal training as a bilevel optimization problem and propose the {\em Direct Prediction Set Minimization (DPSM)} algorithm to solve it. The key insight behind DPSM is to minimize a measure of the prediction set size (upper level) that is conditioned on the learned quantile of conformity scores (lower level). We analyze that DPSM has a learning bound of $O(1/\sqrt{n})$ (with $n$ training samples), while prior conformal training methods based on stochastic approximation for the quantile has a bound of $Ω(1/s)$ (with batch size $s$ and typically $s \ll \sqrt{n}$). Experiments on various benchmark datasets and deep models show that DPSM significantly outperforms the best prior conformal training baseline with $20.46\%\downarrow$ in the prediction set size and validates our theory.

cs.LG↗

Uniformly consistent proportion estimation for composite hypotheses via integral equations: "the case of location-shift families"

We consider estimating the proportion of random variables for two types of composite null hypotheses: (i) the means or medians of the random variables belonging to a non-empty, bounded interval; (ii) the means or medians of the random variables belonging to an unbounded interval that is not the whole real line. For each type of composite null hypotheses, uniformly consistent estimators of the proportion of false null hypotheses are constructed for random variables whose distributions are members of a Type I location-shift family. Further, uniformly consistent estimators of certain functions of a bounded null on the means or medians are provided for the random variables mentioned earlier; these functions are continuous and of bounded variation. The estimators are constructed via solutions to Lebesgue-Stieltjes integral equations and harmonic analysis, do not rely on a concept of p-value, and have various applications.

math.ST↗

Uniformly consistent proportion estimation for composite hypotheses via integral equations: "the case of Gamma random variables"

We consider estimating the proportion of random variables for two types of composite null hypotheses: (i) the means of the random variables belonging to a non-empty, bounded interval; (ii) the means of the random variables belonging to an unbounded interval that is not the whole real line. For each type of composite null hypotheses, uniformly consistent estimators of the proportion of false null hypotheses are constructed for random variables whose distributions are members of the Gamma family. Further, uniformly consistent estimators of certain functions of a bounded null on the means are provided for the random variables mentioned earlier. These functions are continuous and of bounded variation. The estimators are constructed via solutions to Lebesgue-Stieltjes integral equations and harmonic analysis, do not rely on a concept of p-value, and have various applications.ce via mixture models, and may be used to estimate the sparsity level in high-dimensional Gaussian linear models.

math.ST↗

Consistent estimation of the proportion of false nulls and FDR for adaptive multiple testing Normal means under weak dependence

We consider multiple testing means of many dependent Normal random variables that do not necessarily follow a joint Normal distribution. Under weak dependence, we show the uniform consistency of proportion estimators that are constructed as solutions to Lebesgue-Stieltjes equations for the setting of a point, bounded and one-sided null, respectively, and characterize via the index of weak dependence the sparsest proportion these estimators can consistently estimate. On the other hand, under a principal correlation structure and employing a suitable definition of p-value for composite null hypotheses, we show that three key empirical processes induced by a single-step multiple testing procedure (MTP) satisfy the strong law of large numbers for testing each of the three types of nulls. Further, under this structure and for testing a point null and a one-sided null respectively, we construct an adaptive single-step MTP that employs a proportion estimator mentioned earlier, and show that the false discovery proportion of this procedure satisfies the weak law of large numbers and hence consistently estimates the false discovery rate of the procedure. In addition, we report some findings on the estimators of Jin and of Meinshausen and Rice of the proportion of false nulls in the critically and very sparse regimes under weak dependence and model misspecifications, respectively.

stat.ME↗

Variable Selection via Adaptive False Negative Control in Linear Regression

Variable selection methods have been developed in linear regression to provide sparse solutions. Recent studies have focused on further interpretations on the sparse solutions in terms of false positive control. In this paper, we consider false negative control for variable selection with the goal to efficiently select a high proportion of relevant predictors. Different from existing studies in power analysis and sure screening, we propose to directly estimate the false negative proportion (FNP) of a decision rule and select the smallest subset of predictors that has the estimated FNP less than a user-specified control level. The proposed method is adaptive to the user-specified control level on FNP by selecting less candidates if a higher level is implemented. On the other hand, when data has stronger effect size or larger sample size, the proposed method controls FNP more efficiently with less false positives. New analytic techniques are developed to cope with the major challenge of FNP control when relevant predictors cannot be consistently separated from irrelevant ones. Our numerical results are in line with the theoretical findings.

math.ST↗

A strong law of large numbers related to multiple testing Normal means

Assessing the stability of a multiple testing procedure under dependence is important but very challenging. Even for multiple testing which among a set of Normal random variables have mean zero, which we refer to as the "Normal means problem", to date there lacks a classification of the type of dependence under which the strong law of large numbers (SLLN) holds for the numbers of rejections and false rejections. We introduce the concept of "principal correlation structure (PCS)" that characterizes the type of dependence for which such SLLN holds, and establish the law. Further, we show that PCS ensures the SLLN for the false discover proportion when there is always a positive proportion of zero Normal means. We also investigate the stability of two conditional multiple testing procedures for the Normal means problem, and show that the associated SLLN holds when in addition the decomposition of the covariance matrix of the Normal random variables that induces PCS is homogeneous in certain sense. Our results also provide a formal way to check if the "weak dependence" assumption, a widely used assumption in the multiple testing literature, holds for the Normal means problem. As by-products, we establish a universal bound on Hermite polynomials and a universal comparison result on the covariance of the indicator functions of the two p-values of testing the marginal means of a bivariate Normal random vector and the correlation between the two components of the vector. These are of their own interests.

math.ST↗

A grouped, selectively weighted false discovery rate procedure

False discovery rate (FDR) control in structured hypotheses testing is an important topic in simultaneous inference. Most existing methods that aim to utilize group structure among hypotheses either employ the groupwise mixture model or weight all p-values or hypotheses. Thus, their powers can be improved when the groupwise mixture model is inappropriate or when most groups contain only true null hypotheses. Motivated by this, we propose a grouped, selectively weighted FDR procedure, which we refer to as "sGBH". Specifically, without employing the groupwise mixture model, sGBH identifies groups of hypotheses of interest, weights p-values in each such group only, and tests only the selected hypotheses using the weighted p-values. The sGBH subsumes a standard grouped, weighted FDR procedure which we refer to as "GBH". We provide simple conditions to ensure the conservativeness of sGBH, together with empirical evidence on its much improved power over GBH. The new procedure is applied to a gene expression study.

stat.ME↗

A weighted FDR procedure under discrete and heterogeneous null distributions

Multiple testing with false discovery rate (FDR) control has been widely conducted in the ``discrete paradigm" where p-values have discrete and heterogeneous null distributions. However, in this scenario existing FDR procedures often lose some power and may yield unreliable inference, and for this scenario there does not seem to be an FDR procedure that partitions hypotheses into groups, employs data-adaptive weights and is non-asymptotically conservative. We propose a weighted FDR procedure for multiple testing in the discrete paradigm that efficiently adapts to both the heterogeneity and discreteness of p-value distributions. We theoretically justify the non-asymptotic conservativeness of the weighted FDR procedure under independence, and show via simulation studies that, for multiple testing based on p-values of Binomial test or Fisher's exact test, it is more powerful than six other procedures. The weighted FDR procedure is applied to a drug safety study and a differential methylation study based on discrete data, where it makes more discoveries than two existing methods.

stat.ME↗

On Benjamini-Hochberg procedure applied to mid p-values

Multiple testing with discrete p-values routinely arises in various scientific endeavors. However, procedures, including the false discovery rate (FDR) controlling Benjamini-Hochberg (BH) procedure, often used in such settings, being developed originally for p-values with continuous distributions, are too conservative, and so may not be as powerful as one would hope for. Therefore, improving the BH procedure by suitably adapting it to discrete p-values without losing its FDR control is currently an important path of research. This paper studies the FDR control of the BH procedure when it is applied to mid p-values and derive conditions under which it is conservative. Our simulation study reveals that the BH procedure applied to mid p-values may be conservative under much more general settings than characterized in this work, and that an adaptive version of the BH procedure applied to mid p-values is as powerful as an existing adaptive procedure based on randomized p-values.

stat.ME↗

Uniformly consistently estimating the proportion of false null hypotheses via Lebesgue-Stieltjes integral equations

The proportion of false null hypotheses is a very important quantity in statistical modelling and inference based on the two-component mixture model and its extensions, and in control and estimation of the false discovery rate and false non-discovery rate. Most existing estimators of this proportion threshold p-values, deconvolve the mixture model under constraints on its components, or depend heavily on the location-shift property of distributions. Hence, they usually are not consistent, applicable to non-location-shift distributions, or applicable to discrete statistics or p-values. To eliminate these shortcomings, we construct uniformly consistent estimators of the proportion as solutions to Lebesgue-Stieltjes integral equations. In particular, we provide such estimators respectively for random variables whose distributions have Riemann-Lebesgue type characteristic functions, form discrete natural exponential families with infinite supports, and form natural exponential families with separable moment sequences. We provide the speed of convergence and uniform consistency class for each such estimator under independence. In addition, we provide example distribution families for which a consistent estimator of the proportion cannot be constructed using our techniques.

math.PR↗

Efficient Predictor Ranking and False Discovery Proportion Control in High-Dimensional Regression

We propose a ranking and selection procedure to prioritize relevant predictors and control false discovery proportion (FDP) of variable selection. Our procedure utilizes a new ranking method built upon the de-sparsified Lasso estimator. We show that the new ranking method achieves the optimal order of minimum non-zero effects in ranking relevant predictors ahead of irrelevant ones. Adopting the new ranking method, we develop a variable selection procedure to asymptotically control FDP at a user-specified level. We show that our procedure can consistently estimate the FDP of variable selection as long as the de-sparsified Lasso estimator is asymptotically normal. In numerical analyses, our procedure compares favorably to existing methods in ranking efficiency and FDP control when the regression model is relatively sparse.

stat.ME↗

False discovery rate control for multiple testing based on p-values with càdlàg distribution functions

For multiple testing based on p-values with càdlàg distribution functions, we propose an FDR procedure "BH+" with proven conservativeness. BH+ is at least as powerful as the BH procedure when they are applied to super-uniform p-values. Further, when applied to mid p-values, BH+ is more powerful than it is applied to conventional p-values. An easily verifiable necessary and sufficient condition for this is provided. BH+ is perhaps the first conservative FDR procedure applicable to mid p-values. BH+ is applied to multiple testing based on discrete p-values in a methylation study, an HIV study and a clinical safety study, where it makes considerably more discoveries than the BH procedure.

stat.ME↗

Multiple testing with discrete data: proportion of true null hypotheses and two adaptive FDR procedures

We consider multiple testing with false discovery rate (FDR) control when p-values have discrete and heterogeneous null distributions. We propose a new estimator of the proportion of true null hypotheses and demonstrate that it is less upwardly biased than Storey's estimator and two other estimators. The new estimator induces two adaptive procedures, i.e., an adaptive Benjamini-Hochberg (BH) procedure and an adaptive Benjamini-Hochberg-Heyse (BHH) procedure. We prove that the the adaptive BH procedure is conservative non-asymptotically. Through simulation studies, we show that these procedures are usually more powerful than their non-adaptive counterparts and that the adaptive BHH procedure is usually more powerful than the adaptive BH procedure and a procedure based on randomized p-value. The adaptive procedures are applied to a study of HIV vaccine efficacy, where they identify more differentially polymorphic positions than the BH procedure at the same FDR level.

stat.ME↗

Stopping time property of thresholds of Storey-type FDR procedures

For multiple testing, we introduce Storey-type FDR procedures and the concept of "regular estimator of the proportion of true nulls". We show that the rejection threshold of a Storey-type FDR procedure is a stopping time with respect to the backward filtration generated by the p-values and that a Storey-type FDR estimator at this rejection threshold equals the pre-specified FDR level, when the estimator of the proportion of true nulls is regular. These results hold regardless of the dependence among or the types of distributions of the p-values. They directly imply that a Storey-type FDR procedure is conservative when the null p-values are independent and uniformly distributed.

math.ST↗

Natural Exponential Families: Resolution of A Conjecture and Existence of Reduction Functions

One-parameter natural exponential family (NEF) plays fundamental roles in probability and statistics. This article contains two independent results: (a) A conjecture of Bar-Lev, Bshouty and Enis states that a polynomial with a simple root at $0$ and a complex root with positive imaginary part is the variance function of some NEF with mean domain $\left(0,\infty\right)$ if and only if the real part of the complex root is not positive. This conjecture is resolved. The positive answer to this conjecture enlarges existing family of polynomials that are able to generate NEFs, and it helps prevent practitioners from choosing incompatible functions as variance functions for statistical modeling using NEFs. (b) if a random variable $ξ$ has parametric distributions that form a infinitely divisible NEF whose induced measure is absolutely continuous with respect to its basis measure, then there exists a deterministic function $h$, called "reduction function", such that $\mathbb{E} \left(h\left(ξ\right)\right)=\mathbb{V}\left(ξ\right)$, i.e., $h\left(ξ\right)$ is an unbiased estimator of the variance of $ξ$. The reduction function has applications to estimating latent, low-dimensional structures and to dimension reduction in the first and/or second moments in high-dimensional data.

math.ST↗

Explicit solutions to a vector time series model and its induced model for business cycles

This article gives the explicit solution to a general vector time series model that describes interacting, heterogeneous agents that operate under uncertainties but according to Keynesian principles, from which a model for business cycle is induced by a weighted average of the growth rates of the agents in the model. The explicit solution enables a direct simulation of the time series defined by the model and better understanding of the joint behavior of the growth rates. In addition, the induced model for business cycles and its solutions are explicitly given and analyzed. The explicit solutions provide a better understanding of the mathematics of these models and the econometric properties they try to incorporate.

stat.ME↗