SearcharxivSearch

arXiv subjects

Daniel Xiang

Publications and source records attributed to Daniel Xiang.

11 recordsLinked to original sources

Hierarchical Bayesian Estimation of Covariance Matrices

We develop a hierarchical Bayesian framework for covariance matrix estimation built on a key observation: while equivariance under the full general linear group GL(p) is well known, it is an extremely restrictive property -- estimators equivariant to GL(p) are limited to scalar multiples of the sample covariance matrix and carry considerably larger risks than shrinkage estimators. By contrast, commonly used shrinkage estimators, including the Haff empirical Bayes estimator, and the Ledoit--Wolf estimators, are all equivariant under the smaller orthogonal group O(p). Exploiting this structure, we establish that the Haar measure Bayes rule in an oracle eigenvalue model is the minimum risk estimator within the class of O(p)-equivariant estimators, and derive oracle Bayes rules for the covariance and precision matrices under the squared Frobenius, Stein, and squared Stein loss functions. These oracle rules serve as theoretical benchmarks that dominate all commonly used estimators. To approximate them when the true eigenvalues are unknown, we introduce a hierarchical Bayes model that places a finite P'olya tree prior on the eigenvalue distribution and uses Gibbs sampling to generate posterior draws, yielding both shrinkage estimates for the eigenvalues and approximations to the oracle Bayes rules. Simulations suggest that the finite P'olya tree prior is able to recover the general form of the distribution of the eigenvalues, and confirm that the resulting estimators closely approach oracle performance, substantially outperforming classical competitors for both covariance and precision matrix estimation.

stat.ME

Estimating the local false discovery rate under an unknown symmetric null

This paper is concerned with estimating the local false discovery rate (lfdr) in a two-groups model where the only assumption regarding the null distribution is symmetry about zero. Our motivation comes from the contemporary framework for multiple hypothesis testing, particularly relevant in variable selection problems, which transforms any user-specified scores into statistics whose null distributions are symmetric about zero, whereas enrichment to the right of zero is generally expected for the non-nulls. While modern methods such as the knockoff filter (Barber and Candes; 2015) are able to exploit the null property for controlling the false discovery rate (FDR), an arguably more appropriate goal is to target control of the local false discovery rate for the rejected hypotheses, as proposed in Soloff et al. (2024) where the standard two-groups model (known $f_0$ and independence) is analyzed. Here, we take a step in this direction and propose to estimate the lfdr by targeting the surrogate density ratio $f(-w)/f(w)$, for $w>0$, where $f$ is the marginal density in the aforementioned ``stripped-down'' two-groups model. We study several estimators and propose a logistic regression based method with natural cubic spline basis. We also show that any consistent estimator of this surrogate yields asymptotic lfdr control of the multiple testing procedure that thresholds the estimate at the nominal level.

stat.ME

Adaptive procedures for boundary FDR control

A cornerstone of the multiple testing literature is the Benjamini-Hochberg (BH) procedure, which guarantees control of the FDR when $p$-values are independent or positively dependent. While BH controls the average quality of rejections, it does not provide guarantees for individual discoveries, particularly those near the rejection threshold, which are more likely to be false than the average rejection. For independent $p$-values with Uniform$(0,1)$ null distribution, the Support Line procedure (SL; arXiv:2207.07299) provably controls the error probability for the rejection at the edge of the discovery set (i.e. the one with largest $p$-value) at level $q m_0/m$, where $m_0$ is the number of true null hypotheses and $q$ is a tuning parameter. In this work, we study adaptive versions of the SL procedure that operate in two steps: the first step estimates $m_0$ from non-significant statistics, and the second step runs the SL procedure at an adjusted level $q m / \hat{m}_0$. The adaptive procedures are shown to control the false discovery probability for the "boundary'' rejection under an independence assumption. Simulation studies suggest that some but not all of the two-stage procedures maintain error control under positive dependence, and that substantial power is gained relative to the original SL procedure. We illustrate differences between the procedures on meta-data from the recent literature in behavioral psychology on growth mindset and nudge interventions.

stat.ME

Reliable conformal novelty detection at the decision boundary

Novelty detection via conformal $p$-values and BH procedure provides distribution-free global false discovery rate (FDR) control. We present here fundamental limits of this approach by showing that it does not produce reliable detection at the decision boundary. We study boundary false discovery rate (bFDR), the probability that the least extreme reported novelty is in fact a null observation. We first show that the support line (SL) procedure, controlling the bFDR in the continuous independent framework, fails to control the bFDR in the conformal case. We then present several modifications of the SL procedure that restore reliability at the decision boundary, by controlling the bFDR, with specific improvements in situations where many novelties are expected (adaptive procedures) and when the calibration sample is too small with respect to the test sample (subsampled procedures). Numerical experiments with both synthetic and real data support our findings and show the relevance of the new proposed approach.

stat.ME

A frequentist local false discovery rate

The local false discovery rate (lfdr) of Efron et al. (2001) enjoys major conceptual and decision-theoretic advantages over the false discovery rate (FDR) as an error criterion in multiple testing, but is only well-defined in Bayesian models where the truth status of each null hypothesis is random. We define a frequentist counterpart to the lfdr based on the relative frequency of nulls at each point in the sample space. The frequentist lfdr is defined without reference to any prior, but preserves several important properties of the Bayesian lfdr: For continuous test statistics, $\text{lfdr}(t)$ gives the probability, conditional on observing some statistic equal to $t$, that the corresponding null hypothesis is true. Evaluating the lfdr at an individual test statistic also yields a calibrated forecast of whether its null hypothesis is true. Finally, thresholding the lfdr at $\frac{1}{1+\lambda}$ gives the best separable rejection rule under the weighted classification loss where Type I errors are $\lambda$ times as costly as Type II errors. The lfdr can be estimated efficiently using parametric or non-parametric methods, and a closely related error criterion can be provably controlled in finite samples under independence assumptions. Whereas the FDR measures the average quality of all discoveries in a given rejection region, our lfdr measures how the quality of discoveries varies across the rejection region, allowing for a more fine-grained analysis without requiring the introduction of a prior.

math.ST

Non-standard boundary behaviour in two-component mixture models

Consider a binary mixture model of the form $F_\theta = (1-\theta)F_0 + \theta F_1$, where $F_0$ is standard Gaussian and $F_1$ is a completely specified heavy-tailed distribution with the same support. For a sample of $n$ independent and identically distributed values $X_i \sim F_\theta$, the maximum likelihood estimator $\hat\theta_n$ is asymptotically normal provided that $0 < \theta < 1$ is an interior point. This paper investigates the large-sample behaviour for boundary points, which is entirely different and strikingly asymmetric for $\theta=0$ and $\theta=1$. The reason for the asymmetry has to do with typical choices such that $F_0$ is an extreme boundary point and $F_1$ is usually not extreme. On the right boundary, well known results on boundary parameter problems are recovered, giving $\lim \mathbb{P}_1(\hat\theta_n < 1)=1/2$. On the left boundary, $\lim\mathbb{P}_0(\hat\theta_n > 0)=1-1/\alpha$, where $1\leq \alpha \leq 2$ indexes the domain of attraction of the density ratio $f_1(X)/f_0(X)$ when $X\sim F_0$. For $\alpha=1$, which is the most important case in practice, we show how the tail behaviour of $F_1$ governs the rate at which $\mathbb{P}_0(\hat\theta_n > 0)$ tends to zero. A new limit theorem for the joint distribution of the sample maximum and sample mean conditional on positivity establishes multiple inferential anomalies. Most notably, given $\hat\theta_n > 0$, the likelihood ratio statistic has a conditional null limit distribution $G\neq\chi^2_1$ determined by the joint limit theorem. We show through this route that no advantage is gained by extending the single distribution $F_1$ to the nonparametric composite mixture generated by the same tail-equivalence class.

math.ST

Sharp phase transitions in high-dimensional changepoint detection

We study a hypothesis testing problem in the context of high-dimensional changepoint detection. Given a matrix $X \in \R^{p \times n}$ with independent Gaussian entries, the goal is to determine whether or not a sparse, non-null fraction of rows in $X$ exhibits a shift in mean at a common index between $1$ and $n$. We focus on three aspects of this problem: the sparsity of non-null rows, the presence of a single, common changepoint in the non-null rows, and the signal strength associated with the changepoint. Within an asymptotic regime relating the data dimensions $n$ and $p$ to the signal sparsity and strength, the information-theoretic limits of this testing problem are characterized by a formula that determines whether or not there exists a testing procedure whose sum of Type I and II errors tends to zero as $n,p \to \infty$. The formula, called the \emph{detection boundary}, partitions the parameter space into a two regions: one where it is possible to detect the presence of a single aligned changepoint (detectable region), and another where no test is able to consistently distinguish the mean matrix from one with constant rows (undetectable region).

math.ST

Interpretation of local false discovery rates under the zero assumption

In large-scale studies with parallel signal-plus-noise observations, the local false discovery rate is a summary statistic that is often presumed to be equal to the posterior probability that the signal is null. We prefer to call the latter quantity the local null-signal rate to emphasize our view that a null signal and a false discovery are not identical events. The local null-signal rate is commonly estimated through empirical Bayes procedures that build on the `zero density assumption,' which attributes the density of observations near zero entirely to null signals. In this paper, we argue that this strategy does not furnish estimates of the local null-signal rate, but instead of a quantity we call the complementary local activity rate (clar). Although it is likely to be small, an inactive signal is not necessarily zero. The clar dominates both the local null-signal rate and the local false sign rate and is a weakly continuous functional of the signal distribution. As a consequence, it takes on sensible values when the signal is sparse but not exactly zero. Our findings clarify the interpretation of local false discovery rates estimated under the zero density assumption.

math.ST

Sparse-limit approximation for t-statistics

In a range of genomic applications, it is of interest to quantify the evidence that the signal at site~$i$ is active given conditionally independent replicate observations summarized by the sample mean and variance $(\bar Y, s^2)$ at each site. We study the version of the problem in which the signal distribution is sparse, and the error distribution has an unknown site-specific variance so that the null distribution of the standardized statistic is Student-$t$ rather than Gaussian. The main contribution of this paper is a sparse-mixture approximation to the non-null density of the $t$-ratio. This formula demonstrates the effect of low degrees of freedom on the Bayes factor, or the conditional probability that the site is active. We illustrate some differences on a HIV dataset for gene-expression data previously analyzed by Efron (2012).

math.ST

The edge of discovery: Controlling the local false discovery rate at the margin

Despite the popularity of the false discovery rate (FDR) as an error control metric for large-scale multiple testing, its close Bayesian counterpart the local false discovery rate (lfdr), defined as the posterior probability that a particular null hypothesis is false, is a more directly relevant standard for justifying and interpreting individual rejections. However, the lfdr is difficult to work with in small samples, as the prior distribution is typically unknown. We propose a simple multiple testing procedure and prove that it controls the expectation of the maximum lfdr across all rejections; equivalently, it controls the probability that the rejection with the largest p-value is a false discovery. Our method operates without knowledge of the prior, assuming only that the p-value density is uniform under the null and decreasing under the alternative. We also show that our method asymptotically implements the oracle Bayes procedure for a weighted classification risk, optimally trading off between false positives and false negatives. We derive the limiting distribution of the attained maximum lfdr over the rejections, and the limiting empirical Bayes regret relative to the oracle procedure.

stat.ME

Permanental Graphs

The two components for infinite exchangeability of a sequence of distributions $(P_n)$ are (i) consistency, and (ii) finite exchangeability for each $n$. A consequence of the Aldous-Hoover theorem is that any node-exchangeable, subselection-consistent sequence of distributions that describes a randomly evolving network yields a sequence of random graphs whose expected number of edges grows quadratically in the number of nodes. In this note, another notion of consistency is considered, namely, delete-and-repair consistency; it is motivated by the sense in which infinitely exchangeable permutations defined by the Chinese restaurant process (CRP) are consistent. A goal is to exploit delete-and-repair consistency to obtain a nontrivial sequence of distributions on graphs $(P_n)$ that is sparse, exchangeable, and consistent with respect to delete-and-repair, a well known example being the Ewens permutations \cite{tavare}. A generalization of the CRP$(\alpha)$ as a distribution on a directed graph using the $\alpha$-weighted permanent is presented along with the corresponding normalization constant and degree distribution; it is dubbed the Permanental Graph Model (PGM). A negative result is obtained: no setting of parameters in the PGM allows for a consistent sequence $(P_n)$ in the sense of either subselection or delete-and-repair.

math.ST