SearcharxivSearch

arXiv subjects

Ran Xie

Publications and source records attributed to Ran Xie.

5 recordsLinked to original sources

Method of Moments Estimation of High-Dimensional Covariance Using a Parametric Model

We propose method-of-moments estimators for the eigenvalues of variance component covariance matrices in multivariate mixed effects models. Assuming a parametric form for the eigenvalue distribution, we focus on the high-dimensional regime where the number of predictors is large and comparable to the number of realizations of each random effect. In this setting, we show that the empirical moments of sum-of-squares matrices (e.g., MANOVA estimators of the covariance matrices) can be closely approximated by deterministic functions of the underlying parameters. This relationship enables the construction of consistent and asymptotically normal estimators via moment matching. Our approach is motivated by applications in quantitative genetics, where estimating genetic covariance components across multiple phenotypic traits is of central interest. We implement our method in a new python package mlmm-mom, and demonstrate how our method adapts to several common experimental designs in this domain.

stat.ME

Boosted Conformal Prediction Intervals

This paper introduces a boosted conformal procedure designed to tailor conformalized prediction intervals toward specific desired properties, such as enhanced conditional coverage or reduced interval length. We employ machine learning techniques, notably gradient boosting, to systematically improve upon a predefined conformity score function. This process is guided by carefully constructed loss functions that measure the deviation of prediction intervals from the targeted properties. The procedure operates post-training, relying solely on model predictions and without modifying the trained model (e.g., the deep network). Systematic experiments demonstrate that starting from conventional conformal methods, our boosted procedure achieves substantial improvements in reducing interval length and decreasing deviation from target conditional coverage.

stat.ME

CLT for Linear Spectral Statistics in High-Dimensional Random Effects Models

We study sample covariance matrices arising from multi-level components of variance. Thus, let $ B_n=\frac{1}{N}\sum_{j=1}^NT_{j}^{1/2}x_jx_j^TT_{j}^{1/2}$, where $x_j\in R^n$ are i.i.d. standard Gaussian, and $T_{j}=\sum_{r=1}^kl_{jr}^2Σ_{r}$ are $n\times n$ real symmetric matrices with bounded spectral norm, corresponding to $k$ levels of variation. As the matrix dimensions $n$ and $N$ increase proportionally, we show that the linear spectral statistics (LSS) of $B_n$ have Gaussian limits. The CLT is expressed as the convergence of a set of LSS to a standard multivariate Gaussian after centering by a mean vector $Γ_n$ and a covariance matrix $Λ_n$ which depend on $n$ and $N$ and may be evaluated numerically. Our work is motivated by the estimation of high-dimensional covariance matrices between phenotypic traits in quantitative genetics, particularly within nested linear random-effects models with up to $k$ levels of randomness. Our proof builds on the Bai-Silverstein \cite{baisilverstein2004} martingale method with some innovation to handle the multi-level setting.

math.PR

They may look and look, yet not see: BMDs cannot be tested adequately

Bugs, misconfiguration, and malware can cause ballot-marking devices (BMDs) to print incorrect votes. Several approaches to testing BMDs have been proposed. In logic and accuracy testing (LAT) and parallel or live testing, auditors input known test votes into the BMD and check the printout. Passive testing monitors the rate of "spoiled" BMD printout, on the theory that if BMDs malfunction, the rate will increase noticeably. We show that these approaches cannot reliably detect outcome-altering problems, because: (i) The number of possible interactions with BMDs is enormous, so testing interactions uniformly at random is hopeless. (ii) To probe the space of interactions intelligently requires an accurate model of voter behavior, but because the space of interactions is so large, building an accurate model requires observing a huge number of voters in every jurisdiction in every election--more voters than there are in most jurisdictions. (iii) Even with a perfect model of voter behavior, the number of tests needed exceeds the number of voters in most jurisdictions. (iv) An attacker can target interactions that are expensive to test, e.g., because they involve voting slowly; or interactions for which tampering is less likely to be noticed, e.g., because the voter uses the audio interface. (v) Whether BMDs misbehave or not, the distribution of spoiled ballots is unknown and varies by election and possibly by ballot style: historical data do not help much. Hence, there is no way to calibrate a threshold for passive testing, e.g., to guarantee at least a 95% chance of noticing that 5% of the votes were altered, with at most a 5% false alarm rate. (vi) Even if the distribution of spoiled ballots were known to be Poisson, the vast majority of jurisdictions do not have enough voters for passive testing to have a large chance of detecting problems but only a small chance of false alarms.

stat.AP

Self-similar curve shortening flow in hyperbolic 2-space

We find and classify self-similar solutions of the curve shortening flow in standard hyperbolic 2-space. Together with earlier work of Halldórsson on curve shortening flow in the plane and Santos dos Reis and Tenenblat in the 2-sphere, this completes the classification of self-similar curve shortening flows in the constant curvature model spaces in 2-dimensions.

math.DG