SearcharxivSearch

arXiv subjects

Shuting Shen

Publications and source records attributed to Shuting Shen.

6 recordsLinked to original sources

Post-Learning Inference for Combinatorial Optimizers with High-Dimensional Sparse Contextual Information via Minimal Directional Perturbation

We study post-learning inference for structural properties of data-dependent combinatorial optimizers. The target is whether an oracle optimizer, rather than a latent parameter or smooth functional, belongs to a prescribed class, such as a category-mix, inventory, or resource-feasibility class. We focus on a high-dimensional contextual multinomial logit model with sequentially adaptive data collection, where the parameter-to-optimizer map is discontinuous and the policy induces temporal dependence. We propose a novel perturbation test based on a nonsmooth max-difference revenue statistic comparing the best null assortment with the best alternative assortment. The test perturbs the estimated terminal revenue surface on the selected support: random unit directions capture directional uncertainty, while the minimal perturbation radius captures magnitude uncertainty and yields a p-value. This localizes inference near the null--alternative boundary and avoids uniform error control over the full candidate class. The data are collected by an \(\ell_1\)-penalized online likelihood policy that performs variable selection while controlling regret. Using a new anti-concentration argument for Gaussian maxima differences and martingale Gaussian coupling, we establish uniform estimation rates, effective support recovery, and asymptotic validity of the proposed p-value under adaptive assortment selection. We prove asymptotic size control and power consistency under a localized signal condition.

math.ST

Anti-Concentration Inequalities for the Difference of Maxima of Gaussian Random Vectors

We derive novel anti-concentration bounds for the difference between the maximal values of two Gaussian random vectors across various settings. Our bounds are dimension-free, scaling with the dimension of the Gaussian vectors only through the smaller expected maximum of the Gaussian subvectors. In addition, our bounds hold under the degenerate covariance structures, which previous results do not cover. In addition, we show that our conditions are sharp under the homogeneous component-wise variance setting, while we only impose some mild assumptions on the covariance structures under the heterogeneous variance setting. We apply the new anti-concentration bounds to derive the central limit theorem for the maximizers of discrete empirical processes. Finally, we back up our theoretical findings with comprehensive numerical studies.

math.ST

Inference of Dependency Knowledge Graph for Electronic Health Records

The effective analysis of high-dimensional Electronic Health Record (EHR) data, with substantial potential for healthcare research, presents notable methodological challenges. Employing predictive modeling guided by a knowledge graph (KG), which enables efficient feature selection, can enhance both statistical efficiency and interpretability. While various methods have emerged for constructing KGs, existing techniques often lack statistical certainty concerning the presence of links between entities, especially in scenarios where the utilization of patient-level EHR data is limited due to privacy concerns. In this paper, we propose the first inferential framework for deriving a sparse KG with statistical guarantee based on the dynamic log-linear topic model proposed by \cite{arora2016latent}. Within this model, the KG embeddings are estimated by performing singular value decomposition on the empirical pointwise mutual information matrix, offering a scalable solution. We then establish entrywise asymptotic normality for the KG low-rank estimator, enabling the recovery of sparse graph edges with controlled type I error. Our work uniquely addresses the under-explored domain of statistical inference about non-linear statistics under the low-rank temporal dependent models, a critical gap in existing research. We validate our approach through extensive simulation studies and then apply the method to real-world EHR data in constructing clinical KGs and generating clinical feature embeddings.

stat.ME

Dimension Reduction for Large-Scale Federated Data: Statistical Rate and Asymptotic Inference

In light of the rapidly growing large-scale data in federated ecosystems, the traditional principal component analysis (PCA) is often not applicable due to privacy protection considerations and large computational burden. Algorithms were proposed to lower the computational cost, but few can handle both high dimensionality and massive sample size under distributed settings. In this paper, we propose the FAst DIstributed (FADI) PCA method for federated data when both the dimension $d$ and the sample size $n$ are ultra-large, by simultaneously performing parallel computing along $d$ and distributed computing along $n$. Specifically, we utilize $L$ parallel copies of $p$-dimensional fast sketches to divide the computing burden along $d$ and aggregate the results distributively along the split samples. We present a general framework applicable to multiple statistical problems, and establish comprehensive theoretical results under the general framework. We show that FADI accelerates the computation while enjoying the same non-asymptotic error rate as the traditional PCA when $Lp \ge d$. We also derive inferential results that characterize the asymptotic distribution of FADI, and show a phase-transition phenomenon as $Lp$ increases. We perform extensive simulations to empirically validate our theoretical findings, and apply FADI to the 1000 Genomes data to study the population structure.

stat.ME

Combinatorial Inference on the Optimal Assortment in Multinomial Logit Models

Assortment optimization has received active explorations in the past few decades due to its practical importance. Despite the extensive literature dealing with optimization algorithms and latent score estimation, uncertainty quantification for the optimal assortment still needs to be explored and is of great practical significance. Instead of estimating and recovering the complete optimal offer set, decision-makers may only be interested in testing whether a given property holds true for the optimal assortment, such as whether they should include several products of interest in the optimal set, or how many categories of products the optimal set should include. This paper proposes a novel inferential framework for testing such properties. We consider the widely adopted multinomial logit (MNL) model, where we assume that each customer will purchase an item within the offered products with a probability proportional to the underlying preference score associated with the product. We reduce inferring a general optimal assortment property to quantifying the uncertainty associated with the sign change point detection of the marginal revenue gaps. We show the asymptotic normality of the marginal revenue gap estimator, and construct a maximum statistic via the gap estimators to detect the sign change point. By approximating the distribution of the maximum statistic with multiplier bootstrap techniques, we propose a valid testing procedure. We also conduct numerical experiments to assess the performance of our method.

stat.ML

Combinatorial-Probabilistic Trade-Off: Community Properties Test in the Stochastic Block Models

In this paper, we propose an inferential framework testing the general community combinatorial properties of the stochastic block model. Instead of estimating the community assignments, we aim to test the hypothesis on whether a certain community property is satisfied. For instance, we propose to test whether a given set of nodes belong to the same community or whether different network communities have the same size. We propose a general inference framework that can be applied to all symmetric community properties. To ease the challenges caused by the combinatorial nature of communities properties, we develop a novel shadowing bootstrap testing method. By utilizing the symmetry, our method can find a shadowing representative of the true assignment and the number of assignments to be tested in the alternative can be largely reduced. In theory, we introduce a combinatorial distance between two community classes and show a combinatorial-probabilistic trade-off phenomenon in the community properties test. Our test is honest as long as the product of combinatorial distance between two communities and the probabilistic distance between two assignment probabilities is sufficiently large. On the other hand, we shows that such trade-off also exists in the information-theoretic lower bound of the community property test. We also implement numerical experiments on both the synthetic data and the protein interaction application to show the validity of our method.

math.ST