SearcharxivSearch

arXiv subjects

Xiaocong Xu

Publications and source records attributed to Xiaocong Xu.

14 recordsLinked to original sources

Harnessing Unimodality in Semiparametric Contextual Pricing via Oracle Price Map Learning

We study contextual dynamic pricing in a semiparametric scalar-index valuation model where the latent value is $v_t=μ_\ast(\mathsf c_t)+ξ_t$, with an unknown utility map $μ_\ast$ and an unknown additive noise distribution. The key decision object is the one-dimensional oracle price map $u\mapsto p^\ast(u)$ induced by the scalar index $u=μ_\ast(\mathsf c)$ and the noise tail. Under the $β$-Hölder smoothness of the tail function for $β\geq 2$ and a revenue-geometry condition that gives a unique, stable, interior maximizer, this oracle map is itself $(β-1)$-smooth. We exploit such structure through $\mathsf{ORBIT}$, a modular coarse-to-fine policy that takes a scalar pilot index as input, localizes a benchmark price in each active bin, and learns a local polynomial approximation of the oracle map inside a trust region via bandit convex optimization. For the baseline linear utility model $μ_\ast(\mathsf c)=\mathsf c^\topθ_\ast$, an adaptive elliptical exploration scheme constructs the required scalar pilot online without distributional assumptions on the contexts. The resulting policy achieves regret $\widetilde{O}\big(T^{\frac{2β-1}{4β-3}}+\sqrt{dT}\big)$. For fixed $d$, we establish a matching lower bound in the horizon dependence, unveiling that the nonparametric oracle-map learning term is minimax sharp. The same scalar-pilot interface also yields extensions to sparse high-dimensional linear utility and nonparametric Hölder utility.

stat.ML

The distribution of Ridgeless least squares interpolators

The Ridgeless minimum $\ell_2$-norm interpolator in overparametrized linear regression has attracted considerable attention in recent years in both machine learning and statistics communities. While it seems to defy conventional wisdom that overfitting leads to poor prediction, recent theoretical research on its $\ell_2$-type risks reveals that its norm minimizing property induces an `implicit regularization' that helps prediction in spite of interpolation. This paper takes a further step that aims at understanding its precise stochastic behavior as a statistical estimator. Specifically, we characterize the distribution of the Ridgeless interpolator in high dimensions, in terms of a Ridge estimator in an associated Gaussian sequence model with positive regularization, which provides a precise quantification of the prescribed implicit regularization in the most general distributional sense. Our distributional characterizations hold for general non-Gaussian random designs and extend uniformly to positively regularized Ridge estimators. As a direct application, we obtain a complete characterization for a general class of weighted $\ell_q$ risks of the Ridge(less) estimators that are previously only known for $q=2$ by random matrix methods. These weighted $\ell_q$ risks not only include the standard prediction and estimation errors, but also include the non-standard covariate shift settings. Our uniform characterizations further reveal a surprising feature of the commonly used generalized and $k$-fold cross-validation schemes: tuning the estimated $\ell_2$ prediction risk by these methods alone lead to simultaneous optimal $\ell_2$ in-sample, prediction and estimation risks, as well as the optimal length of debiased confidence intervals.

math.ST

Gradient descent inference in empirical risk minimization

Gradient descent is one of the most widely used iterative algorithms in modern statistical learning. However, its precise algorithmic dynamics in high-dimensional settings remain only partially understood, which has limited its broader potential for statistical inference applications. This paper provides a precise, non-asymptotic joint distributional characterization of gradient descent iterates and their debiased statistics in a broad class of empirical risk minimization problems, in the so-called mean-field regime where the sample size is proportional to the signal dimension. Our non-asymptotic state evolution theory holds for both general non-convex loss functions and non-Gaussian data, and reveals the central role of two Onsager correction matrices that precisely characterize the non-trivial dependence among all gradient descent iterates in the mean-field regime. Leveraging the joint state evolution characterization, we show that the gradient descent iterate retrieves approximate normality after a debiasing correction via a linear combination of all past iterates, where the debiasing coefficients can be estimated by the proposed gradient descent inference algorithm. This leads to a new algorithmic statistical inference framework based on debiased gradient descent, which (i) applies to a broad class of models with both convex and non-convex losses, (ii) remains valid at each iteration without requiring algorithmic convergence, and (iii) exhibits a certain robustness to possible model misspecification. As a by-product, our framework also provides algorithmic estimates of the generalization error at each iteration. As canonical examples, we demonstrate our theory and inference methods in the single-index regression model and a generalized logistic regression model, where the natural loss functions may exhibit arbitrarily non-convex landscapes.

math.ST

A leave-one-out approach to approximate message passing

Approximate message passing (AMP) has emerged both as a popular class of iterative algorithms and as a powerful analytic tool in a wide range of statistical estimation problems and statistical physics models. A well established line of AMP theory proves Gaussian approximations for the empirical distributions of the AMP iterate in the high dimensional limit, under the GOE random matrix model and its variants. This paper provides a non-asymptotic, leave-one-out representation for the AMP iterate that holds under a broad class of Gaussian random matrix models with general variance profiles. In contrast to the typical AMP theory that describes the empirical distributions of the AMP iterate via a low dimensional state evolution, our leave-one-out representation yields an intrinsically high dimensional state evolution formula which provides non-asymptotic characterizations for the possibly heterogeneous, entrywise behavior of the AMP iterate under the prescribed random matrix models. To exemplify some distinct features of our AMP theory in applications, we analyze, in the context of regularized linear estimation, the precise stochastic behavior of the Ridge estimator for independent and non-identically distributed observations whose covariates exhibit general variance profiles. We find that its finite-sample distribution is characterized via a weighted Ridge estimator in a heterogeneous Gaussian sequence model. Notably, in contrast to the i.i.d. sampling scenario, the effective noise and regularization are now full dimensional vectors determined via a high dimensional system of equations. Our leave-one-out method of proof differs significantly from the widely adopted conditioning approach for rotational invariant ensembles, and relies instead on an inductive method that utilizes almost solely integration-by-parts and concentration techniques.

math.ST

Precise Asymptotics and Refined Regret of Variance-Aware UCB

In this paper, we study the behavior of the Upper Confidence Bound-Variance (UCB-V) algorithm for the Multi-Armed Bandit (MAB) problems, a variant of the canonical Upper Confidence Bound (UCB) algorithm that incorporates variance estimates into its decision-making process. More precisely, we provide an asymptotic characterization of the arm-pulling rates for UCB-V, extending recent results for the canonical UCB in Kalvit and Zeevi (2021) and Khamaru and Zhang (2024). In an interesting contrast to the canonical UCB, our analysis reveals that the behavior of UCB-V can exhibit instability, meaning that the arm-pulling rates may not always be asymptotically deterministic. Besides the asymptotic characterization, we also provide non-asymptotic bounds for the arm-pulling rates in the high probability regime, offering insights into the regret analysis. As an application of this high probability result, we establish that UCB-V can achieve a more refined regret bound, previously unknown even for more complicate and advanced variance-aware online decision-making algorithms.

stat.ML

Asymptotic FDR Control with Model-X Knockoffs: Is Moments Matching Sufficient?

We propose a unified theoretical framework for studying the robustness of the model-X knockoffs framework by investigating the asymptotic false discovery rate (FDR) control of the practically implemented approximate knockoffs procedure. This procedure deviates from the model-X knockoffs framework by substituting the true covariate distribution with a user-specified distribution that can be learned using in-sample observations. By replacing the distributional exchangeability condition of the model-X knockoff variables with three conditions on the approximate knockoff statistics, we establish that the approximate knockoffs procedure achieves the asymptotic FDR control. Using our unified framework, we further prove that an arguably most popularly used knockoff variable generation method--the Gaussian knockoffs generator based on the first two moments matching--achieves the asymptotic FDR control when the two-moment-based knockoff statistics are employed in the knockoffs inference procedure. For the first time in the literature, our theoretical results justify formally the effectiveness and robustness of the Gaussian knockoffs generator. Simulation and real data examples are conducted to validate the theoretical findings.

stat.ML

Phase transition for the bottom singular vector of rectangular random matrices

In this paper, we consider the rectangular random matrix $X=(x_{ij})\in \mathbb{R}^{N\times n}$ whose entries are iid with tail $\mathbb{P}(|x_{ij}|>t)\sim t^{-α}$ for some $α>0$. We consider the regime $N(n)/n\to \mathsf{a}>1$ as $n$ tends to infinity. Our main interest lies in the right singular vector corresponding to the smallest singular value, which we will refer to as the "bottom singular vector", denoted by $\mathfrak{u}$. In this paper, we prove the following phase transition regarding the localization length of $\mathfrak{u}$: when $α<2$ the localization length is $O(n/\log n)$; when $α>2$ the localization length is of order $n$. Similar results hold for all right singular vectors around the smallest singular value. The variational definition of the bottom singular vector suggests that the mechanism for this localization-delocalization transition when $α$ goes across $2$ is intrinsically different from the one for the top singular vector when $α$ goes across $4$.

math.PR

Phase transition for the smallest eigenvalue of covariance matrices

In this paper, we study the smallest non-zero eigenvalue of the sample covariance matrices $\mathcal{S}(Y)=YY^*$, where $Y=(y_{ij})$ is an $M\times N$ matrix with iid mean $0$ variance $N^{-1}$ entries. We prove a phase transition for its distribution, induced by the fatness of the tail of $y_{ij}$'s. More specifically, we assume that $y_{ij}$ is symmetrically distributed with tail probability $\mathbb{P}(|\sqrt{N}y_{ij}|\geq x)\sim x^{-α}$ when $x\to \infty$, for some $α\in (2,4)$. We show the following conclusions: (i). When $α>\frac83$, the smallest eigenvalue follows the Tracy-Widom law on scale $N^{-\frac23}$; (ii). When $2<α<\frac83$, the smallest eigenvalue follows the Gaussian law on scale $N^{-\fracα{4}}$; (iii). When $α=\frac83$, the distribution is given by an interpolation between Tracy-Widom and Gaussian; (iv). In case $α\leq \frac{10}{3}$, in addition to the left edge of the MP law, a deterministic shift of order $N^{1-\fracα{2}}$ shall be subtracted from the smallest eigenvalue, in both the Tracy-Widom law and the Gaussian law. Overall speaking, our proof strategy is inspired by \cite{ALY} which is originally done for the bulk regime of the Lévy Wigner matrices. In addition to various technical complications arising from the bulk-to-edge extension, two ingredients are needed for our derivation: an intermediate left edge local law based on a simple but effective matrix minor argument, and a mesoscopic CLT for the linear spectral statistic with asymptotic expansion for its expectation.

math.PR

Extreme eigenvalues of Log-concave Ensemble

In this paper, we consider the log-concave ensemble of random matrices, a class of covariance-type matrices $XX^*$ with isotropic log-concave $X$-columns. A main example is the covariance estimator of the uniform measure on isotropic convex body. Non-asymptotic estimates and first order asymptotic limits for the extreme eigenvalues have been obtained in the literature. In this paper, with the recent advancements on log-concave measures \cite{chen, KL22}, we take a step further to locate the eigenvalues with a nearly optimal precision, namely, the spectral rigidity of this ensemble is derived. Based on the spectral rigidity and an additional ``unconditional" assumption, we further derive the Tracy-Widom law for the extreme eigenvalues of $XX^*$, and the Gaussian law for the extreme eigenvalues in case strong spikes are present.

math.PR

Spectral Statistics of Sample Block Correlation Matrices

A fundamental concept in multivariate statistics, sample correlation matrix, is often used to infer the correlation/dependence structure among random variables, when the population mean and covariance are unknown. A natural block extension of it, {\it sample block correlation matrix}, is proposed to take on the same role, when random variables are generalized to random sub-vectors. In this paper, we establish a spectral theory of the sample block correlation matrices and apply it to group independent test and related problem, under the high-dimensional setting. More specifically, we consider a random vector of dimension $p$, consisting of $k$ sub-vectors of dimension $p_t$'s, where $p_t$'s can vary from $1$ to order $p$. Our primary goal is to investigate the dependence of the $k$ sub-vectors. We construct a random matrix model called sample block correlation matrix based on $n$ samples for this purpose. The spectral statistics of the sample block correlation matrix include the classical Wilks' statistic and Schott's statistic as special cases. It turns out that the spectral statistics do not depend on the unknown population mean and covariance. Further, under the null hypothesis that the sub-vectors are independent, the limiting behavior of the spectral statistics can be described with the aid of the Free Probability Theory. Specifically, under three different settings of possibly $n$-dependent $k$ and $p_t$'s, we show that the empirical spectral distribution of the sample block correlation matrix converges to the free Poisson binomial distribution, free Poisson distribution (Marchenko-Pastur law) and free Gaussian distribution (semicircle law), respectively. We then further derive the CLTs for the linear spectral statistics of the block correlation matrix under general setting.

math.ST

Improving Global Forest Mapping by Semi-automatic Sample Labeling with Deep Learning on Google Earth Images

Global forest cover is critical to the provision of certain ecosystem services. With the advent of the google earth engine cloud platform, fine resolution global land cover mapping task could be accomplished in a matter of days instead of years. The amount of global forest cover (GFC) products has been steadily increasing in the last decades. However, it's hard for users to select suitable one due to great differences between these products, and the accuracy of these GFC products has not been verified on global scale. To provide guidelines for users and producers, it is urgent to produce a validation sample set at the global level. However, this labeling task is time and labor consuming, which has been the main obstacle to the progress of global land cover mapping. In this research, a labor-efficient semi-automatic framework is introduced to build a biggest ever Forest Sample Set (FSS) contained 395280 scattered samples categorized as forest, shrubland, grassland, impervious surface, etc. On the other hand, to provide guidelines for the users, we comprehensively validated the local and global mapping accuracy of all existing 30m GFC products, and analyzed and mapped the agreement of them. Moreover, to provide guidelines for the producers, optimal sampling strategy was proposed to improve the global forest classification. Furthermore, a new global forest cover named GlobeForest2020 has been generated, which proved to improve the previous highest state-of-the-art accuracies (obtained by Gong et al., 2017) by 2.77% in uncertain grids and by 1.11% in certain grids.

cs.CV

General Implicit Iterative Method for Unified Gas-kinetic Scheme

In order to further enhance the computational efficiency of the implicit unified gas-kinetic scheme (IUGKS, JCP 315 (2016) 16-38) for multi-scale flow simulation, a two-step IUGKS is proposed in this paper. The multiscale solution of the UGKS is determined by the integral solution of the kinetic model equation, which is composed of the Lagrangian integration of the equilibrium and the free particle transport of the nonequilibrium state. With the implicit evaluation of the macroscopic variables in the first iterative step, the integration of the equilibrium can be directly used in the flux calculation of macroscopic flow variables in the second iterative step. This is equivalent to include the viscous flux in the implicit scheme to accelerate the convergence of the solution instead of using the Euler flux in the iterative process of the original IUGKS. In the present IUGKS, the update of macroscopic flow variables are closely coupled with the implicit evolution of the gas distribution function. At the same time, in order to get the more accurate solution, the full Boltzmann collision term is incorporated into the current scheme through the penalty method. Different iterative techniques, such as LU-SGS and multi-grid, are used for solving the linear algebraic system of coupled macroscopic and microscopic equations. The efficiency of the IUGKS has reached a favorable level among all implicit schemes for the kinetic equations in the literature. Several numerical examples are used to validate the performance of the IUGKS. Accurate solutions have been obtained efficiently in all flow regimes from low speed to hypersonic ones.

physics.comp-ph

Modeling and computation for non-equilibrium gas dynamics: beyond kinetic relaxation model

The non-equilibrium gas dynamics is described by the Boltzmann equation, which can be solved numerically through the deterministic and stochastic methods. Due to the complicated collision term of the Boltzmann equation, many kinetic relaxation models have been proposed and used in the past seventy years for the study of rarefied flow. In order to develop a multiscale method for the rarefied and continuum flow simulation, by adopting the integral solution of the kinetic model equation a DVM-type unified gas-kinetic scheme (UGKS) has been constructed. The UGKS models the gas dynamics on the cell size and time step scales while the accumulating effect from particle transport and collision has been taken into account within a time step. Under the UGKS framework, a unified gas-kinetic wave-particle (UGKWP) method has been further developed for non-equilibrium flow simulation, where the time evolution of gas distribution function is composed of analytical wave and individual particle. In the highly rarefied regime, particle transport and collision will play a dominant role. Due to the single relaxation time model for particle collision, there is a noticeable discrepancy between the UGKWP solution and the full Boltzmann or DSMC result, especially in the high Mach and Knudsen number cases. In this paper, besides the kinetic relaxation model, a modification of particle collision time according to the particle velocity will be implemented in UGKWP. As a result, the new model greatly improves the performance of UGKWP in the capturing of non-equilibrium flow. There is a perfect match between UGKWP and DSMC or Boltzmann solution in the highly rarefied regime. In the near continuum and continuum flow regime, the UGKWP will gradually get back to the macroscopic variables based Navier-Stokes flow solver at small cell Knudsen number.

physics.comp-ph

Unified gas-kinetic wave-particle methods V: diatomic molecular flow

In this paper, the unified gas-kinetic wave-particle (UGKWP) method is further developed for diatomic gas with the energy exchange between translational and rotational modes for flow study in all regimes. The multiscale transport mechanism in UGKWP is coming from the direct modeling in a discretized space, where the cell's Knudsen number, defined by the ratio of particle mean free path over the numerical cell size, determines the flow physics simulated by the wave particle formulation. The non-equilibrium distribution function in UGKWP is tracked by the discrete particle and analytical wave. The weights of distributed particle and wave in different regimes are controlled by the accumulating evolution solution of particle transport and collision within a time step, where distinguishable macroscopic flow variables of particle and wave are updated inside each control volume. With the variation of local cell's Knudsen number, the UGKWP becomes a particle method in the highly rarefied flow regime and converges to the gas-kinetic scheme (GKS) for the Navier-Stokes solution in the continuum flow regime without particles. Even targeting on the same solution as the discrete velocity method (DVM)-based unified gas-kinetic scheme (UGKS), the computational cost and memory requirement in UGKWP could be reduced by several orders of magnitude for the high speed and high temperature flow simulation, where the translational and rotational non-equilibrium becomes important in the transition and rarefied regime. As a result, 3D hypersonic computations around a flying vehicle in all regimes can be conducted using a personal computer. The UGKWP method for diatomic gas will be validated in various cases from one dimensional shock structure to three dimensional flow over a sphere, and the numerical solutions will be compared with the reference DSMC results and experimental measurements.

physics.comp-ph