SearcharxivSearch

arXiv subjects

Shunan Sheng

Publications and source records attributed to Shunan Sheng.

9 recordsLinked to original sources

Empirical optimal transport potentials: fast rates and a functional central limit theorem

Optimal transport potentials are fundamental objects in statistics, economics, and machine learning: their gradients generate optimal transport maps, while the potentials themselves act as location-dependent dual prices and sensitivity variables. We study the estimation of the quadratic optimal transport potential when a fixed absolutely continuous reference distribution $\mu$ is transported to an unknown distribution $\nu$, accessed to via its empirical measure. Our main ingredient is a stability inequality that controls the $L^1(\mu)$ distance, modulo additive constants, between a strongly convex potential $\varphi$ and a convex potential $\widetilde\varphi$ by a weak dual norm of $(\nabla\widetilde\varphi)_\#\mu-(\nabla\varphi)_\#\mu,$ together with a second-order Wasserstein remainder of logarithmic type. This separation between the leading empirical-process term and the Wasserstein remainder yields faster convergence for potentials than for the corresponding transport maps. Under smoothness and uniform convexity assumptions, the exact semidiscrete Brenier potential converges in $L^1(\mu)$ at rate $n^{-1/2}$ for $d\leq3$, at rate $n^{-1/2}(\log n)^{5/2}$ for $d=4$, and at rate $n^{-2/d}(\log n)^{(d+2)/d}$ for $d\geq5$. The polynomial exponents are sharp. In dimensions $d\leq3$, we further establish a nondegenerate function-space central limit theorem and prove consistency of the nonparametric bootstrap. These results yield joint root-$n$ inference for every fixed finite collection of normalization-invariant weighted contrasts of the potential, including regional shadow premia in reference-based risk problems. Finally, we prove matching upper and lower bounds of order $\varepsilon\log(1/\varepsilon)$ for the normalization-invariant sum of the entropic dual potentials.

math.ST

The Influence Function of Transport-based Quantiles

Transport-based quantiles extend univariate quantiles to multivariate distributions via optimal transport. We study the influence function of the transport quantile map $\mathbf{Q}_P$, defined as the optimal transport map pushing a fixed reference measure $\mu$ forward to a target distribution $P$. For the Huber contamination $P_t=(1-t)P+t\delta_{x_0}$, we prove that the first-order limit $\mathbf{I}(x_0;\mathbf{Q}_P(z)) := \lim_{t\downarrow 0} [\mathbf{Q}_{(1-t)P+t\delta_{x_0}}(z)-\mathbf{Q}_P(z)]/t$ exists whenever $x_0\ne \mathbf{Q}_P(z)$ and characterize it uniquely. Specifically, $\mathbf{I}(x_0;\mathbf{Q}_P(z))=\nabla G_{x_0}(z)$, where $G_{x_0}$ is characterized by a uniformly elliptic equation with a Dirac source and a Neumann boundary condition. In every dimension $d\ge 2$, this influence function has a pole-type singularity. For fixed $z\in\operatorname{int}(\Omega_\mu)$, it remains bounded when $\mathbf{F}_P(x_0)$ stays away from $z$, where $\mathbf{F}_P=\mathbf{Q}_P^{-1}$ is the transport-based distribution function, but diverges as $x_0\to\mathbf{Q}_P(z)$, equivalently as $\mathbf{F}_P(x_0)\to z$. In fact, $\|\mathbf{I}(x_0;\mathbf{Q}_P(z))\|\asymp\|z-\mathbf{F}_P(x_0)\|^{-(d-1)}$. This contrasts with the bounded influence function of univariate quantiles and implies that $\mathbf{I}(X;\mathbf{Q}_P(z))$, for $X\sim P$, has infinite second moment. Numerical experiments further suggest that empirical transport quantiles may exhibit stable-type non-Gaussian fluctuations.

math.ST

Bid--Ask Martingale Optimal Transport

Martingale Optimal Transport (MOT) provides a framework for robust pricing and hedging of illiquid derivatives. Classical MOT enforces exact calibration of model marginals to the mid-prices of vanilla options. Motivated by the industry practice of fitting bid and ask marginals to vanilla prices, we introduce a relaxation of MOT in which model-implied volatilities are only required to lie within observed bid--ask spreads; equivalently, model marginals lie between the bid and ask marginals in convex order. The resulting Bid--Ask MOT (BAMOT) yields realistic price bounds for illiquid derivatives and, via strong duality, can be interpreted as the superhedging price when short and long positions in vanilla options are priced at the bid and ask, respectively. We further establish convergence of BAMOT to classical MOT as bid--ask spreads vanish, and quantify the convergence rate using a novel distance intrinsically linked to bid--ask spreads. Finally, we support our findings with several synthetic and real-data examples.

q-fin.MF

Theory and computation for structured variational inference

Structured variational inference constitutes a core methodology in modern statistical applications. Unlike mean-field variational inference, the approximate posterior is assumed to have interdependent structure. We consider the natural setting of star-structured variational inference, where a root variable impacts all the other ones. We prove the first results for existence, uniqueness, and self-consistency of the variational approximation. In turn, we derive quantitative approximation error bounds for the variational approximation to the posterior, extending prior work from the mean-field setting to the star-structured setting. We also develop a gradient-based algorithm with provable guarantees for computing the variational approximation using ideas from optimal transport theory. We explore the implications of our results for Gaussian measures and hierarchical Bayesian models, including generalized linear models with location family priors and spike-and-slab priors with one-dimensional debiasing. As a by-product of our analysis, we develop new stability results for star-separable transport maps which might be of independent interest.

stat.ML

Mode Collapse of Mean-Field Variational Inference

Mean-field variational inference (MFVI) is a widely used method for approximating high-dimensional probability distributions by product measures. It has been empirically observed that MFVI optimizers often suffer from mode collapse. Specifically, when the target measure $\pi$ is a mixture $\pi = w P_0 + (1 - w) P_1$, the MFVI optimizer tends to place most of its mass near a single component of the mixture. This work provides the first theoretical explanation of mode collapse in MFVI. We introduce the notion to capture the separatedness of the two mixture components -- called $\varepsilon$-separateness -- and derive explicit bounds on the fraction of mass that any MFVI optimizer assigns to each component when $P_0$ and $P_1$ are $\varepsilon$-separated for sufficiently small $\varepsilon$. Our results suggest that the occurrence of mode collapse crucially depends on the relative position of the components. To address this issue, we propose the rotational variational inference (RoVI), which augments MFVI with a rotation matrix. The numerical studies support our theoretical findings and demonstrate the benefits of RoVI.

stat.ML

Stability of Mean-Field Variational Inference

Mean-field variational inference (MFVI) is a widely used method for approximating high-dimensional probability distributions by product measures. This paper studies the stability properties of the mean-field approximation when the target distribution varies within the class of strongly log-concave measures. We establish dimension-free Lipschitz continuity of the MFVI optimizer with respect to the target distribution, measured in the 2-Wasserstein distance, with Lipschitz constant inversely proportional to the log-concavity parameter. Under additional regularity conditions, we further show that the MFVI optimizer depends differentiably on the target potential and characterize the derivative by a partial differential equation. Methodologically, we follow a novel approach to MFVI via linearized optimal transport: the non-convex MFVI problem is lifted to a convex optimization over transport maps with a fixed base measure, enabling the use of calculus of variations and functional analysis. We discuss several applications of our results to robust Bayesian inference and empirical Bayes, including a quantitative Bernstein--von Mises theorem for MFVI, as well as to distributed stochastic control.

math.PR

Linearization of Monge-Amp\`ere Equations and Statistical Applications

Optimal transport has found numerous applications across data science, many of which require differentiating the optimal transport map with respect to the underlying probability densities in the Fr\'echet sense. In this work, we show that when the reference measure $Q$ is sufficiently regular in space and the curve of target measures $\{P_t\}_{t\in I}$ is both spatially regular and $\mathcal{C}^1$ in time, then the associated curve of optimal transport maps $\{\nabla \phi_t\}_{t\in I}$ pushing $Q$ toward $P_t$ is itself a $\mathcal{C}^1$ curve. Moreover, we identify its time derivative as the solution to the \emph{linearized Monge--Amp\`ere equation}, a second-order elliptic PDE with strictly oblique boundary conditions and a vanishing zero-order term. Our proof relies on applying the implicit function theorem to the Monge--Amp\`ere equation with natural boundary conditions. As consequences, we establish regularity of the transport-based quantile regressor with respect to the covariates and derive a central limit theorem for smooth optimal transport maps.

math.AP

Binary Spatial Random Field Reconstruction from Non-Gaussian Inhomogeneous Time-series Observations

We develop a new model for spatial random field reconstruction of a binary-valued spatial phenomenon. In our model, sensors are deployed in a wireless sensor network across a large geographical region. Each sensor measures a non-Gaussian inhomogeneous temporal process which depends on the spatial phenomenon. Two types of sensors are employed: one collects point observations at specific time points, while the other collects integral observations over time intervals. Subsequently, the sensors transmit these time-series observations to a Fusion Center (FC), and the FC infers the spatial phenomenon from these observations. We show that the resulting posterior predictive distribution is intractable and develop a tractable two-step procedure to perform inference. Firstly, we develop algorithms to perform approximate Likelihood Ratio Tests on the time-series observations, compressing them to a single bit for both point sensors and integral sensors. Secondly, once the compressed observations are transmitted to the FC, we utilize a Spatial Best Linear Unbiased Estimator (S-BLUE) to reconstruct the binary spatial random field at any desired spatial location. The performance of the proposed approach is studied using simulation. We further illustrate the effectiveness of our method using a weather dataset from the National Environment Agency (NEA) of Singapore with fields including temperature and relative humidity.

eess.SP

Balanced Meta-Softmax for Long-Tailed Visual Recognition

Deep classifiers have achieved great success in visual recognition. However, real-world data is long-tailed by nature, leading to the mismatch between training and testing distributions. In this paper, we show that the Softmax function, though used in most classification tasks, gives a biased gradient estimation under the long-tailed setup. This paper presents Balanced Softmax, an elegant unbiased extension of Softmax, to accommodate the label distribution shift between training and testing. Theoretically, we derive the generalization bound for multiclass Softmax regression and show our loss minimizes the bound. In addition, we introduce Balanced Meta-Softmax, applying a complementary Meta Sampler to estimate the optimal class sample rate and further improve long-tailed learning. In our experiments, we demonstrate that Balanced Meta-Softmax outperforms state-of-the-art long-tailed classification solutions on both visual recognition and instance segmentation tasks.

cs.LG