SearcharxivSearch

arXiv subjects

Martin Chak

Publications and source records attributed to Martin Chak.

7 recordsLinked to original sources

Strong log-concavity in probit regression

We show that strong log-concavity emerges in probit regression likelihoods without ridge penalization (i.e. Gaussian priors), unlike for the logistic case. Specifically, we provide: (a) a non-asymptotic characterization of strong log-concavity for fixed designs, similar to that for the existence of the maximum likelihood estimator (MLE) and (b) an asymptotic analysis for random designs, in the proportional regime when the sample size $n$ and the number of covariates $d$ grow proportionally $d/n\rightarrow r\in[0,1)$. In the latter case we show that, provided $r$ is small enough, the resulting condition number is finite and independent of $n,d,r$ with high-probability. Numerically tractable estimates are given in case $r=0$. Thus, probit regression provides a non-trivial example of high-dimensional, well-conditioned log-concave objective.

math.ST

Complexity of Markov Chain Monte Carlo for Generalized Linear Models

Markov Chain Monte Carlo (MCMC), Laplace approximation (LA) and variational inference (VI) methods are popular approaches to Bayesian inference, each with trade-offs between computational cost and accuracy. However, a theoretical understanding of these differences is missing, particularly when both the sample size $n$ and the dimension $d$ are large. LA and Gaussian VI are justified by Bernstein-von Mises (BvM) theorems, and recent work has derived the characteristic condition $n\gg d^2$ for their validity, improving over the condition $n\gg d^3$. In this paper, we show for linear, logistic and Poisson regression that for $n\gtrsim d$, MCMC attains the same complexity scaling in $n$, $d$ as first-order optimization algorithms, up to sub-polynomial factors. Thus MCMC is competitive with LA and Gaussian VI in complexity, under a scaling between $n$ and $d$ more general than BvM regimes. Our complexities apply to appropriately scaled priors that are not necessarily Gaussian-tailed, including Student-$t$ and flat priors, with log-posteriors that are not necessarily globally concave or gradient-Lipschitz.

stat.CO

On theoretical guarantees and a blessing of dimensionality for nonconvex sampling

Guarantees for algorithms sampling from nonlogconcave target measures on $\mathbb{R}^d$ are studied. For the class of measures with logdensities that have bounded Hessians and are strongly concave outside a Euclidean ball of radius $R$, it is shown that complete polynomial complexity can in fact be achieved if $R\leq c\sqrt{d}$. On the other hand, an exponential number of point evaluations is shown to be generally necessary for any algorithm as soon as $R\geq C\sqrt{d}$ for constants $C>c>0$. Importance sampling with a tail-matching proposal achieves the former, owing to a blessing of dimensionality. It is also shown that if strong concavity outside a ball is replaced by a distant dissipativity condition, then sampling guarantees must generally scale exponentially with $d$ in essentially all parameter regimes.

stat.CO

Reflection coupling for unadjusted generalized Hamiltonian Monte Carlo in the nonconvex stochastic gradient case

Contraction in Wasserstein 1-distance with explicit rates is established for generalized Hamiltonian Monte Carlo with stochastic gradients under possibly nonconvex conditions. The algorithms considered include splitting schemes of kinetic Langevin diffusion commonly used in molecular dynamics simulations. To accommodate the degenerate noise structure corresponding to inertia existing in the chain, a characteristically discrete-in-time coupling and contraction proof is devised. As consequence, quantitative Gaussian concentration bounds are provided for empirical averages. Convergence in Wasserstein 2-distance and total variation are also given, together with numerical bias estimates.

math.PR

Regularity preservation in Kolmogorov equations for non-Lipschitz coefficients under Lyapunov conditions

Given global Lipschitz continuity and differentiability of high enough order on the coefficients in It\^{o}'s equation, differentiability of associated semigroups, existence of twice differentiable solutions to Kolmogorov equations and weak convergence rates of numerical approximations are known results. In this work and against the counterexamples of Hairer et al.(2015), the drift and diffusion coefficients having Lipschitz constants that are $o(\log V)$ and $o(\sqrt{\log V})$ respectively for a function $V$ satisfying $(\partial_t + L)V\leq CV$ is shown to be a generalizing condition in place of global Lipschitz continuity for the above.

math.PR

Optimal friction matrix for underdamped Langevin sampling

A systematic procedure for optimising the friction coefficient in underdamped Langevin dynamics as a sampling tool is given by taking the gradient of the associated asymptotic variance with respect to friction. We give an expression for this gradient in terms of the solution to an appropriate Poisson equation and show that it can be approximated by short simulations of the associated first variation/tangent process under concavity assumptions on the log density. Our algorithm is applied to the estimation of posterior means in Bayesian inference problems and reduced variance is demonstrated when compared to the original underdamped and overdamped Langevin dynamics in both full and stochastic gradient cases.

stat.CO

On the Generalised Langevin Equation for Simulated Annealing

In this paper, we consider the generalised (higher order) Langevin equation for the purpose of simulated annealing and optimisation of nonconvex functions. Our approach modifies the underdamped Langevin equation by replacing the Brownian noise with an appropriate Ornstein-Uhlenbeck process to account for memory in the system. Under reasonable conditions on the loss function and the annealing schedule, we establish convergence of the continuous time dynamics to a global minimum. In addition, we investigate the performance numerically and show better performance and higher exploration of the state space compared to the underdamped and overdamped Langevin dynamics with the same annealing schedule.

math.PR