SearcharxivSearch

arXiv subjects

Anthony Lee

Publications and source records attributed to Anthony Lee.

At least 19 recordsLinked to original sources

Resampling in conditional SMC algorithms

Conditional sequential Monte Carlo (CSMC) algorithms arise in particle Markov chain Monte Carlo and a number of related settings. As in standard sequential Monte Carlo (SMC) algorithms, it is possible to employ a number of approaches to resampling within CSMC, but some additional care is required to arrive at a valid algorithm. We present a simple framework for implementing valid SMC and CSMC algorithms which (a) covers most known resampling schemes including those with a complicated dependence structure like systematic resampling, but also adaptive resampling, and even more `exotic' schemes like a version of chopthin resampling; (b) explains how to implement conditional analogues of these and other well known resampling schemes without randomly permuting/shifting the order of the ancestor indices; (c) requires only very weak assumptions which include neither (marginal) "unbiasedness" nor exchangeability.

stat.CO

Explicit integral representations and quantitative bounds for two-layer ReLU networks

An approach to construct explicit integral representations for two-layer ReLU networks is presented, which provides relatively simple representations for any multivariate polynomial. Quantitative bounds are provided for a particular, sharpened ReLU integral representation, which involves a harmonic extension and a projection. The bounds demonstrate that functions can be approximated with $L^{2}(\mathcal{D})$ errors that do not depend explicitly on dimension or degree, but rather the coefficients of their monomial expansions and the distribution $\mathcal{D}$. We also present a connection to the RKHS of the exponential kernel $K(x,y)=\exp\left(\left\langle x,y\right\rangle \right)$, and a very simple integral representation involving additionally multiplication via a fixed function which has better quantitative bounds.

stat.ML

Robust, partially alive particle Metropolis-Hastings via the Frankenfilter

When a hidden Markov model permits the conditional likelihood of an observation given the hidden process to be zero, all particle simulations from one observation time to the next could produce zeros. If so, the filtering distribution cannot be estimated and the estimated parameter likelihood is zero. The alive particle filter addresses this by simulating a random number of particles for each inter-observation interval, stopping after a target number of non-zero conditional likelihoods. For outlying observations or poor parameter values, a non-zero result can be extremely unlikely, and computational costs prohibitive. We introduce the Frankenfilter, a principled, partially alive particle filter that targets a user-defined amount of success whilst fixing lower and upper bounds on the number of simulations. The Frankenfilter produces unbiased estimators of the likelihood, suitable for pseudo-marginal Metropolis--Hastings (PMMH). We demonstrate that PMMH with the Frankenfilter is more robust to outliers and mis-specified initial parameter values than PMMH using standard particle filters, and is typically at least 2-3 times more efficient. We also provide advice for choosing the amount of success. In the case of n exact observations, this is particularly simple: target n successes.

stat.ME

Maximal Fuchsian subgroups of the $d=2$ Bianchi group

Let $\Gamma$ denote the $d = 2$ Bianchi group $\operatorname{PSL}(2,\mathbb{Z}[\sqrt{-2}])$. We give an explicit description of all conjugacy classes of maximal nonelementary Fuchsian subgroups of $\Gamma$ as integral orders of certain indefinite quaternion algebras over $\mathbb{Q}$. Using this description, we also provide the covolumes corresponding to each conjugacy class. As an application, we compute the limit $\lim_{x\to\infty} \frac{\Pi(x)}{x}$ where $\Pi(x)$ counts the number of primitive totally geodesic immersed surfaces in $\Gamma\backslash\mathbb{H}^3$ with areas less than $x$.

math.NT

Knots and variance ordering of sequential Monte Carlo algorithms

Sequential Monte Carlo algorithms, or particle filters, are widely used for approximating intractable integrals, particularly those arising in Bayesian inference and state-space models. We introduce a new variance reduction technique, the knot operator, which improves the efficiency of particle filters by incorporating potential function information into part, or all, of a transition kernel. The knot operator induces a partial ordering of Feynman-Kac models that implies an order on the asymptotic variance of particle filters, offering a new approach to algorithm design. We discuss connections to existing strategies for designing efficient particle filters, including model marginalisation. Our theory generalises such techniques and provides quantitative asymptotic variance ordering results. We revisit the fully-adapted (auxiliary) particle filter using our theory of knots to show how a small modification guarantees an asymptotic variance ordering for all relevant test functions.

stat.CO

How to identify earth pressures on in-service tunnel linings: Insights from Bayesian inversion to address non-uniqueness

Identifying earth pressures on in-service transportation tunnel linings is essential for their health monitoring and performance prediction, particularly in structures that exhibit poor performance. Due to the high costs associated with pressure gauges, pressure inversion based on easily observed structural responses, such as deformations, is preferred. A significant challenge lies in the non-uniqueness of inversion results, where various pressures can yield similar structural responses. Existing approaches often overlook detailed discussions on this critical issue. In addressing this gap, this study introduces a Bayesian approach. The proposed statistical framework effectively quantifies the uncertainty induced by non-uniqueness. Further analysis identifies the uniform component in distributed pressures as the primary source of non-uniqueness. Insights into mitigation strategies are provided, including increasing the quantity of deformation data or incorporating an observation of internal normal force within the tunnel lining -- the latter proving to be notably more effective. A practical application in a numerical case study demonstrates the effectiveness of this approach. In addition, our investigation recommends maintaining deformation measurement accuracy within the range of [-1, 1] mm to ensure satisfactory outcomes. Finally, deficiencies and potential future extensions of this approach are discussed.

stat.AP

Adaptive tempered reversible jump algorithm for Bayesian curve fitting

Bayesian curve fitting plays an important role in inverse problems, and is often addressed using the Reversible Jump Markov Chain Monte Carlo (RJMCMC) algorithm. However, this algorithm can be computationally inefficient without appropriately tuned proposals. As a remedy, we present an adaptive RJMCMC algorithm for the curve fitting problems by extending the Adaptive Metropolis sampler from a fixed-dimensional to a trans-dimensional case. In this presented algorithm, both the size and orientation of the proposal function can be automatically adjusted in the sampling process. Specifically, the curve fitting setting allows for the approximation of the posterior covariance of the a priori unknown function on a representative grid of points. This approximation facilitates the definition of efficient proposals. In addition, we introduce an auxiliary-tempered version of this algorithm via non-reversible parallel tempering. To evaluate the algorithms, we conduct numerical tests involving a series of controlled experiments. The results demonstrate that the adaptive algorithms exhibit significantly higher efficiency compared to the conventional ones. Even in cases where the posterior distribution is highly complex, leading to ineffective convergence in the auxiliary-tempered conventional RJMCMC, the proposed auxiliary-tempered adaptive RJMCMC performs satisfactorily. Furthermore, we present a realistic inverse example to test the algorithms. The successful application of the adaptive algorithm distinguishes it again from the conventional one that fails to converge effectively even after millions of iterations.

stat.CO

Mixing time of the conditional backward sampling particle filter

The conditional backward sampling particle filter (CBPF) is a powerful Markov chain Monte Carlo sampler for general state space hidden Markov model (HMM) smoothing. It was proposed as an improvement over the conditional particle filter (CPF), which has an $O(T^2)$ complexity under a general `strong' mixing assumption, where $T$ is the time horizon. Empirical evidence of the superiority of the CBPF over the CPF has never been theoretically quantified. We show that the CBPF has $O(T \log T)$ time complexity under strong mixing: its mixing time is upper bounded by $O(\log T)$, for any sufficiently large number of particles $N$ independent of $T$. This $O(\log T)$ mixing time is optimal. To prove our main result, we introduce a novel coupling of two CBPFs, which employs a maximal coupling of two particle systems at each time instant. The coupling is implementable and we use it to construct unbiased, finite variance, estimates of functionals which have arbitrary dependence on the latent state's path, with a total expected cost of $O(T \log T)$. We use this to construct unbiased estimates of the HMM's score function, and also investigate other couplings which can exhibit improved behaviour. We demonstrate our methods on financial and calcium imaging applications.

stat.CO

Weak Poincar\'e Inequalities for Markov chains: theory and applications

We investigate the application of Weak Poincar\'e Inequalities (WPI) to Markov chains to study their rates of convergence and to derive complexity bounds. At a theoretical level we investigate the necessity of the existence of WPIs to ensure \mathrm{L}^{2}-convergence, in particular by establishing equivalence with the Resolvent Uniform Positivity-Improving (RUPI) condition and providing a counterexample. From a more practical perspective, we extend the celebrated Cheeger's inequalities to the subgeometric setting, and further apply these techniques to study random-walk Metropolis algorithms for heavy-tailed target distributions and to obtain lower bounds on pseudo-marginal algorithms.

math.PR

On the Forgetting of Particle Filters

We study the forgetting properties of the particle filter when its state - the collection of particles - is regarded as a Markov chain. Under a strong mixing assumption on the particle filter's underlying Feynman-Kac model, we find that the particle filter is exponentially mixing, and forgets its initial state in $O(\log N )$ 'time', where $N$ is the number of particles and time refers to the number of particle filter algorithm steps, each comprising a selection (or resampling) and mutation (or prediction) operation. We present an example which shows that this rate is optimal. In contrast to our result, available results to-date are extremely conservative, suggesting $O(\alpha^N)$ time steps are needed, for some $\alpha>1$, for the particle filter to forget its initialisation. We also study the conditional particle filter (CPF) and extend our forgetting result to this context. We establish a similar conclusion, namely, CPF is exponentially mixing and forgets its initial state in $O(\log N )$ time. To support this analysis, we establish new time-uniform $L^p$ error estimates for CPF, which can be of independent interest. We also establish new propagation of chaos type results using our proof techniques, discuss implications to couplings of particle filters and an application to processing out-of-sequence measurements.

math.PR

Explicit convergence bounds for Metropolis Markov chains: isoperimetry, spectral gaps and profiles

We derive the first explicit bounds for the spectral gap of a random walk Metropolis algorithm on $R^d$ for any value of the proposal variance, which when scaled appropriately recovers the correct $d^{-1}$ dependence on dimension for suitably regular invariant distributions. We also obtain explicit bounds on the ${\rm L}^2$-mixing time for a broad class of models. In obtaining these results, we refine the use of isoperimetric profile inequalities to obtain conductance profile bounds, which also enable the derivation of explicit bounds in a much broader class of models. We also obtain similar results for the preconditioned Crank--Nicolson Markov chain, obtaining dimension-independent bounds under suitable assumptions.

math.PR

Generalizing the German Tank Problem

The German Tank Problem dates back to World War II when the Allies used a statistical approach to estimate the number of enemy tanks produced or on the field from observed serial numbers after battles. Assuming that the tanks are labeled consecutively starting from 1, if we observe $k$ tanks from a total of $N$ tanks with the maximum observed tank being $m$, then the best estimate for $N$ is $m(1 + 1/k) - 1$. We explore many generalizations. We looked at the discrete and continuous one dimensional case. We explored different estimators such as the $L$\textsuperscript{th} largest tank, and applied motivation from portfolio theory and studied a weighted average; however, the original formula was the best. We generalized the problem in two dimensions, with pairs instead of points, studying the discrete and continuous square and circle variants. There were complications from curvature issues and that not every number is representable as a sum of two squares. We often concentrated on the large $N$ limit. For the discrete and continuous square, we tested various statistics, finding the largest observed component did best; the scaling factor for both cases is $(2k+1)/2k$. The discrete case was especially involved because we had to use approximation formulas that gave us the number of lattice points inside the circle. Interestingly, the scaling factors were different for the cases. Lastly, we generalized the problem into $L$ dimensional squares and circles. The discrete and continuous square proved similar to the two dimensional square problem. However, for the $L$\textsuperscript{th} dimensional circle, we had to use formulas for the volume of the $L$-ball, and had to approximate the number of lattice points inside it. The formulas for the discrete circle were particularly interesting, as there was no $L$ dependence in the formula.

math.PR

Monte Carlo twisting for particle filters

We consider the problem of designing efficient particle filters for twisted Feynman--Kac models. Particle filters using twisted models can deliver low error approximations of statistical quantities and such twisting functions can be learnt iteratively. Practical implementations of these algorithms are complicated by the need to (i) sample from the twisted transition dynamics, and (ii) calculate the twisted potential functions. We expand the class of applicable models using rejection sampling for (i) and unbiased approximations for (ii) using a random weight particle filter. We characterise the average acceptance rates within the particle filter in order to control the computational cost, and analyse the asymptotic variance. Empirical results show the mean squared error of the normalising constant estimate in our method is smaller than a memory-equivalent particle filter but not a computation-equivalent filter. Both comparisons are improved when more efficient sampling is possible which we demonstrate on a stochastic volatility model.

stat.ME

Volume-Independent Music Matching by Frequency Spectrum Comparison

Often, I hear a piece of music and wonder what the name of the piece is. Indeed, there are applications such as Shazam app that provides music matching. However, the limitations of those apps are that the same piece performed by the same musician cannot be identified if it is not the same recording. Shazam identifies the recording of it, not the music. This is because Shazam matches the variation in volume, not the frequencies of the sound. This research attempts to match music the way humans understand it: by the frequency spectrum of music, not the volume variation. Essentially, the idea is to precompute the frequency spectrums of all the music in the database, then take the unknown piece and try to match its frequency spectrum against every segment of every music in the database. I did it by matching the frequency spectrum of the unknown piece to our database by sliding the window by 0.1 seconds and calculating the error by taking Absolute value, normalizing the audio, subtracting the normalized arrays, and taking the sum of absolute differences. The segment that shows the least error is considered the candidate for the match. The matching performance proved to be dependent on the complexity of the music. Matching simple music, such as single note pieces, was successful. However, more complex pieces, such as Chopins Ballade 4, were not successful, that is, the algorithm could not produce low error values in any of the music in the database. I suspect that it has to do with having too many notes: mismatches in the higher harmonics added up to a significant amount of errors, which swamps the calculations.

cs.SD

Solving the Poisson equation using coupled Markov chains

This article shows how coupled Markov chains that meet exactly after a random number of iterations can be used to generate unbiased estimators of the solutions of the Poisson equation. Through this connection, we re-derive known unbiased estimators of expectations with respect to the stationary distribution of a Markov chain and provide conditions for the finiteness of their moments. We further construct unbiased estimators of the asymptotic variance of Markov chain ergodic averages, and provide conditions for the finiteness of the estimators' moments of any order. If their second moment is finite, the average of independent copies of such estimators converges to the asymptotic variance at the Monte Carlo rate, comparing favorably to known rates for batch means and spectral variance estimators. The results are illustrated with numerical experiments.

stat.CO

Comparison of Markov chains via weak Poincar\'e inequalities with application to pseudo-marginal MCMC

We investigate the use of a certain class of functional inequalities known as weak Poincar\'e inequalities to bound convergence of Markov chains to equilibrium. We show that this enables the straightforward and transparent derivation of subgeometric convergence bounds for methods such as the Independent Metropolis--Hastings sampler and pseudo-marginal methods for intractable likelihoods, the latter being subgeometric in many practical settings. These results rely on novel quantitative comparison theorems between Markov chains. Associated proofs are simpler than those relying on drift/minorization conditions and the tools developed allow us to recover and further extend known results as particular cases. We are then able to provide new insights into the practical use of pseudo-marginal algorithms, analyse the effect of averaging in Approximate Bayesian Computation (ABC) and the use of products of independent averages, and also to study the case of lognormal weights relevant to particle marginal Metropolis--Hastings (PMMH).

stat.CO

A general perspective on the Metropolis-Hastings kernel

Since its inception the Metropolis-Hastings kernel has been applied in sophisticated ways to address ever more challenging and diverse sampling problems. Its success stems from the flexibility brought by the fact that its verification and sampling implementation rests on a local ``detailed balance'' condition, as opposed to a global condition in the form of a typically intractable integral equation. While checking the local condition is routine in the simplest scenarios, this proves much more difficult for complicated applications involving auxiliary structures and variables. Our aim is to develop a framework making establishing correctness of complex Markov chain Monte Carlo kernels a purely mechanical or algebraic exercise, while making communication of ideas simpler and unambiguous by allowing a stronger focus on essential features -- a choice of embedding distribution, an involution and occasionally an acceptance function -- rather than the induced, boilerplate structure of the kernels that often tends to obscure what is important. This framework can also be used to validate kernels that do not satisfy detailed balance, i.e. which are not reversible, but a modified version thereof.

stat.CO