SearcharxivSearch

arXiv subjects

Henrik Hult

Publications and source records attributed to Henrik Hult.

At least 19 recordsLinked to original sources

A dual process for the coupled Wright-Fisher diffusion

The coupled Wright-Fisher diffusion is a multi-dimensional Wright-Fisher diffusion for multi-locus and multi-allelic genetic frequencies, expressed as the strong solution to a system of stochastic differential equations that are coupled in the drift, where the pairwise interaction among loci is modelled by an inter-locus selection. In this paper, an ancestral process, which is dual to the coupled Wright-Fisher diffusion, is derived. The dual process corresponds to the block counting process of coupled ancestral selection graphs, one for each locus. Jumps of the dual process arise from coalescence, mutation, single-branching, which occur at one locus at the time, and double-branching, which occur simultaneously at two loci. The coalescence and mutation rates have the typical structure of the transition rates of the Kingman coalescent process. The single-branching rate not only contains the one-locus selection parameters in a form that generalises the rates of an ancestral selection graph, but it also contains the two-locus selection parameters to include the effect of the pairwise interaction on the single loci. The double-branching rate reflects the particular structure of pairwise selection interactions of the coupled Wright-Fisher diffusion. Moreover, in the special case of two loci, two alleles, with selection and parent independent mutation, the stationary density for the coupled Wright-Fisher diffusion and the transition rates of the dual process are obtained in an explicit form.

math.PR

Weak convergence of the scaled jump chain and number of mutations of the Kingman coalescent

The Kingman coalescent is a fundamental process in population genetics modelling the ancestry of a sample of individuals backwards in time. In this paper, in a large-sample-size regime, we study asymptotic properties of the coalescent under neutrality and a general finite-alleles mutation scheme, i.e. including both parent independent and parent dependent mutation. In particular, we consider a sequence of Markov chains that is related to the coalescent and consists of block-counting and mutation-counting components. We show that these components, suitably scaled, converge weakly to deterministic components and Poisson processes with varying intensities, respectively. Along the way, we develop a novel approach, based on a change of measure, to generalise the convergence result from the parent independent to the parent dependent mutation setting, in which several crucial quantities are not known explicitly.

math.PR

A weak convergence approach to large deviations for stochastic approximations

The theory of stochastic approximations form the theoretical foundation for studying convergence properties of many popular recursive learning algorithms in statistics, machine learning and statistical physics. Large deviations for stochastic approximations provide asymptotic estimates of the probability that the learning algorithm deviates from its expected path, given by a limit ODE, and the large deviation rate function gives insights to the most likely way that such deviations occur. In this paper we prove a large deviation principle for general stochastic approximations with state-dependent Markovian noise and decreasing step size. Using the weak convergence approach to large deviations, we generalize previous results for stochastic approximations and identify the appropriate scaling sequence for the large deviation principle. We also give a new representation for the rate function, in which the rate function is expressed as an action functional involving the family of Markov transition kernels. Examples of learning algorithms that are covered by the large deviation principle include stochastic gradient descent, persistent contrastive divergence and the Wang-Landau algorithm.

math.PR

Marginal Thresholding in Noisy Image Segmentation

This work presents a study on label noise in medical image segmentation by considering a noise model based on Gaussian field deformations. Such noise is of interest because it yields realistic looking segmentations and because it is unbiased in the sense that the expected deformation is the identity mapping. Efficient methods for sampling and closed form solutions for the marginal probabilities are provided. Moreover, theoretically optimal solutions to the loss functions cross-entropy and soft-Dice are studied and it is shown how they diverge as the level of noise increases. Based on recent work on loss function characterization, it is shown that optimal solutions to soft-Dice can be recovered by thresholding solutions to cross-entropy with a particular a priori unknown threshold that efficiently can be computed. This raises the question whether the decrease in performance seen when using cross-entropy as compared to soft-Dice is caused by using the wrong threshold. The hypothesis is validated in 5-fold studies on three organ segmentation problems from the TotalSegmentor data set, using 4 different strengths of noise. The results show that changing the threshold leads the performance of cross-entropy to go from systematically worse than soft-Dice to similar or better results than soft-Dice.

cs.CV

Noisy Image Segmentation With Soft-Dice

This paper presents a study on the soft-Dice loss, one of the most popular loss functions in medical image segmentation, for situations where noise is present in target labels. In particular, the set of optimal solutions are characterized and sharp bounds on the volume bias of these solutions are provided. It is further shown that a sequence of soft segmentations converging to optimal soft-Dice also converges to optimal Dice when converted to hard segmentations using thresholding. This is an important result because soft-Dice is often used as a proxy for maximizing the Dice metric. Finally, experiments confirming the theoretical results are provided.

cs.CV

On Image Segmentation With Noisy Labels: Characterization and Volume Properties of the Optimal Solutions to Accuracy and Dice

We study two of the most popular performance metrics in medical image segmentation, Accuracy and Dice, when the target labels are noisy. For both metrics, several statements related to characterization and volume properties of the set of optimal segmentations are proved, and associated experiments are provided. Our main insights are: (i) the volume of the solutions to both metrics may deviate significantly from the expected volume of the target, (ii) the volume of a solution to Accuracy is always less than or equal to the volume of a solution to Dice and (iii) the optimal solutions to both of these metrics coincide when the set of feasible segmentations is constrained to the set of segmentations with the volume equal to the expected volume of the target.

cs.CV

Asymptotic behaviour of sampling and transition probabilities in coalescent models under selection and parent dependent mutations

The results in this paper provide new information on asymptotic properties of classical models: the neutral Kingman coalescent under a general finite-alleles, parent-dependent mutation mechanism, and its generalisation, the ancestral selection graph. Several relevant quantities related to these fundamental models are not explicitly known when mutations are parent dependent. Examples include the probability that a sample taken from a population has a certain type configuration, and the transition probabilities of their block counting jump chains. In this paper, asymptotic results are derived for these quantities, as the sample size goes to infinity. It is shown that the sampling probabilities decay polynomially in the sample size with multiplying constant depending on the stationary density of the Wright-Fisher diffusion and that the transition probabilities converge to the limit of frequencies of types in the sample.

math.PR

Importance sampling for a simple Markovian intensity model using subsolutions

This paper considers importance sampling for estimation of rare-event probabilities in a specific collection of Markovian jump processes used for e.g. modelling of credit risk. Previous attempts at designing importance sampling algorithms have resulted in poor performance and the main contribution of the paper is the design of efficient importance sampling algorithms using subsolutions. The dynamics of the jump processes causes the corresponding Hamilton-Jacobi equations to have an intricate state-dependence, which makes the design of efficient algorithms difficult. We provide theoretical results that quantify the performance of importance sampling algorithms in general and construct asymptotically optimal algorithms for some examples. The computational gain compared to standard Monte Carlo is illustrated by numerical examples.

math.PR

Almost sure convergence of the accelerated weight histogram algorithm

The accelerated weight histogram (AWH) algorithm is an iterative extended ensemble algorithm, developed for statistical physics and computational biology applications. It is used to estimate free energy differences and expectations with respect to Gibbs measures. The AWH algorithm is based on iterative updates of a design parameter, which is closely related to the free energy, obtained by matching a weight histogram with a specified target distribution. The weight histogram is constructed from samples of a Markov chain on the product of the state space and parameter space. In this paper almost sure convergence of the AWH algorithm is proved, for estimating free energy differences as well as estimating expectations with adaptive ergodic averages. The proof is based on identifying the AWH algorithm as a stochastic approximation and studying the properties of the associated limit ordinary differential equation.

math.PR

Variational Auto Encoder Gradient Clustering

Clustering using deep neural network models have been extensively studied in recent years. Among the most popular frameworks are the VAE and GAN frameworks, which learns latent feature representations of data through encoder / decoder neural net structures. This is a suitable base for clustering tasks, as the latent space often seems to effectively capture the inherent essence of data, simplifying its manifold and reducing noise. In this article, the VAE framework is used to investigate how probability function gradient ascent over data points can be used to process data in order to achieve better clustering. Improvements in classification is observed comparing with unprocessed data, although state of the art results are not obtained. Processing data with gradient descent however results in more distinct cluster separation, making it simpler to investigate suitable hyper parameter settings such as the number of clusters. We propose a simple yet effective method for investigating suitable number of clusters for data, based on the DBSCAN clustering algorithm, and demonstrate that cluster number determination is facilitated with gradient processing. As an additional curiosity, we find that our baseline model used for comparison; a GMM on a t-SNE latent space for a VAE structure with weight one on reconstruction during training (autoencoder), yield state of the art results on the MNIST data, to our knowledge not beaten by any other existing model.

cs.LG

Particle Filter Bridge Interpolation

Auto encoding models have been extensively studied in recent years. They provide an efficient framework for sample generation, as well as for analysing feature learning. Furthermore, they are efficient in performing interpolations between data-points in semantically meaningful ways. In this paper, we build further on a previously introduced method for generating canonical, dimension independent, stochastic interpolations. Here, the distribution of interpolation paths is represented as the distribution of a bridge process constructed from an artificial random data generating process in the latent space, having the prior distribution as its invariant distribution. As a result the stochastic interpolation paths tend to reside in regions of the latent space where the prior has high mass. This is a desirable feature since, generally, such areas produce semantically meaningful samples. In this paper, we extend the bridge process method by introducing a discriminator network that accurately identifies areas of high latent representation density. The discriminator network is incorporated as a change of measure of the underlying bridge process and sampling of interpolation paths is implemented using sequential Monte Carlo. The resulting sampling procedure allows for greater variability in interpolation paths and stronger drift towards areas of high data density.

stat.ML

Exact simulation of coupled Wright-Fisher diffusions

In this paper an exact rejection algorithm for simulating paths of the coupled Wright-Fisher diffusion is introduced. The coupled Wright-Fisher diffusion is a family of multidimensional Wright-Fisher diffusions that have drifts depending on each other through a coupling term and that find applications in the study of interacting genes' networks as those encountered in studies of antibiotic resistance. Our algorithm uses independent neutral Wright-Fisher diffusions as candidate proposals, which can be sampled exactly by means of existing algorithms and are only needed at a finite number of points. Once a candidate is accepted, the remaining of the path can be recovered by sampling from a neutral multivariate Wright-Fisher bridge, for which we also provide an exact sampling strategy. The technique relies on a modification of the alternating series method and extends existing algorithms that are currently available for the one-dimensional case. Finally, the algorithm's complexity is derived and its performance demonstrated in a simulation study.

math.PR

Estimates of the proportion of SARS-CoV-2 infected individuals in Sweden

In this paper a Bayesian SEIR model is studied to estimate the proportion of the population infected with SARS-CoV-2, the virus responsible for COVID-19. To capture heterogeneity in the population and the effect of interventions to reduce the rate of epidemic spread, the model uses a time-varying contact rate, whose logarithm has a Gaussian process prior. A Poisson point process is used to model the occurrence of deaths due to COVID-19 and the model is calibrated using data of daily death counts in combination with a snapshot of the the proportion of individuals with an active infection, performed in Stockholm in late March. The methodology is applied to regions in Sweden. The results show that the estimated proportion of the population who has been infected is around 13.5% in Stockholm, by 2020-05-15, and ranges between 2.5% - 15.6% in the other investigated regions. In Stockholm where the peak of daily death counts is likely behind us, parameter uncertainty does not heavily influence the expected daily number of deaths, nor the expected cumulative number of deaths. It does, however, impact the estimated cumulative number of infected individuals. In the other regions, where random sampling of the number of active infections is not available, parameter sharing is used to improve estimates, but the parameter uncertainty remains substantial.

q-bio.PE

Exact and efficient simulation of tail probabilities of heavy-tailed infinite series

We develop an efficient simulation algorithm for computing the tail probabilities of the infinite series $S = \sum_{n \geq 1} a_n X_n$ when random variables $X_n$ are heavy-tailed. As $S$ is the sum of infinitely many random variables, any simulation algorithm that stops after simulating only fixed, finitely many random variables is likely to introduce a bias. We overcome this challenge by rewriting the tail probability of interest as a sum of a random number of telescoping terms, and subsequently developing conditional Monte Carlo based low variance simulation estimators for each telescoping term. The resulting algorithm is proved to result in estimators that a) have no bias, and b) require only a fixed, finite number of replications irrespective of how rare the tail probability of interest is. Thus, by combining a traditional variance reduction technique such as conditional Monte Carlo with more recent use of auxiliary randomization to remove bias in a multi-level type representation, we develop an efficient and unbiased simulation algorithm for tail probabilities of $S$. These have many applications including in analysis of financial time-series and stochastic recurrence equations arising in models in actuarial risk and population biology.

math.PR

Large deviations for weighted empirical measures arising in importance sampling

Importance sampling is a popular method for efficient computation of various properties of a distribution such as probabilities, expectations, quantiles etc. The output of an importance sampling algorithm can be represented as a weighted empirical measure, where the weights are given by the likelihood ratio between the original distribution and the sampling distribution. In this paper the efficiency of an importance sampling algorithm is studied by means of large deviations for the weighted empirical measure. The main result, which is stated as a Laplace principle for the weighted empirical measure arising in importance sampling, can be viewed as a weighted version of Sanov's theorem. The main theorem is applied to quantify the performance of an importance sampling algorithm over a collection of subsets of a given target set as well as quantile estimates. The analysis yields an estimate of the sample size needed to reach a desired precision as well as of the reduction in cost for importance sampling compared to standard Monte Carlo.

math.PR

Min-max representations of viscosity solutions of Hamilton-Jacobi equations and applications in rare-event simulation

In this paper a duality relation between the Mañé potential and Mather's action functional is derived in the context of convex and state-dependent Hamiltonians. The duality relation is used to obtain min-max representations of viscosity solutions of first order Hamilton-Jacobi equations. These min-max representations naturally suggest class\-es of subsolutions of Hamilton-Jacobi equations that arise in the theory of large deviations. The subsolutions, in turn, are good candidates for designing efficient rare-event simulation algorithms.

math.AP

A simple time-consistent model for the forward density process

In this paper a simple model for the evolution of the forward density of the future value of an asset is proposed. The model allows for a straightforward initial calibration to option prices and has dynamics that are consistent with empirical findings from option price data. The model is constructed with the aim of being both simple and realistic, and avoid the need for frequent re-calibration. The model prices of $n$ options and a forward contract are expressed as time-varying functions of an $(n+1)$-dimensional Brownian motion and it is investigated how the Brownian trajectory can be determined from the trajectories of the price processes. An approach based on particle filtering is presented for determining the location of the driving Brownian motion from option prices observed in discrete time. A simulation study and an empirical study of call options on the S&P 500 index illustrates that the model provides a good fit to option price data.

q-fin.PR

Markov chain Monte Carlo for computing rare-event probabilities for a heavy-tailed random walk

In this paper a method based on a Markov chain Monte Carlo (MCMC) algorithm is proposed to compute the probability of a rare event. The conditional distribution of the underlying process given that the rare event occurs has the probability of the rare event as its normalizing constant. Using the MCMC methodology a Markov chain is simulated, with that conditional distribution as its invariant distribution, and information about the normalizing constant is extracted from its trajectory. The algorithm is described in full generality and applied to the problem of computing the probability that a heavy-tailed random walk exceeds a high threshold. An unbiased estimator of the reciprocal probability is constructed whose normalized variance vanishes asymptotically. The algorithm is extended to random sums and its performance is illustrated numerically and compared to existing importance sampling algorithms.

math.PR