Searcharxiv⌕ Search

arXiv subjects

Dario Spanò

Publications and source records attributed to Dario Spanò.

At least 19 recordsLinked to original sources

A debiased Bernoulli factory and unbiased estimation of a probability

Given a known function $f : [0, 1] \to (0, 1)$ and a random but almost surely finite number of independent, Ber$(x)$-distributed random variables with unknown $x \in [0, 1]$, we prove the existence of an unbiased, $[0, 1]$-valued estimator of the probability $f(x) \in (0, 1)$. Our estimator is based on so-called debiasing, or randomly truncating a telescopic series of consistent estimators. Debiased estimators of a probability are not typically constrained to $[0, 1]$, or even bounded, even when all consistent estimators used as inputs are. We show that constructing the series of consistent estimators from the coefficients of a particular Bernoulli factory yields provable boundedness provided $f \in C^ρ[0, 1]$ for $ρ> 3$. Our result can be thought of as a novel Bernoulli factory with the appealing property that the required number of Ber$(x)$-distributed random variates is independent of their outcomes.

math.PR↗

Genealogical processes of sequential Monte Carlo methods and other non-neutral population models under rapid mutation

We show that genealogical trees arising from a broad class of non-neutral models of population evolution converge to the Kingman coalescent under a suitable rescaling of time. As well as non-neutral biological evolution, our results apply to genetic algorithms encompassing the prominent class of sequential Monte Carlo (SMC) methods. The time rescaling we need differs slightly from that used in classical results for convergence to the Kingman coalescent, which has implications for the performance of different resampling schemes in SMC algorithms. In addition, our work substantially simplifies earlier proofs of convergence to the Kingman coalescent, and corrects an error common to several earlier results.

math.PR↗

Splitting schemes and estimators for stochastic differential equations with Hölder multiplicative noise

We study parameter estimation for univariate stochastic differential equations with locally Lipschitz drift and Hölder continuous multiplicative diffusion, a class commonly arising in several applications. Existing inference methods typically rely on either the Euler-Maruyama discretisation, despite its lack of strong convergence and failure to preserve the state space, or on approximations, e.g. Gaussian approximation or truncation of Hermite's expansions, impacting on their stability and computational efficiency. We introduce the first explicit pseudo-likelihood estimators based on numerical splitting schemes that are both strong mean-square convergent and state space preserving for this class of SDEs. Our approach is based on a novel decomposition of the SDE that exploits reducibility and the Lamperti transform, leading to Lie-Trotter (LT) and Strang splitting schemes yielding explicit pseudo-likelihoods and maximum likelihood estimators based on them. We prove strong mean-square convergence, state space preservation, and improved robustness with respect to the discretisation step compared to Euler-Maruyama-based methods. We further establish consistency and asymptotic normality of the LT estimator. Because the proposed numerical scheme couples drift and diffusion parameters in the pseudo-likelihood, the asymptotic analysis requires new proof techniques. Extensive simulations demonstrate that the proposed estimators outperform existing methods in both accuracy and computational efficiency.

stat.ME↗

Exact inference via quasi-conjugacy in two-parameter Poisson-Dirichlet hidden Markov models

We introduce a nonparametric model for inferring time-evolving, unobserved probability distributions from discrete-time data consisting of unlabelled partitions. The latent process is a two-parameter Poisson-Dirichlet diffusion, and observations arise via exchangeable sampling. Applications include social and genetic data where only aggregate clustering summaries are observed. To address the intractable likelihood, we develop a tractable inferential framework that avoids label enumeration and direct simulation of the latent state. We exploit a duality between the diffusion and a pure-death process on partitions, together with coagulation operators that encode the effect of new data. These yield closed-form, recursive updates for forward and backward inference. We compute exact posterior distributions of the latent state at arbitrary times and predictive distributions of future or interpolated partitions. This enables online and offline inference and forecasting with full uncertainty quantification, bypassing MCMC and sequential Monte Carlo. Compared to particle filtering, our method achieves higher accuracy, lower variance, and substantial computational gains. We illustrate the methodology with synthetic experiments and a social network application, recovering interpretable patterns in time-varying heterozygosity.

stat.ME↗

Subordinated Wright-Fisher Priors

A new class of time-dependent Dirichlet priors is introduced as a generalisation of the Wright-Fisher diffusion, allowing discontinuities in the trajectories, as well as non-Markovian memory. This class is obtained as a simple stochastic time-change (subordination), interpreted as a hyper-prior assigned to the operational time-clock of a Wright-Fisher diffusion. Explicit representations and exact sampling algorithms are obtained for prior and posterior distributions of the process and of its clock, given partially exchangeable data sampled at discrete time-points. Computability and conjugacy rely on a novel class of discrete dual processes, generalising existing results on duality and computable filters.

math.ST↗

The effective strength of selection in random environment

We analyse a family of two-types Wright-Fisher models with selection in a random environment and skewed offspring distribution. We provide a calculable criterion to quantify the impact of different shapes of selection on the fate of the weakest allele, and thus compare them. The main mathematical tool is duality, which we prove to hold, also in presence of random environment (quenched and in some cases annealed), between the population's allele frequencies and genealogy, both in the case of finite population size and in the scaling limit for large size. Duality also yields new insight on properties of branching-coalescing processes in random environment, such as their long term behaviour.

math.PR↗

Dual process in the two-parameter Poisson-Dirichlet diffusion

The two-parameter Poisson-Dirichlet diffusion takes values in the infinite ordered simplex and extends the celebrated infinitely-many-neutral-alleles model, having a two-parameter Poisson-Dirichlet stationary distribution. Here we identify a dual process for this diffusion and obtain its transition probabilities. The dual is shown to be given by Kingman's coalescent with mutation, conditional on a given configuration of leaves. Interestingly, the dual depends on the additional parameter of the stationary distribution only through the test functions and not through the transition rates. After discussing the sampling probabilities of a two-parameter Poisson-Dirichlet partition drawn conditionally on another partition, we use these notions together with the dual process to derive the transition density of the diffusion. Our derivation provides a new probabilistic proof of this result, leveraging on an extension of Pitman's Polya urn scheme, whereby the urn is split after a finite number of steps and two urns are run independently onwards. The proof strategy exemplifies the power of duality and could be exported to other models where a dual is available.

math.PR↗

Bernoulli factories and duality in Wright-Fisher and Allen-Cahn models of population genetics

Mathematical models of genetic evolution often come in pairs, connected by a so-called duality relation. The most seminal example are the Wright-Fisher diffusion and the Kingman coalescent, where the former describes the stochastic evolution of neutral allele frequencies in a large population forwards in time, and the latter describes the genetic ancestry of randomly sampled individuals from the population backwards in time. As well as providing a richer description than either model in isolation, duality often yields equations satisfied by quantities of interest. We employ the so-called Bernoulli factory - a celebrated tool in simulation-based computing - to derive duality relations for broad classes of genetics models. As concrete examples, we present Wright-Fisher diffusions with general drift functions, and Allen-Cahn equations with general, nonlinear forcing terms. The drift and forcing functions can be interpreted as the action of frequency-dependent selection. To our knowledge, this work is the first time a connection has been drawn between Bernoulli factories and duality in models of population genetics.

math.PR↗

EWF : simulating exact paths of the Wright--Fisher diffusion

The Wright--Fisher diffusion is important in population genetics in modelling the evolution of allele frequencies over time subject to the influence of biological phenomena such as selection, mutation, and genetic drift. Simulating paths of the process is challenging due to the form of the transition density. We present EWF, a robust and efficient sampler which returns exact draws for the diffusion and diffusion bridge processes, accounting for general models of selection including those with frequency-dependence. Given a configuration of selection, mutation, and endpoints, EWF returns draws at the requested sampling times from the law of the corresponding Wright--Fisher process. Output was validated by comparison to approximations of the transition density via the Kolmogorov--Smirnov test and QQ plots. All software is available at https://github.com/JaroSant/EWF

q-bio.PE↗

Diffusion Limits at Small Times for Coalescent Processes with Mutation and Selection

The Ancestral Selection Graph (ASG) is an important genealogical process which extends the well-known Kingman coalescent to incorporate natural selection. We show that the number of lineages of the ASG with and without mutation is asymptotic to $2/t$ as $t\to 0$, in agreement with the limiting behaviour of the Kingman coalescent. We couple these processes on the same probability space using a Poisson random measure construction that allows us to precisely compare their hitting times. These comparisons enable us to characterise the speed of coming down from infinity of the ASG as well as its fluctuations in a functional central limit theorem. This extends similar results for the Kingman coalescent.

math.PR↗

Wright-Fisher diffusion bridges

{\bf Abstract} The trajectory of the frequency of an allele which begins at $x$ at time $0$ and is known to have frequency $z$ at time $T$ can be modelled by the bridge process of the Wright-Fisher diffusion. Bridges when $x=z=0$ are particularly interesting because they model the trajectory of the frequency of an allele which appears at a time, then is lost by random drift or mutation after a time $T$. The coalescent genealogy back in time of a population in a neutral Wright-Fisher diffusion process is well understood. In this paper we obtain a new interpretation of the coalescent genealogy of the population in a bridge from a time $t\in (0,T)$. In a bridge with allele frequencies of 0 at times 0 and $T$ the coalescence structure is that the population coalesces in two directions from $t$ to $0$ and $t$ to $T$ such that there is just one lineage of the allele under consideration at times $0$ and $T$. The genealogy in Wright-Fisher diffusion bridges with selection is more complex than in the neutral model, but still with the property of the population branching and coalescing in two directions from time $t\in (0,T)$. The density of the frequency of an allele at time $t$ is expressed in a way that shows coalescence in the two directions. A new algorithm for exact simulation of a neutral Wright-Fisher bridge is derived. This follows from knowing the density of the frequency in a bridge and exact simulation from the Wright-Fisher diffusion. The genealogy of the neutral Wright-Fisher bridge is also modelled by branching Pólya urns, extending a representation in a Wright-Fisher diffusion. This is a new very interesting representation that relates Wright-Fisher bridges to classical urn models in a Bayesian setting.

math.PR↗

Duality and Fixation in $Ξ$-Wright-Fisher processes with frequency-dependent selection

A two-types, discrete-time population model with finite, constant size is constructed, allowing for a general form of frequency-dependent selection and skewed offspring distribution. Selection is defined based on the idea that individuals first choose a (random) number of $\textit{potential}$ parents from the previous generation and then, from the selected pool, they inherit the type of the fittest parent. The probability distribution function of the number of potential parents per individual thus parametrises entirely the selection mechanism. Using sampling- and moment-duality, weak convergence is then proved both for the allele frequency process of the selectively weak type and for the population's ancestral process. The scaling limits are, respectively, a two-types $Ξ$-Fleming-Viot jump-diffusion process with frequency-dependent selection, and a branching-coalescing process with general branching and simultaneous multiple collisions. Duality also leads to a characterisation of the probability of extinction of the selectively weak allele, in terms of the ancestral process' ergodic properties.

math.PR↗

Bayesian non-parametric inference for $Λ$-coalescents: consistency and a parametric method

We investigate Bayesian non-parametric inference of the $Λ$-measure of $Λ$-coalescent processes with recurrent mutation, parametrised by probability measures on the unit interval. We give verifiable criteria on the prior for posterior consistency when observations form a time series, and prove that any non-trivial prior is inconsistent when all observations are contemporaneous. We then show that the likelihood given a data set of size $n \in \mathbb{N}$ is constant across $Λ$-measures whose leading $n - 2$ moments agree, and focus on inferring truncated sequences of moments. We provide a large class of functionals which can be extremised using finite computation given a credible region of posterior truncated moment sequences, and a pseudo-marginal Metropolis-Hastings algorithm for sampling the posterior. Finally, we compare the efficiency of the exact and noisy pseudo-marginal algorithms with and without delayed acceptance acceleration using a simulation study.

stat.ME↗

Conjugacy properties of time-evolving Dirichlet and gamma random measures

We extend classic characterisations of posterior distributions under Dirichlet process and gamma random measures priors to a dynamic framework. We consider the problem of learning, from indirect observations, two families of time-dependent processes of interest in Bayesian nonparametrics: the first is a dependent Dirichlet process driven by a Fleming-Viot model, and the data are random samples from the process state at discrete times; the second is a collection of dependent gamma random measures driven by a Dawson-Watanabe model, and the data are collected according to a Poisson point process with intensity given by the process state at discrete times. Both driving processes are diffusions taking values in the space of discrete measures whose support varies with time, and are stationary and reversible with respect to Dirichlet and gamma priors respectively. A common methodology is developed to obtain in closed form the time-marginal posteriors given past and present data. These are shown to belong to classes of finite mixtures of Dirichlet processes and gamma random measures for the two models respectively, yielding conjugacy of these classes to the type of data we consider. We provide explicit results on the parameters of the mixture components and on the mixing weights, which are time-varying and drive the mixtures towards the respective priors in absence of further data. Explicit algorithms are provided to recursively compute the parameters of the mixtures. Our results are based on the projective properties of the signals and on certain duality properties of their projections.

math.ST↗

Canonical correlations for dependent gamma processes

The present paper provides a characterisation of exchangeable pairs of random measures $(\widetildeμ_1,\widetildeμ_2)$ whose identical margins are fixed to coincide with the distribution of a gamma completely random measure, and whose dependence structure is given in terms of canonical correlations. It is first shown that canonical correlation sequences for the finite-dimensional distributions of $(\widetildeμ_1,\widetildeμ_2)$ are moments of means of a Dirichlet process having random base measure. Necessary and sufficient conditions are further given for canonically correlated gamma completely random measures to have independent joint increments. Finally, time-homogeneous Feller processes with gamma reversible measure and canonical autocorrelations are characterised as Dawson--Watanabe diffusions with independent homogeneous immigration, time-changed via an independent subordinator. It is thus shown that Dawson--Watanabe diffusions subordinated by pure drift are the only processes in this class whose time-finite-dimensional distributions have, jointly, independent increments.

math.PR↗

Filtering hidden Markov measures

We consider the problem of learning two families of time-evolving random measures from indirect observations. In the first model, the signal is a Fleming--Viot diffusion, which is reversible with respect to the law of a Dirichlet process, and the data is a sequence of random samples from the state at discrete times. In the second model, the signal is a Dawson--Watanabe diffusion, which is reversible with respect to the law of a gamma random measure, and the data is a sequence of Poisson point configurations whose intensity is given by the state at discrete times. A common methodology is developed to obtain the filtering distributions in a computable form, which is based on the projective properties of the signals and duality properties of their projections. The filtering distributions take the form of mixtures of Dirichlet processes and gamma random measures for each of the two families respectively, and an explicit algorithm is provided to compute the parameters of the mixtures. Hence, our results extend classic characterisations of the posterior distribution under Dirichlet process and gamma random measures priors to a dynamic framework.

math.ST↗

The ancestral process of long term seed bank models

We present a new model for seed banks, where direct ancestors of individuals may have lived in the near as well as the very far past. The classical Wright-Fisher model, as well as a seed bank model with bounded age distribution considered by Kaj, Krone and Lascoux (2001) are special cases of our model. We discern three parameter regimes of the seed bank age distribution, which lead to substantially different behaviour in terms of genetic variability, in particular with respect to fixation of types and time to the most recent common ancestor. We prove that for age distributions with finite mean, the ancestral process converges to a time-changed Kingman coalescent, while in the case of infinite mean, ancestral lineages might not merge at all with positive probability. Further, we present a construction of the forward in time process in equilibrium. The mathematical methods are based on renewal theory, the urn process introduced by Kaj et al., as well as on a paper by Hammond and Sheffield (2011).

math.PR↗

Orthogonal polynomial kernels and canonical correlations for Dirichlet measures

We consider a multivariate version of the so-called Lancaster problem of characterizing canonical correlation coefficients of symmetric bivariate distributions with identical marginals and orthogonal polynomial expansions. The marginal distributions examined in this paper are the Dirichlet and the Dirichlet multinomial distribution, respectively, on the continuous and the N-discrete d-dimensional simplex. Their infinite-dimensional limit distributions, respectively, the Poisson-Dirichlet distribution and Ewens's sampling formula, are considered as well. We study, in particular, the possibility of mapping canonical correlations on the d-dimensional continuous simplex (i) to canonical correlation sequences on the d+1-dimensional simplex and/or (ii) to canonical correlations on the discrete simplex, and vice versa. Driven by this motivation, the first half of the paper is devoted to providing a full characterization and probabilistic interpretation of n-orthogonal polynomial kernels (i.e., sums of products of orthogonal polynomials of the same degree n) with respect to the mentioned marginal distributions. We establish several identities and some integral representations which are multivariate extensions of important results known for the case d=2 since the 1970s. These results, along with a common interpretation of the mentioned kernels in terms of dependent Polya urns, are shown to be key features leading to several non-trivial solutions to Lancaster's problem, many of which can be extended naturally to the limit as $d\rightarrow\infty$.

math.PR↗