SearcharxivSearch

arXiv subjects

Peter Pfaffelhuber

Publications and source records attributed to Peter Pfaffelhuber.

At least 19 recordsLinked to original sources

A small noise approximation for Muller's Ratchet

We consider an infinite system of SDEs with Fleming-Viot noise indexed by $k=0,1,2,\dots$, whose parameters $\alpha,\lambda$, and $\nu$ are the (deleterious) selection coefficient, the (uni-directional) mutation rate, and a quantity which determines the size of the system's fluctuations. The SDE's unique weak solution $X(t) = (X_k(t))_{k=0,1,2,...}$ models what is known in population genetics as Muller's ratchet. Here, $X_k(t)$ stands for the frequency of individuals carrying $k$ deleterious mutations. Since the mutation process is uni-directional, $t\mapsto \inf\{k: X_k(t)> 0\}$ is non-decreasing for almost every path of $X$, and we refer to an increase as a click of Muller's ratchet. A long standing question concerns the clicking rate of Muller's ratchet. Using Duhamel's principle for semigroups, we give a partial answer by approximating $E(\sum_{k=1}^\infty kX_k(t) )$ and $E\big(X_0(t)\big)$ up to $O(1/\nu^2)$ for fixed $\alpha$, $\lambda$ and $t>0$. Our results suggest that $\psi:=\nu \alpha e^{-\lambda/\alpha}$ is a crucial quantity also when the mutation/selection ratio $\theta = \lambda/\alpha$ is moderately large: for large $\nu \alpha$, clicking of the ratchet on the time scale $\frac 1\alpha \log \theta$ becomes rare as soon as $\psi$ becomes large.

math.PR

Typical Healthcare Pathways as a Basis for Admixture Modeling of Patient Trajectories

Background: Understanding whether patients follow similar or distinct patterns of care is important for characterizing clinical practice, identifying patient subgroups, and supporting quality improvement. However, routine healthcare trajectories are difficult to compare directly because patients may differ in their diagnostic workup, treatment sequencing, timing of clinical events, and documentation practices. Despite this variation, trajectories often contain recurring patterns at the cohort level. Methods: To address this challenge, we present a framework that explicitly separates cohort-level typical pathway identification from patient-level inference. At the cohort level, we derive an interpretable representation of care processes using a rule-based algorithm to identify typical healthcare pathways, resulting in a compact pathway graph. These pathways are then modeled as Markov chains and used as structured components in an admixture model, allowing each patient to be represented as a probabilistic mixture of typical pathways rather than being assigned to a single pathway component. The resulting admixture weights provide a compact representation of patient trajectories for subgroup characterization. We further assess the stability of the identified pathways and inferred admixture representations across multiple train-test splits. Results: Across train-test splits, the framework demonstrated consistent pathway structures and patient-level mixture patterns. Applied to routine care data from prostate cancer patients undergoing radical prostatectomy, the framework identified interpretable care patterns and supported the identification of patient subgroups with similar clinical event patterns. Conclusion: Overall, the proposed framework provides an interpretable and stable approach for summarizing treatment pathways and characterizing patient subgroups in real-world practice.

stat.ME

Testing for Single-Population Ancestry in the Admixture Model

The Admixture Model describes genetic marker data by representing each individual's genome as a mixture of contributions from $K$ ancestral populations, with the individual admixture vector summarizing the corresponding ancestry proportions. In population and forensic genetics, a key question is whether an individual's genome supports a predominantly single-ancestry interpretation or whether an admixed interpretation is more appropriate. We propose a statistical test for single-population ancestry in the supervised Admixture Model, where ancestral allele frequencies are treated as known. The test assesses whether the largest admixture component exceeds a practitioner-chosen dominance threshold, giving precise meaning to the notion of a sufficiently strong single-population contribution. To calibrate the test, we develop a constrained parametric bootstrap procedure that generates data under a null-constrained maximum likelihood estimator, accounting for the constrained hypothesis structure, the marker-wise heterogeneity and small sample sizes. Under standard regularity conditions, we prove that the proposed test has asymptotic level $\alpha$ and is consistent, ensuring control of false single-ancestry declarations while reliably detecting dominant ancestry components. Simulation studies demonstrate good finite-sample performance across different numbers of ancestral populations, marker-panel sizes, dominance thresholds, and allele-frequency distributions. We further illustrate the practical utility of the method using data from the 1000 Genomes Project. The proposed framework delivers interpretable, threshold-based ancestry assessment with rigorous error control, and extends constrained bootstrap methodology to the independent but non-identically distributed setting of genetic marker data.

stat.ME

Optimal strategies in the all-heads coin game

We study a sequential coin-flipping game: a player starts with $n$~coins, each heads with probability~$p$, and in each round flips all remaining coins and must set aside at least one head, losing if none shows. The player wins once all coins have been set aside. The optimal winning probability~$w_{n,p}$ obeys a Bellman equation with a nonlinear suffix-maximum operator. For $p=\tfrac12$ every strategy achieves $w_{n,1/2}=\tfrac12$. For $p>\tfrac12$ the strategy~\One{} (set aside a single head) is optimal, $n\mapsto w_{n,p}$ is strictly increasing, and the limit $W(p):=\lim_n w_{n,p}$ has an explicit series representation with $p\le W(p)<1$. For $p<\tfrac12$ near~$\tfrac12$ we give a first-order perturbation expansion in $\delta:=\tfrac12-p$: the deficit satisfies $\tfrac12-w_{n,\,1/2-\delta}\approx\delta\,c_n$, where $c_n$ obeys a linear recursion for $n\ge7$ with limit $L\approx1.7035$. To first order the optimal-value sequence has a strict local minimum at $n=5$ and no local maximum.

math.PR

Markov processes forced on a subspace by a large drift, with applications to population genetics

Consider a sequence of Markov processes $X^1, X^2,...$ with state space $E$, where $X^N$ has a strong drift to $D \subseteq E$, such that $\Phi(X^N)$ is slow for some appropriate $\Phi: E\to D$. Using the method of martingale problems, we give a limit result, such that $\Phi(X^N) \xRightarrow{N\to\infty} Z$ in the space of c\`adl\`ag paths, and $X^N \xRightarrow{N\to\infty} X$ in measure. \\ We apply the general limit result to models for copy number variation of genetic elements in a diploid Moran model of size $N$. The population by time $t$ is described by $X^N \in \mathcal P(\mathbb N_0)$, where $X^N_k$ is the frequency of individuals with copy number $k$, and $\Phi: \mathcal P(\mathbb

math.PR

Formalization of Brownian motion in Lean

Brownian motion is a building block in modern probability theory. In this paper, we describe a formalization of Brownian motion using the Lean theorem prover. We build on the existing measure-theoretic foundations in Lean's mathematical library, Mathlib, and we develop several key components needed for the construction of Brownian motion, including the Carath\'eodory and Kolmogorov extension theorems, Gaussian measures in Banach spaces, and the Kolmogorov-Chentsov theorem for path continuity.

math.PR

Duality and the well-posedness of a martingale problem

For two Polish state spaces $E_X$ and $E_Y$, and an operator $G_X$, we obtain existence and uniqueness of a $G_X$-martingale problem provided there is a bounded continuous duality function $H$ on $E_X \times E_Y$ together with a dual process $Y$ on $E_Y$ which is the unique solution of a $G_Y$-martingale problem. For the corresponding solutions $(X_t)_{t\ge 0}$ and $(Y_t)_{t\ge 0}$, duality with respect to a function $H$ in its simplest form means that the relation $\mathbb E_x[H(X_t,y)] = \mathbb E_y[H(x,Y_t)]$ holds for all $(x,y) \in E_X \times E_Y$ and $t\ge 0$. While duality is well-known to imply uniqueness of the $G_X$-martingale problem, we give here a set of conditions under which duality also implies existence without using approximating sequences of processes of a different kind (e.g.\ jump processes to approximate diffusions) which is a widespread strategy for proving existence of solutions of martingale problems. Given the process $(Y_t)_{t\ge 0}$ and a duality function $H$, to prove existence of $(X_t)_{t\ge 0}$ one has to show that the r.h.s.\ of the duality relation defines for each $y$ a measure on $E_X$, i.e.\ there are transition kernels $(μ_t)_{t\geq 0}$ from $E_X$ to $E_X$ such that $\mathbb E_y[H(x,Y_t)] = \int μ_t(x,dx')\, H(x',y)$ for all $(x,y) \in E_X \times E_Y$ and all $t\geq 0$. As examples, we treat resampling and branching models, such as the Fleming-Viot measure-valued diffusion and its spatial counterparts (with both, discrete and continuum space), as well as branching systems, such as Feller's branching diffusion. While our main result as well as all examples come with (locally) compact state spaces, we discuss the strategy to lift our results to genealogy-valued processes or historical processes, leading to non-compact (discrete and continuum) state spaces. Such applications will be tackled in forthcoming work based on the present article.

math.PR

A diploid population model for copy number variation of genetic elements

We study the following model for a diploid population of constant size $N$: Every individual carries a random number of (genetic) elements. Upon a reproduction event each of the two parents passes each element independently with probability $\tfrac 12$ on to the offspring. We study the process $X^N = (X^N(1), X^N(2),...)$, where $X_t^N(k)$ is the frequency of individuals at time $t$ that carry $k$ elements, and prove convergence (in some weak sense) of $X^N$ jointly with its empirical first moment $Z^N$ to the ``slow-fast'' system $(Z,X)$, where $X_t = \text{Poi}(Z_t)$ and $Z$ evolves according to a critical Feller branching process. We discuss heuristics explaining this finding and some extensions and limitations.

math.PR

A central limit theorem concerning uncertainty in estimates of individual admixture

The concept of individual admixture (IA) assumes that the genome of individuals is composed of alleles inherited from $K$ ancestral populations. Each copy of each allele has the same chance $q_k$ to originate from population $k$, and together with the allele frequencies $p$ in all populations at all $M$ markers, comprises the admixture model. Here, we assume a supervised scheme, i.e.\ allele frequencies $p$ are given through a reference database of size $N$, and $q$ is estimated via maximum likelihood for a single sample. We study laws of large numbers and central limit theorems describing effects of finiteness of both, $M$ and $N$, on the estimate of $q$. We recall results for the effect of finite $M$, and provide a central limit theorem for the effect of finite $N$, introduce a new way to express the uncertainty in estimates in standard barplots, give simulation results, and discuss applications in forensic genetics.

q-bio.PE

A unified framework for limit results in Chemical Reaction Networks on multiple time-scales

If $(X^N)_{N=1,2,...}$ is a sequence of Markov processes which solve the martingale problems for some operators $(G^N)_{N=1,2,...}$, it is a classical task to derive a limit result as $N\to\infty$, in particular a weak process limit with limiting operator $G$. For slow-fast systems $X^N = (V^N, Z^N)$ where $V^N$ is slow and $Z^N$ is fast, $G^N$ consists of two (or more) terms, and we are interested in weak convergence of $V^N$ to some Markov process $V$. In this case, for some $f\in \mathcal D(G)$, the domain of $G$, depending only on $v$, the limit $Gf$ can sometimes be derived by using some $g_N\to 0$ (depending on $v$ and $z$), and study convergence of $G^N (f + g_N) \to Gf$. We develop this method further in order to obtain functional Laws of Large Numbers (LLNs) and Central Limit Theorems (CLTs). We then apply our general result to various examples from Chemical Reaction Network theory. We show that we can rederive most limits previously obtained, but also provide new results in the case when the fast-subsystem is first order. In particular, we allow that fast species to be consumed faster than they are produced, and we derive a CLT for Hill dynamics with coefficient~2.

math.PR

Mean-field limits for non-linear Hawkes processes with excitation and inhibition

We study a multivariate, non-linear Hawkes process $Z^N$ on the complete graph with $N$ nodes. Each vertex is either excitatory (probability $p$) or inhibitory (probability $1-p$). We take the mean-field limit of $Z^N$, leading to a multivariate point process $\bar Z$. If $p\neq\tfrac12$, we rescale the interaction intensity by $N$ and find that the limit intensity process solves a deterministic convolution equation and all components of $\bar Z$ are independent. In the critical case, $p=\tfrac12$, we rescale by $N^{1/2}$ and obtain a limit intensity, which solves a stochastic convolution equation and all components of $\bar Z$ are conditionally independent.

math.PR

The Martingale Problem Method Revisited

We use the abstract method of (local) martingale problems in order to give criteria for convergence of stochastic processes. Extending previous notions, the formulation we use is neither restricted to Markov processes (or semimartingales), nor to continuous or cadlag paths. We illustrate our findings both, by finding generalizations of known results, and proving new results. For the latter, we work on processes with fixed times of discontinuity.

math.PR

Modifiers of mutation rate in selectively fluctuating environments

We study a mutation-selection model with a fluctuating environment. More precisely, individuals in a large population are assumed to have a modifier locus determining the mutation rate $u \in [0,\vartheta]$ at a second locus with types $v \in [0,1]$. In addition, the environment fluctuates, meaning that individual types change their fitness at some high rate. Fitness only depends on the type of the second locus. We obtain general limit results for the evolution of the allele frequency distribution for rapidly fluctuating environments. As an application, we make use of the resulting Fleming-Viot process and compute the fixation probabilities for higher mutation rates in the special case of two bi-allelic loci in the limit of small fitness differences at the second locus.

math.PR

The partial duplication random graph with edge deletion

We study a random graph model in continuous time. Each vertex is partially copied with the same rate, i.e.\ an existing vertex is copied and every edge leading to the copied vertex is copied with independent probability $p$. In addition, every edge is deleted at constant rate, a mechanism which extends previous partial duplication models. In this model, we obtain results on the degree distribution, which shows a phase transition such that either -- if $p$ is small enough -- the frequency of isolated vertices converges to 1, or there is a positive fraction of vertices with unbounded degree. We derive results on the degrees of the initial vertices as well as on the sub-graph of non-isolated vertices. In particular, we obtain expressions for the number of star-like subgraphs and cliques.

math.PR

Fixation probabilities and hitting times under low levels of frequency-dependent selection

In population genetics, diffusions on the unit interval are often used to model the frequency path of an allele. In this setting we derive approximations for fixation probabilities, expected hitting times and the expected site-frequency-spectrum under small frequency-dependent selection. Specifically, we rederive and extend the one-third rule of evolutionary game theory (Nowak et al., 2004) and effects of stochastic slowdown (Altrock and Traulsen, 2009). Since similar effects are of interest in other application areas, we formulate our results for general one-dimensional diffusions.

math.PR

Stochastic averaging for multiscale Markov processes with an application to a Wright-Fisher model with fluctuating selection

Let $Z = (Z_t)_{t\in[0,\infty)}$ be an ergodic Markov process and, for every $n\in\mathbb{N}$, let $Z^n = (Z_{n^2 t})_{t\in[0,\infty)}$ drive a process $X^n$. Classical results show under suitable conditions that the sequence of non-Markovian processes $(X^n)_{n\in\mathbb{N}}$ converges to a Markov process and give its infinitesimal characteristics. Here, we consider a general sequence $(Z^n)_{n\in\mathbb{N}}$. Using a general result on stochastic averaging from [Kur92], we derive conditions which ensure that the sequence $(X^n)_{n\in\mathbb{N}}$ converges as in the classical case. As an application, we consider the diffusion limit of a Wright-Fisher model with fluctuating selection.

math.PR

The fixation probability and time for a doubly beneficial mutant

For a highly beneficial mutant $A$ entering a randomly reproducing population of constant size, we study the situation when a second beneficial mutant $B$ arises before $A$ has fixed. If the selection coefficient of $B$ is greater than the selection coefficient of $A$, and if $A$ and $B$ can recombine at some rate $ρ$, there is a chance that the double beneficial mutant $AB$ forms and eventually fixes. We give a convergence result for the fixation probability of $AB$ and its fixation time for large selection coefficients.

math.PR