Searcharxiv⌕ Search

arXiv subjects

Lancelot F. James

Publications and source records attributed to Lancelot F. James.

At least 19 recordsLinked to original sources

Learning from Neighbors with PHIBP: Predicting Infectious Disease Dynamics in Data-Sparse Environments

Modeling sparse count data, which arise across numerous scientific fields, presents significant statistical challenges. This chapter addresses these challenges in the context of infectious disease prediction, with a focus on predicting outbreaks in geographic regions that have historically reported zero cases. To this end, we present the detailed computational framework and experimental application of the Poisson Hierarchical Indian Buffet Process (PHIBP), with demonstrated success in handling sparse count data in microbiome and ecological studies. The PHIBP's architecture, grounded in the concept of absolute abundance, systematically borrows statistical strength from related regions and circumvents the known sensitivities of relative-rate methods to zero counts. Through a series of experiments on infectious disease data, we show that this principled approach provides a robust foundation for generating coherent predictive distributions and for the effective use of comparative measures such as alpha and beta diversity. The chapter's emphasis on algorithmic implementation and experimental results confirms that this unified framework delivers both accurate outbreak predictions and meaningful epidemiological insights in data-sparse settings.

stat.ML↗

Coagulation-Fragmentation Duality of Infinitely Exchangeable Partitions from Coupled Mixed Poisson Species Sampling Models

We generalize the celebrated coagulation-fragmentation duality of Pitman (1999), originally established for the PD$(α,θ)$ laws of Pitman and Yor (1997), resolving a two-decade open problem. Our framework extends the duality to processes driven by arbitrary non-negative L'evy subordinators and, for the first time, to multi-group settings with coupled dynamics. The solution is a novel four-component system built from the PHIBP, a framework developed for modeling complex microbiome species sampling (arXiv:2502.01919), which circumvents intractable analysis on traditional partition spaces. Crucially, this architecture embeds naturally within Bertoin's (2006) continuous-time fragmentation framework, resolving a foundational impasse he highlighted [Ch. 4, p. 213]: where time-reversal fails, we introduce simultaneous structural duality, where fragmentation and coalescence evolve in physical time while maintaining a pointwise dual relationship via coupled subordinators. This provides a new generative modeling framework enabling continuous-time representations of coupled genealogical and mutational dynamics -- opening avenues for complex ancestral processes such as recombination graphs. The architecture yields exact compound Poisson representations providing tractable paths to explicit joint EPPF laws, with exact sampling inherited from the PHIBP framework. Section 7 develops the $h$-biased Lévy-Itô coupled duality constructor, extending our framework to arbitrary Polish spaces by adapting the $h$-biased PRM structure of Pitman and Yor to Cox processes driven by a common PRM. This demonstrates that duality arises from point-process regrouping rather than features specific to interval partitions.

math.PR↗

Poisson Hierarchical Indian Buffet Processes-With Indications for Microbiome Species Sampling Models

We introduce the Poisson Hierarchical Indian Buffet Process (PHIBP), a new class of species sampling models designed to address the challenges of complex, sparse count data by facilitating information sharing across and within groups. Our theoretical developments enable a tractable Bayesian nonparametric framework with machine learning elements, accommodating a potentially infinite number of species (taxa) whose parameters are learned from data. Focusing on microbiome analysis, we address key gaps by providing a flexible multivariate count model that accounts for overdispersion and robustly handles diverse data types (OTUs, ASVs). We introduce novel parameters reflecting species abundance and diversity. The model borrows strength across groups while explicitly distinguishing between technical and biological zeros to interpret sparse co-occurrence patterns. This results in a framework with tractable posterior inference, exact generative sampling, and a principled solution to the unseen species problem. We describe extensions where domain experts can incorporate knowledge through covariates and structured priors, with potential for strain-level analysis. While motivated by ecology, our work provides a broadly applicable methodology for hierarchical count modeling in genetics, commerce, and text analysis, and has significant implications for the broader theory of species sampling models arising in probability and statistics.

stat.ML↗

Network and interaction models for data with hierarchical granularity via fragmentation and coagulation

We introduce a nested family of Bayesian nonparametric models for network and interaction data with a hierarchical granularity structure that naturally arises through finer and coarser population labelings. In the case of network data, the structure is easily visualized by merging and shattering vertices, while respecting the edge structure. We further develop Bayesian inference procedures for the model family, and apply them to synthetic and real data. The family provides a connection of practical and theoretical interest between the Hollywood model of Crane and Dempsey, and the generalized-gamma graphex model of Caron and Fox. A key ingredient for the construction of the family is fragmentation and coagulation duality for integer partitions, and for this we develop novel duality relations that generalize those of Pitman and Dong, Goldschmidt and Martin. The duality is also crucially used in our inferential procedures.

math.ST↗

Modelling financial volume curves with hierarchical Poisson processes

Modeling the trading volume curves of financial instruments throughout the day is of key interest in financial trading applications. Predictions of these so-called volume profiles guide trade execution strategies, for example, a common strategy is to trade a desired quantity across many orders in line with the expected volume curve throughout the day so as not to impact the price of the instrument. The volume curves (for each day) are naturally grouped by stock and can be further gathered into higher-level groupings, such as by industry. In order to model such admixtures of volume curves, we introduce a hierarchical Poisson process model for the intensity functions of admixtures of inhomogenous Poisson processes, which represent the trading times of the stock throughout the day. The model is based on the hierarchical Dirichlet process, and an efficient Markov Chain Monte Carlo (MCMC) algorithm is derived following the slice sampling framework for Bayesian nonparametric mixture models. We demonstrate the method on datasets of different stocks from the Trade and Quote repository maintained by Wharton Research Data Services, including the most liquid stock on the NASDAQ stock exchange, Apple, demonstrating the scalability of the approach.

q-fin.ST↗

Posterior distributions of Gibbs-type priors

Gibbs type priors have been shown to be natural generalizations of Dirichlet process (DP) priors used for intricate applications of Bayesian nonparametric methods. This includes applications to mixture models and to species sampling models arising in populations genetics. Notably these latter applications, and also applications where power law behavior such as that arising in natural language models are exhibited, provide instances where the DP model is wholly inadequate. Gibbs type priors include the DP, the also popular Pitman-Yor process and closely related normalized generalized gamma process as special cases. However, there is in fact a richer infinite class of such priors, where, despite knowledge about the exchangeable marginal structures produced by sampling $n$ observations, descriptions of the corresponding posterior distribution, a crucial component in Bayesian analysis, remain unknown. This paper presents descriptions of the posterior distributions for the general class, utilizing a novel proof that leverages the exclusive Gibbs properties of these models. The results are applied to several specific cases for further illustration.

math.ST↗

Inverse clustering of Gibbs Partitions via independent fragmentation and dual dependent coagulation operators

Gibbs partitions of the integers generated by stable subordinators of index $α\in(0,1)$ form remarkable classes of random partitions where in principle much is known about their properties, including practically effortless obtainment of otherwise complex asymptotic results potentially relevant to applications in general combinatorial stochastic processes, random tree/graph growth models and Bayesian statistics. This class includes the well-known models based on the two-parameter Poisson-Dirichlet distribution which forms the bulk of explicit applications. This work continues efforts to provide interpretations for a larger classes of Gibbs partitions by embedding important operations within this framework. Here we address the formidable problem of extending the dual, infinite-block, coagulation/fragmentation results of Jim Pitman (1999, Annals of Probability), where in terms of coagulation they are based on independent two-parameter Poisson-Dirichlet distributions, to all such Gibbs (stable Poisson-Kingman) models. Our results create nested families of Gibbs partitions, and corresponding mass partitions, over any $0<β<α<1.$ We primarily focus on the fragmentation operations, which remain independent in this setting, and corresponding remarkable calculations for Gibbs partitions derived from that operation. We also present definitive results for the dual coagulation operations, now based on our construction of dependent processes, and demonstrate its relatively simple application in terms of Mittag-Leffler and generalized gamma models. The latter demonstrates another approach to recover the duality results in Pitman (1999).

math.PR↗

Posterior distributions for Hierarchical Spike and Slab Indian Buffet processes

Bayesian nonparametric hierarchical priors are highly effective in providing flexible models for latent data structures exhibiting sharing of information between and across groups. Most prominent is the Hierarchical Dirichlet Process (HDP), and its subsequent variants, which model latent clustering between and across groups. The HDP, may be viewed as a more flexible extension of Latent Dirichlet Allocation models (LDA), and has been applied to, for example, topic modelling, natural language processing, and datasets arising in health-care. We focus on analogous latent feature allocation models, where the data structures correspond to multisets or unbounded sparse matrices. The fundamental development in this regard is the Hierarchical Indian Buffet process (HIBP), which utilizes a hierarchy of Beta processes over J groups, where each group generates binary random matrices, reflecting within group sharing of features, according to beta-Bernoulli IBP priors. To encompass HIBP versions of non-Bernoulli extensions of the IBP, we introduce hierarchical versions of general spike and slab IBP. We provide explicit novel descriptions of the marginal, posterior and predictive distributions of the HIBP and its generalizations which allow for exact sampling and simpler practical implementation. We highlight common structural properties of these processes and establish relationships to existing IBP type and related models arising in the literature. Examples of potential applications may involve topic models, Poisson factorization models, random count matrix priors and neural network models

math.ST↗

A Bayesian model for sparse graphs with flexible degree distribution and overlapping community structure

We consider a non-projective class of inhomogeneous random graph models with interpretable parameters and a number of interesting asymptotic properties. Using the results of Bollobás et al. [2007], we show that i) the class of models is sparse and ii) depending on the choice of the parameters, the model is either scale-free, with power-law exponent greater than 2, or with an asymptotic degree distribution which is power-law with exponential cut-off. We propose an extension of the model that can accommodate an overlapping community structure. Scalable posterior inference can be performed due to the specific choice of the link probability. We present experiments on five different real-world networks with up to 100,000 nodes and edges, showing that the model can provide a good fit to the degree distribution and recovers well the latent community structure.

stat.ML↗

Gibbs Partitions, Riemann-Liouville Fractional Operators, Mittag-Leffler Functions, and Fragmentations Derived From Stable Subordinators

Pitman(2003)(and subsequently Gnedin and Pitman (2006) showed that a large class of random partitions of the integers derived from a stable subordinator of index $α\in(0,1)$ have infinite Gibbs (product) structure as a characterizing feature. The most notable case are random partitions derived from the two-parameter Poisson-Dirichlet distribution, $\mathrm{PD}(α,θ)$, which are induced by mixing over variables with generalized Mittag-Leffler distributions, denoted by $\mathrm{ML}(α,θ).$ Our aim in this work is to provide indications on the utility of the wider class of Gibbs partitions as it relates to a study of Riemann-Liouville fractional integrals and size-biased sampling, decompositions of special functions, and its potential use in the understanding of various constructions of more exotic processes. We provide novel characterizations of general laws associated with two nested families of $\mathrm{PD}(α,θ)$ mass partitions that are constructed from notable fragmentation operations described in Dong, Goldschmidt and Martin(2006) and Pitman(1999), respectively. These operations are known to be related in distribution to various constructions of discrete random trees/graphs in $[n],$ and their scaling limits, such as stable trees. A centerpiece of our work are results related to Mittag-Leffler functions, which play a key role in fractional calculus and are otherwise Laplace transforms of the $\mathrm{ML}(α,θ)$ variables. Notably, this leads to an interpretation of $\mathrm{PD}(α,θ)$ laws within a mixed Poisson waiting time framework based on $\mathrm{ML}(α,θ)$ variables, which suggests connections to recent construction of Pólya urn models with random immigration by Peköz, Röllin and Ross(2018). Simplifications in the Brownian case are highlighted.

math.PR↗

Bayesian inference on random simple graphs with power law degree distributions

We present a model for random simple graphs with a degree distribution that obeys a power law (i.e., is heavy-tailed). To attain this behavior, the edge probabilities in the graph are constructed from Bertoin-Fujita-Roynette-Yor (BFRY) random variables, which have been recently utilized in Bayesian statistics for the construction of power law models in several applications. Our construction readily extends to capture the structure of latent factors, similarly to stochastic blockmodels, while maintaining its power law degree distribution. The BFRY random variables are well approximated by gamma random variables in a variational Bayesian inference routine, which we apply to several network datasets for which power law degree distributions are a natural assumption. By learning the parameters of the BFRY distribution via probabilistic inference, we are able to automatically select the appropriate power law behavior from the data. In order to further scale our inference procedure, we adopt stochastic gradient ascent routines where the gradients are computed on minibatches (i.e., subsets) of the edges in the graph.

stat.ML↗

Independence by Random Scaling

We give conditions under which a scalar random variable T can be coupled to a random scaling factor $ξ$ such that T and $ξ$T are rendered stochastically independent. A similar result is obtained for random measures. One consequence is a generalization of a result by Pitman and Yor on the Poisson-Dirichlet distribution to its negative parameter range. Another application are diffusion excursions straddling an exponential random time.

math.PR↗

Scaled subordinators and generalizations of the Indian buffet process

We study random families of subsets of $\mathbb{N}$ that are similar to exchangeable random partitions, but do not require constituent sets to be disjoint: Each element of ${\mathbb{N}}$ may be contained in multiple subsets. One class of such objects, known as Indian buffet processes, has become a popular tool in machine learning. Based on an equivalence between Indian buffet and scale-invariant Poisson processes, we identify a random scaling variable whose role is similar to that played in exchangeable partition models by the total mass of a random measure. Analogous to the construction of exchangeable partitions from normalized subordinators, random families of sets can be constructed from randomly scaled subordinators. Coupling to a heavy-tailed scaling variable induces a power law on the number of sets containing the first $n$ elements. Several examples, with properties desirable in applications, are derived explicitly. A relationship to exchangeable partitions is made precise as a correspondence between scaled subordinators and Poisson-Kingman measures, generalizing a result of Arratia, Barbour and Tavare on scale-invariant processes.

math.PR↗

Generalized Mittag Leffler distributions arising as limits in preferential attachment models

For $0<α<1,$ and $θ>-α,$ let $(S^{-α}_{α,θ+r})_{\{r\ge 0\}}$ denote an increasing(decreasing) sequence of variables forming a time inhomogeneous Markov chain whose marginal distributions are equivalent to generalized Mittag Leffler distributions. We exploit the property that such a sequence may be connected with the two parameter $(α,θ)$ family of Poisson Dirichlet distributions. We demonstrate that the sequences serve as limits in certain types of preferential attachment models. As one illustrative application, we describe the explicit joint limiting distribution of scaled degree sequences arising under a class of linear weighted preferential attachment models as treated in Móri (2005), with weight $β>-1.$ When $β=0$ this corresponds to the Barbasi-Albert preferential attachment model. We then construct sequences of nested $(α,θ)$ Chinese restaurant partitions of $[n]$. From this, we identify and analyze relevant quantities that may be thought of as mimics for vectors of degree sequences, or differences in tree lengths. We also describe connections to a wide class of continuous time coalescent processes that can be seen as a variation of stochastic flows of bridges related to generalized Fleming-Viot models. Under a change of measure our results suggest the possibilities for identification of limiting distributions related to consistent families of nested Gibbs partitions of $[n]$ that would otherwise be difficult by methods using moments or Laplace transforms. In this regard, we focus on special simplifications obtained in the case of $α=1/2.$ That is to say, limits derived from a $\mathrm{PD}(1/2|t)$ distribution. Throughout we present some distributional results that are relevant to various settings. We describe nestings across the $α$ parameter in section 6

math.PR↗

Poisson Latent Feature Calculus for Generalized Indian Buffet Processes

The purpose of this work is to describe a unified, and indeed simple, mechanism for non-parametric Bayesian analysis, construction and generative sampling of a large class of latent feature models which one can describe as generalized notions of Indian Buffet Processes(IBP). This is done via the Poisson Process Calculus as it now relates to latent feature models. The IBP was ingeniously devised by Griffiths and Ghahramani in (2005) and its generative scheme is cast in terms of customers entering sequentially an Indian Buffet restaurant and selecting previously sampled dishes as well as new dishes. In this metaphor dishes corresponds to latent features, attributes, preferences shared by individuals. The IBP, and its generalizations, represent an exciting class of models well suited to handle high dimensional statistical problems now common in this information age. The IBP is based on the usage of conditionally independent Bernoulli random variables, coupled with completely random measures acting as Bayesian priors, that are used to create sparse binary matrices. This Bayesian non-parametric view was a key insight due to Thibaux and Jordan (2007). One way to think of generalizations is to to use more general random variables. Of note in the current literature are models employing Poisson and Negative-Binomial random variables. However, unlike their closely related counterparts, generalized Chinese restaurant processes, the ability to analyze IBP models in a systematic and general manner is not yet available. The limitations are both in terms of knowledge about the effects of different priors and in terms of models based on a wider choice of random variables. This work will not only provide a thorough description of the properties of existing models but also provide a simple template to devise and analyze new models.

math.ST↗

Exact simulation pricing with Gamma processes and their extensions

Exact path simulation of the underlying state variable is of great practical importance in simulating prices of financial derivatives or their sensitivities when there are no analytical solutions for their pricing formulas. However, in general, the complex dependence structure inherent in most nontrivial stochastic volatility (SV) models makes exact simulation difficult. In this paper, we present a nontrivial SV model that parallels the notable Heston SV model in the sense of admitting exact path simulation as studied by Broadie and Kaya. The instantaneous volatility process of the proposed model is driven by a Gamma process. Extensions to the model including superposition of independent instantaneous volatility processes are studied. Numerical results show that the proposed model outperforms the Heston model and two other Lévy driven SV models in terms of model fit to the real option data. The ability to exactly simulate some of the path-dependent derivative prices is emphasized. Moreover, this is the first instance where an infinite-activity volatility process can be applied exactly in such pricing contexts.

q-fin.CP↗

Stick-breaking PG(α,ζ)-Generalized Gamma Processes

This work centers around results related to Proposition 21 of Pitman and Yor's (1997) paper on the two parameter Poisson Dirichlet distribution indexed by (α,θ) for 0<α<1, also α=0, and θ>-α, denoted PD(α,θ). We develop explicit stick-breaking representations for a class that contains the PD(α,θ) for the range θ=0, and θ>0, we call PG(α,ζ). We also construct a larger class, EPG(α,ζ), containing the entire range. These classes are indexed by α, and an arbitrary non-negative random variable ζ. The bulk of this work focuses on investigating various properties of this larger class, EPG(α,ζ), which lead to connections to other work in the literature. In particular, we develop completely explicit stick-breaking representations for this entire class via size biased sampling as described in Perman, Pitman and Yor (1992). This represents the first case outside of the PD(α,θ) where one obtains explicit results for the entire range of α, the EPG are within the larger class of mass partitions generated by conditioning on the total mass of an α-stable subordinator. Furthermore Markov chains are derived which establish links between Markov chains derived from stick-breaking(insertion/deletion), as described in Perman, Pitman and Yor and Markov chains derived from successive usage of dual coagulation fragmentation operators described in Bertoin and Goldschmidt and Dong, Goldschmidt and Martin. Which have connections to certain types of fragmentation trees and coalescents appearing in the recent literature. Our results are also suggestive of new models and tools, for applications in Bayesian Nonparametrics/Machine Learning, where PD(α,θ) bridges are often referred to as Pitman-Yor processes.

math.PR↗

Quantile clocks

Quantile clocks are defined as convolutions of subordinators $L$, with quantile functions of positive random variables. We show that quantile clocks can be chosen to be strictly increasing and continuous and discuss their practical modeling advantages as business activity times in models for asset prices. We show that the marginal distributions of a quantile clock, at each fixed time, equate with the marginal distribution of a single subordinator. Moreover, we show that there are many quantile clocks where one can specify $L$, such that their marginal distributions have a desired law in the class of generalized $s$-self decomposable distributions, and in particular the class of self-decomposable distributions. The development of these results involves elements of distribution theory for specific classes of infinitely divisible random variables and also decompositions of a gamma subordinator, that is of independent interest. As applications, we construct many price models that have continuous trajectories, exhibit volatility clustering and have marginal distributions that are equivalent to those of quite general exponential Lévy price models. In particular, we provide explicit details for continuous processes whose marginals equate with the popular VG, CGMY and NIG price models. We also show how to perfectly sample the marginal distributions of more general classes of convoluted subordinators when $L$ is in a sub-class of generalized gamma convolutions, which is relevant for pricing of European style options.

math.PR↗