SearcharxivSearch

arXiv subjects

Jesper Møller

Publications and source records attributed to Jesper Møller.

At least 19 recordsLinked to original sources

Sufficient digits and density estimation: A Bayesian nonparametric approach using generalized finite Pólya trees

This paper proposes a novel approach for statistical modelling of a continuous random variable $X$ on $[0, 1)$, based on its digit representation $X=.X_1X_2\ldots$. In general, $X$ can be coupled with a latent random variable $N$ so that $(X_1,\ldots,X_N)$ becomes a sufficient statistics and $.X_{N+1}X_{N+2}\ldots$ is uniformly distributed. In line with this fact, and focusing on binary digits for simplicity, we propose a family of generalized finite P{ó}lya trees that induces a random density for a sample, which becomes a flexible tool for density estimation. Here, the digit system may be random and learned from the data. We provide a detailed Bayesian analysis, including closed form expression for the posterior distribution. We analyse the frequentist properties as the sample size increases, and provide sufficient conditions for consistency of the posterior distributions of the random density and $N$. We consider an extension to data spanning multiple orders of magnitude, and propose a prior distribution that encodes the so-called extended Newcomb-Benford law. Such a model shows promising results for density estimation of human-activity data. Our methodology is illustrated on several synthetic and real datasets.

stat.ME

The asymptotic distribution of the scaled remainder for pseudo golden ratio expansions of a continuous random variable

Let $X=\sum_{k=1}^\infty X_k β^{-k}$ be the base-$β$ expansion of a continuous random variable $X$ on the unit interval where $β$ is the positive solution to $β^n = 1 + β+ \cdots + β^{n-1}$ for an integer $n\ge 2$ (i.e., $β$ is a generalization of the golden mean for which $n=2$). We study the asymptotic distribution and convergence rate of the scaled remainder $\sum_{k=1}^\infty X_{m+k} β^{-k}$ when $m$ tends to infinity.

math.PR

Coupling results and Markovian structures for number representations of continuous random variables

A general setting for nested subdivisions of a bounded real set into intervals defining the digits $X_1,X_2,...$ of a random variable $X$ with a probability density function $f$ is considered. Under the weak condition that $f$ is almost everywhere lower semi-continuous, a coupling between $X$ and a non-negative integer-valued random variable $N$ is established so that $X_1,...,X_N$ have an interpretation as the ``sufficient digits'', since the distribution of $R=(X_{N+1},X_{N+2},...)$ conditioned on $S=(X_1,...,X_N)$ does not depend on $f$. Adding a condition about a Markovian structure of the lengths of the intervals in the nested subdivisions, $R\,|\,S$ becomes a Markov chain of a certain order $s\ge0$. If $s=0$ then $X_{N+1},X_{N+2},...$ are IID with a known distribution. When $s>0$ and the Markov chain is uniformly geometric ergodic, a coupling is established between $(X,N)$ and a random time $M$ so that the chain after time $\max\{N,s\}+M-s$ is stationary and $M$ follows a simple known distribution. The results are related to several examples of number representations generated by a dynamical system, including base-$q$ expansions, generalized Lüroth series, $β$-expansions, and continued fraction representations. The importance of the results and some suggestions and open problems for future research are discussed.

math.PR

Cox processes driven by transformed Gaussian processes on linear networks -- A review and new contributions

There is a lack of point process models on linear networks. For an arbitrary linear network, we consider new models for a Cox process with an isotropic pair correlation function obtained in various ways by transforming an isotropic Gaussian process which is used for driving the random intensity function of the Cox process. In particular we introduce three model classes given by log Gaussian, interrupted, and permanental Cox processes on linear networks, and consider for the first time statistical procedures and applications for parametric families of such models. Moreover, we construct new simulation algorithms for Gaussian processes on linear networks and discuss whether the geodesic metric or the resistance metric should be used for the kind of Cox processes studied in this paper.

math.ST

How many digits are needed?

Let $X_1,X_2,...$ be the digits in the base-$q$ expansion of a random variable $X$ defined on $[0,1)$ where $q\ge2$ is an integer. For $n=1,2,...$, we study the probability distribution $P_n$ of the (scaled) remainder $T^n(X)=\sum_{k=n+1}^\infty X_k q^{n-k}$: If $X$ has an absolutely continuous CDF then $P_n$ converges in the total variation metric to the Lebesgue measure $μ$ on the unit interval. Under weak smoothness conditions we establish first a coupling between $X$ and a non-negative integer valued random variable $N$ so that $T^N(X)$ follows $μ$ and is independent of $(X_1,...,X_N)$, and second exponentially fast convergence of $P_n$ and its PDF $f_n$. We discuss how many digits are needed and show examples of our results. The convergence results are extended to the case of a multivariate random variable defined on a unit cube.

math.PR

Realizability and tameness of fusion systems

A saturated fusion system over a finite $p$-group $S$ is a category whose objects are the subgroups of $S$ and whose morphisms are injective homomorphisms between the subgroups satisfying certain axioms. A fusion system over $S$ is realized by a finite group $G$ if $S$ is a Sylow $p$-subgroup of $G$ and morphisms in the category are those induced by conjugation in $G$. One recurrent question in this subject is to find criteria as to whether a given saturated fusion system is realizable or not. One main result in this paper is that a saturated fusion system is realizable if all of its components (in the sense of Aschbacher) are realizable. Another result is that all realizable fusion systems are tame: a finer condition on realizable fusion systems that involves describing automorphisms of a fusion system in terms of those of some group that realizes it. Stated in this way, these results depend on the classification of finite simple groups, but we also give more precise formulations whose proof is independent of the classification.

math.GR

Singular distribution functions for random variables with stationary digits

Let $F$ be the cumulative distribution function (CDF) of the base-$q$ expansion $\sum_{n=1}^\infty X_n q^{-n}$, where $q\ge2$ is an integer and $\{X_n\}_{n\geq 1}$ is a stationary stochastic process with state space $\{0,\ldots,q-1\}$. In a previous paper we characterized the absolutely continuous and the discrete components of $F$. In this paper we study special cases of models, including stationary Markov chains of any order and stationary renewal point processes, where we establish a law of pure types: $F$ is then either a uniform or a singular CDF on $[0,1]$. Moreover, we study mixtures of such models. In most cases expressions and plots of $F$ are given.

math.PR

Characterization of random variables with stationary digits

Let $q\ge2$ be an integer, $\{X_n\}_{n\geq 1}$ a stochastic process with state space $\{0,\ldots,q-1\}$, and $F$ the cumulative distribution function (CDF) of $\sum_{n=1}^\infty X_n q^{-n}$. We show that stationarity of $\{X_n\}_{n\geq 1}$ is equivalent to a functional equation obeyed by $F$ and use this to characterize the characteristic function of $X$ and the structure of $F$ in terms of its Lebesgue decomposition. More precisely, while the absolutely continuous component of $F$ can only be the uniform distribution on the unit interval, its discrete component can only be a countable convex combination of certain explicitly computable CDFs for probability distributions with finite support. We also show that $\mathrm{d} F$ is a Rajchman measure if and only if $F $ is the uniform CDF on $[0,1]$.

math.PR

Determinantal shot noise Cox processes

We present a new class of cluster point process models, which we call determinantal shot noise Cox processes (DSNCP), with repulsion between cluster centres. They are the special case of generalized shot noise Cox processes where the cluster centres are determinantal point processes. We establish various moment results and describe how these can be used to easily estimate unknown parameters in two particularly tractable cases, namely when the offspring density is isotropic Gaussian and the kernel of the determinantal point process of cluster centres is Gaussian or like in a scaled Ginibre point process. Through a simulation study and the analysis of a real point pattern data set we see that when modelling clustered point patterns, a much lower intensity of cluster centres may be needed in DSNCP models as compared to shot noise Cox processes.

stat.ME

Fitting three-dimensional Laguerre tessellations by hierarchical marked point process models

We present a general statistical methodology for analysing a Laguerre tessellation data set viewed as a realization of a marked point process model. In the first step, for the points we use a nested sequence of multiscale processes which constitute a flexible parametric class of pairwise interaction point process models. In the second step, for the marks/radii conditioned on the points we consider various exponential family models where the canonical sufficient statistic is based on tessellation characteristics. For each step parameter estimation based on maximum pseudolikelihood methods is tractable. Model checking is performed using global envelopes and corresponding tests in the first step and by comparing observed and simulated tessellation characteristics in the second step. We apply our methodology for a 3D Laguerre tessellation data set representing the microstructure of a polycrystalline metallic material, where simulations under a fitted model may substitute expensive laboratory experiments.

stat.ME

Should we condition on the number of points when modelling spatial point patterns?

We discuss the practice of directly or indirectly assuming a model for the number of points when modelling spatial point patterns even though it is rarely possible to validate such a model in practice since most point pattern data consist of only one pattern. We therefore explore the possibility to condition on the number of points instead when fitting and validating spatial point process models. In a simulation study with different popular spatial point process models, we consider model validation using global envelope tests based on functional summary statistics. We find that conditioning on the number of points will for some functional summary statistics lead to more narrow envelopes and that it can also be useful for correcting for some conservativeness in the tests when testing composite hypothesis. However, for other functional summary statistics, it makes little or no difference to condition on the number of points. When estimating parameters in popular spatial point process models, we conclude that for mathematical and computational reasons it is convenient to assume a distribution for the number of points.

stat.ME

MCMC computations for Bayesian mixture models using repulsive point processes

Repulsive mixture models have recently gained popularity for Bayesian cluster detection. Compared to more traditional mixture models, repulsive mixture models produce a smaller number of well separated clusters. The most commonly used methods for posterior inference either require to fix a priori the number of components or are based on reversible jump MCMC computation. We present a general framework for mixture models, when the prior of the `cluster centres' is a finite repulsive point process depending on a hyperparameter, specified by a density which may depend on an intractable normalizing constant. By investigating the posterior characterization of this class of mixture models, we derive a MCMC algorithm which avoids the well-known difficulties associated to reversible jump MCMC computation. In particular, we use an ancillary variable method, which eliminates the problem of having intractable normalizing constants in the Hastings ratio. The ancillary variable method relies on a perfect simulation algorithm, and we demonstrate this is fast because the number of components is typically small. In several simulation studies and an application on sociological data, we illustrate the advantage of our new methodology over existing methods, and we compare the use of a determinantal or a repulsive Gibbs point process prior model.

stat.ME

Approximate Bayesian inference for a spatial point process model exhibiting regularity and random aggregation

In this paper, we propose a doubly stochastic spatial point process model with both aggregation and repulsion. This model combines the ideas behind Strauss processes and log Gaussian Cox processes. The likelihood for this model is not expressible in closed form but it is easy to simulate realisations under the model. We therefore explain how to use approximate Bayesian computation (ABC) to carry out statistical inference for this model. We suggest a method for model validation based on posterior predictions and global envelopes. We illustrate the ABC procedure and model validation approach using both simulated point patterns and a real data example.

stat.ME

Modelling columnarity of pyramidal cells in the human cerebral cortex

For modelling the location of pyramidal cells in the human cerebral cortex we suggest a hierarchical point process in $\mathbb{R}^3$ that exhibits anisotropy in the form of cylinders extending along the $z$-axis. The model consists first of a generalised shot noise Cox process for the $xy$-coordinates, providing cylindrical clusters, and next of a Markov random field model for the $z$-coordinates conditioned on the $xy$-coordinates, providing either repulsion, aggregation, or both within specified areas of interaction. Several cases of these hierarchical point processes are fitted to two pyramidal cell datasets, and of these a final model allowing for both repulsion and attraction between the points seem adequate. We discuss how the final model relates to the so-called minicolumn hypothesis in neuroscience.

stat.ME

Modelling spine locations on dendrite trees using inhomogeneous Cox point processes

Dendritic spines, which are small protrusions on the dendrites of a neuron, are of interest in neuroscience as they are related to cognitive processes such as learning and memory. We analyse the distribution of spine locations on six different dendrite trees from mouse neurons using point process theory for linear networks. Besides some possible small-scale repulsion, { we find that two of the spine point pattern data sets may be described by inhomogeneous Poisson process models}, while the other point pattern data sets exhibit clustering between spines at a larger scale. To model this we propose an inhomogeneous Cox process model constructed by thinning a Poisson process on a linear network with retention probabilities determined by a spatially correlated random field. For model checking we consider network analogues of the empirical $F$-, $G$-, and $J$-functions originally introduced for inhomogeneous point processes on a Euclidean space. The fitted Cox process models seem to catch the clustering of spine locations between spines, but also posses a large variance in the number of points for some of the data sets causing large confidence regions for the empirical $F$- and $G$-functions.

stat.ME

Couplings for determinantal point processes and their reduced Palm distributions with a view to quantifying repulsiveness

For a determinantal point process $X$ with a kernel $K$ whose spectrum is strictly less than one, Andr{é} Goldman has established a coupling to its reduced Palm process $X^u$ at a point $u$ with $K(u,u)>0$ so that almost surely $X^u$ is obtained by removing a finite number of points from $X$. We sharpen this result, assuming weaker conditions and establishing that $X^u$ can be obtained by removing at most one point from $X$, where we specify the distribution of the difference $ξ_u:=X\setminus X^u$. This is used for discussing the degree of repulsiveness in DPPs in terms of $ξ_u$, including Ginibre point processes and other specific parametric models for DPPs.

math.PR

Globally intensity-reweighted estimators for $K$- and pair correlation functions

We introduce new estimators of the inhomogeneous $K$-function and the pair correlation function of a spatial point process as well as the cross $K$-function and the cross pair correlation function of a bivariate spatial point process under the assumption of second-order intensity-reweighted stationarity. These estimators rely on a 'global' normalization factor which depends on an aggregation of the intensity function, whilst the existing estimators depend 'locally' on the intensity function at the individual observed points. The advantages of our new global estimators over the existing local estimators are demonstrated by theoretical considerations and a simulation study.

stat.ME