SearcharxivSearch

arXiv subjects

Peter McCullagh

Publications and source records attributed to Peter McCullagh.

At least 19 recordsLinked to original sources

Non-standard boundary behaviour in two-component mixture models

Consider a binary mixture model of the form $F_\theta = (1-\theta)F_0 + \theta F_1$, where $F_0$ is standard Gaussian and $F_1$ is a completely specified heavy-tailed distribution with the same support. For a sample of $n$ independent and identically distributed values $X_i \sim F_\theta$, the maximum likelihood estimator $\hat\theta_n$ is asymptotically normal provided that $0 < \theta < 1$ is an interior point. This paper investigates the large-sample behaviour for boundary points, which is entirely different and strikingly asymmetric for $\theta=0$ and $\theta=1$. The reason for the asymmetry has to do with typical choices such that $F_0$ is an extreme boundary point and $F_1$ is usually not extreme. On the right boundary, well known results on boundary parameter problems are recovered, giving $\lim \mathbb{P}_1(\hat\theta_n < 1)=1/2$. On the left boundary, $\lim\mathbb{P}_0(\hat\theta_n > 0)=1-1/\alpha$, where $1\leq \alpha \leq 2$ indexes the domain of attraction of the density ratio $f_1(X)/f_0(X)$ when $X\sim F_0$. For $\alpha=1$, which is the most important case in practice, we show how the tail behaviour of $F_1$ governs the rate at which $\mathbb{P}_0(\hat\theta_n > 0)$ tends to zero. A new limit theorem for the joint distribution of the sample maximum and sample mean conditional on positivity establishes multiple inferential anomalies. Most notably, given $\hat\theta_n > 0$, the likelihood ratio statistic has a conditional null limit distribution $G\neq\chi^2_1$ determined by the joint limit theorem. We show through this route that no advantage is gained by extending the single distribution $F_1$ to the nonparametric composite mixture generated by the same tail-equivalence class.

math.ST

Interpretation of local false discovery rates under the zero assumption

In large-scale studies with parallel signal-plus-noise observations, the local false discovery rate is a summary statistic that is often presumed to be equal to the posterior probability that the signal is null. We prefer to call the latter quantity the local null-signal rate to emphasize our view that a null signal and a false discovery are not identical events. The local null-signal rate is commonly estimated through empirical Bayes procedures that build on the `zero density assumption,' which attributes the density of observations near zero entirely to null signals. In this paper, we argue that this strategy does not furnish estimates of the local null-signal rate, but instead of a quantity we call the complementary local activity rate (clar). Although it is likely to be small, an inactive signal is not necessarily zero. The clar dominates both the local null-signal rate and the local false sign rate and is a weakly continuous functional of the signal distribution. As a consequence, it takes on sensible values when the signal is sparse but not exactly zero. Our findings clarify the interpretation of local false discovery rates estimated under the zero density assumption.

math.ST

Sparse-limit approximation for t-statistics

In a range of genomic applications, it is of interest to quantify the evidence that the signal at site~$i$ is active given conditionally independent replicate observations summarized by the sample mean and variance $(\bar Y, s^2)$ at each site. We study the version of the problem in which the signal distribution is sparse, and the error distribution has an unknown site-specific variance so that the null distribution of the standardized statistic is Student-$t$ rather than Gaussian. The main contribution of this paper is a sparse-mixture approximation to the non-null density of the $t$-ratio. This formula demonstrates the effect of low degrees of freedom on the Bayes factor, or the conditional probability that the site is active. We illustrate some differences on a HIV dataset for gene-expression data previously analyzed by Efron (2012).

math.ST

Permanental Graphs

The two components for infinite exchangeability of a sequence of distributions $(P_n)$ are (i) consistency, and (ii) finite exchangeability for each $n$. A consequence of the Aldous-Hoover theorem is that any node-exchangeable, subselection-consistent sequence of distributions that describes a randomly evolving network yields a sequence of random graphs whose expected number of edges grows quadratically in the number of nodes. In this note, another notion of consistency is considered, namely, delete-and-repair consistency; it is motivated by the sense in which infinitely exchangeable permutations defined by the Chinese restaurant process (CRP) are consistent. A goal is to exploit delete-and-repair consistency to obtain a nontrivial sequence of distributions on graphs $(P_n)$ that is sparse, exchangeable, and consistent with respect to delete-and-repair, a well known example being the Ewens permutations \cite{tavare}. A generalization of the CRP$(\alpha)$ as a distribution on a directed graph using the $\alpha$-weighted permanent is presented along with the corresponding normalization constant and degree distribution; it is dubbed the Permanental Graph Model (PGM). A negative result is obtained: no setting of parameters in the PGM allows for a consistent sequence $(P_n)$ in the sense of either subselection or delete-and-repair.

math.ST

A likelihood analysis of quantile-matching transformations

Quantile matching is a strictly monotone transformation that sends the observed response values $\{y_1, . . . , y_n\}$ to the quantiles of a given target distribution. A likelihood based criterion is developed for comparing one target distribution with another in a linear-model setting.

stat.ME

Statistical sparsity

The main contribution of this paper is a mathematical definition of statistical sparsity, which is expressed as a limiting property of a sequence of probability distributions. The limit is characterized by an exceedance measure~$H$ and a rate parameter~$\rho > 0$, both of which are unrelated to sample size. The definition is sufficient to encompass all sparsity models that have been suggested in the signal-detection literature. Sparsity implies that $\rho$~is small, and a sparse approximation is asymptotic in the rate parameter, typically with error $o(\rho)$ in the sparse limit $\rho \to 0$. To first order in sparsity, the sparse signal plus Gaussian noise convolution depends on the signal distribution only through its rate parameter and exceedance measure. This is one of several asymptotic approximations implied by the definition, each of which is most conveniently expressed in terms of the zeta-transformation of the exceedance measure. One implication is that two sparse families having the same exceedance measure are inferentially equivalent, and cannot be distinguished to first order. A converse implication for methodological strategy is that it may be more fruitful to focus on the exceedance measure, ignoring aspects of the signal distribution that have negligible effect on observables and on inferences. From this point of view, scale models and inverse-power measures seem particularly attractive.

stat.ME

Vital variables and survival processes

The focus of a survival study is partly on the distribution of survival times, and partly on the health or quality of life of patients while they live. Health varies over time, and survival is the most basic aspect of health, so the two aspects are closely intertwined. Depending on the nature of the study, a range of variables may be measured; some constant in time, others not; some regarded as responses, others as explanatory risk factors; some directly and personally health-related, others less directly so. This paper begins by classifying variables that may arise in such a setting, emphasizing in particular, the mathematical distinction between vital and non-vital variables. We examine also various types of probabilistic relationships that may exist among variables. Independent evolution is an asymmetric relation, which is intended to encapsulate the notion of one process driving the other; $X$~is a driver of~$Y$ if $X$ evolves independently of the history of~$Y$. This concept arises in several places in the study of survival processes.

stat.ME

M\"obius transformation and a Cauchy family on the sphere

We present some properties of a Cauchy family of distributions on the sphere, which is a spherical extension of the wrapped Cauchy family on the circle. The spherical Cauchy family is closed under the M\"obius transformation on the sphere and there is a similar induced transformation on the parameter space. Stereographic projection transforms the the spherical Cauchy family into a multivariate $t$-family with a certain degree of freedom on Euclidean space. Many tractable properties of the spherical Cauchy are derived using the M\"obius transformation and stereographic projection. A method of moments estimator and an asymptotically efficient estimator are expressed in closed form. The maximum likelihood estimation is also straightforward.

math.ST

The pilgrim process

Pilgrim's monopoly is a probabilistic process giving rise to a non-negative sequence $T_1, T_2,\ldots$ that is infinitely exchangeable, a natural model for time-to-event data. The one-dimensional marginal distributions are exponential. The rules are simple, the process is easy to generate sequentially, and a simple expression is available for both the joint density and the multivariate survivor function. There is a close connection with the Kaplan-Meier estimator of the survival distribution. Embedded within the process is an infinitely exchangeable ordered partition processes connected to Markov branching processes in neutral evolutionary theory. Some aspects of the process, such as the distribution of the number of blocks, can be investigated analytically and confirmed by simulation. By ignoring the order, the embedded process can be considered as an infinitely exchangeable partition process, shown to be closely related to the Chinese restaurant process. Further connection to the Indian buffet process is also provided. Thus we establish a previously unknown link between the well-known Kaplan-Meier estimator and the important Ewens sampling formula.

math.ST

Weak continuity of predictive distribution for Markov survival processes

We explore the concept of a consistent exchangeable survival process - a joint distribution of survival times in which the risk set evolves as a continuous-time Markov process with homogeneous transition rates. We show a correspondence with the de Finetti approach of constructing an exchangeable survival process by generating iid survival times conditional on a completely independent hazard measure. We describe several specific processes, showing how the number of blocks of tied failure times grows asymptotically with the number of individuals in each case. In particular, we show that the set of Markov survival processes with weakly continuous predictive distributions can be characterized by a two-dimensional family called the harmonic process. We end by applying these methods to data, showing how they can be easily extended to handle censoring.

math.ST

A characterization of a Cauchy family on the complex space

It is shown that a family of distributions on the complex space is characterized as the only family such that the orbit of one distribution under a certain group of transformations on the complex space is the same as that under the group of affine transformations. The resulting family is compared with some existing families.

math.ST

Survival models and health sequences

Medical investigations focusing on patient survival often generate not only a failure time for each patient but also a sequence of measurements on patient health at annual or semi-annual check-ups while the patient remains alive. Such a sequence of random length accompanied by a survival time is called a survival process. Ordinarily robust health is associated with longer survival, so the two parts of a survival process cannot be assumed independent. This paper is concerned with a general technique---time reversal---for constructing statistical models for survival processes. A revival model is a regression model in the sense that it incorporates covariate and treatment effects into both the distribution of survival times and the joint distribution of health outcomes. It also allows individual health outcomes to be used clinically for predicting the subsequent survival time.

stat.ME

On Bayes' theorem for improper mixtures

Although Bayes's theorem demands a prior that is a probability distribution on the parameter space, the calculus associated with Bayes's theorem sometimes generates sensible procedures from improper priors, Pitman's estimator being a good example. However, improper priors may also lead to Bayes procedures that are paradoxical or otherwise unsatisfactory, prompting some authors to insist that all priors be proper. This paper begins with the observation that an improper measure on Theta satisfying Kingman's countability condition is in fact a probability distribution on the power set. We show how to extend a model in such a way that the extended parameter space is the power set. Under an additional finiteness condition, which is needed for the existence of a sampling region, the conditions for Bayes's theorem are satisfied by the extension. Lack of interference ensures that the posterior distribution in the extended space is compatible with the original parameter space. Provided that the key finiteness condition is satisfied, this probabilistic analysis of the extended model may be interpreted as a vindication of improper Bayes procedures derived from the original model.

math.ST

Reversible Markov structures on divisible set partitions

We study $k$-divisible partition structures, which are families of random set partitions whose block sizes are divisible by an integer $k=1,2,\ldots$. In this setting, exchangeability corresponds to the usual invariance under relabeling by arbitrary permutations; however, for $k>1$, the ordinary deletion maps on partitions no longer preserve divisibility, and so a random deletion procedure is needed to obtain a partition structure. We describe explicit Chinese restaurant-type seating rules for generating families of exchangeable $k$-divisible partitions that are consistent under random deletion. We further introduce the notion of {\em Markovian partition structures}, which are ensembles of exchangeable Markov chains on $k$-divisible partitions that are consistent under a random process of {\em Markovian deletion}. The Markov chains we study are reversible and refine the class of Markov chains introduced in {\em J.\ Appl.\ Probab.}~{\bf48}(3):778--791.

math.ST

Classification based on a permanental process with cyclic approximation

We introduce a doubly stochastic marked point process model for supervised classification problems. Regardless of the number of classes or the dimension of the feature space, the model requires only 2--3 parameters for the covariance function. The classification criterion involves a permanental ratio for which an approximation using a polynomial-time cyclic expansion is proposed. The approximation is effective even if the feature region occupied by one class is a patchwork interlaced with regions occupied by other classes. An application to DNA microarray analysis indicates that the cyclic approximation is effective even for high-dimensional data. It can employ feature variables in an efficient way to reduce the prediction error significantly. This is critical when the true classification relies on non-reducible high-dimensional features.

stat.ME

Gibbs fragmentation trees

We study fragmentation trees of Gibbs type. In the binary case, we identify the most general Gibbs-type fragmentation tree with Aldous' beta-splitting model, which has an extended parameter range $β>-2$ with respect to the ${\rm beta}(β+1,β+1)$ probability distributions on which it is based. In the multifurcating case, we show that Gibbs fragmentation trees are associated with the two-parameter Poisson--Dirichlet models for exchangeable random partitions of $\mathbb {N}$, with an extended parameter range $0\leα\le1$, $θ\ge-2α$ and $α<0$, $θ=-mα$, $m\in \mathbb {N}$.

math.PR

Marginal likelihood for parallel series

Suppose that $k$ series, all having the same autocorrelation function, are observed in parallel at $n$ points in time or space. From a single series of moderate length, the autocorrelation parameter $β$ can be estimated with limited accuracy, so we aim to increase the information by formulating a suitable model for the joint distribution of all series. Three Gaussian models of increasing complexity are considered, two of which assume that the series are independent. This paper studies the rate at which the information for $β$ accumulates as $k$ increases, possibly even beyond $n$. The profile log likelihood for the model with $k(k+1)/2$ covariance parameters behaves anomalously in two respects. On the one hand, it is a log likelihood, so the derivatives satisfy the Bartlett identities. On the other hand, the Fisher information for $β$ increases to a maximum at $k=n/2$, decreasing to zero for $k\ge n$. In any parametric statistical model, one expects the Fisher information to increase with additional data; decreasing Fisher information is an anomaly demanding an explanation.

math.ST