SearcharxivSearch

arXiv subjects

Steven P. Ellis

Publications and source records attributed to Steven P. Ellis.

7 recordsLinked to original sources

Statistical Methods for the meta-analysis paper by Itzhaky et al

This document describes the statistical methods used in Itzhaky et al ("Systematic Review and Meta-analysis: Twenty-six Years of Randomized Clinical Trials of Psychosocial Interventions to Reduce Suicide Risk in Adolescents"). That paper is a meta-analysis of randomized controlled clinical trials testing methods for preventing suicidal behavior and/or ideation in youth. Particularly on the behavior side the meta-data are challenging to analyze. This paper has two parts. The first is an informal discussion of the statistical methods used. The second gives detailed mathematical derivations of some formulas and methods.

stat.AP

Concurrence Topology of Some Cancer Genomics Data

The topological data analysis method "concurrence topology" is applied to mutation frequencies in 69 genes in glioblastoma data. In dimension 1 some apparent "mutual exclusivity" is found. By simulation of data having approximately the same second order dependence structure as that found in the data, it appears that one triple of mutations, PTEN, RB1, TP53, exhibits mutual exclusivity that depends on special features of the third order dependence and may reflect global dependence among a larger group of genes. A bootstrap analysis suggests that this form of mutual exclusivity is not uncommon in the population from which the data were drawn.

stat.AP

Joining and Independence in Concurrence Topology

"Concurrence topology" (Ellis and Klein \emph{Homology, Homotopy, and Applications,} \textbf{16}) is a TDA method for binary data. The idea is to construct a filtration consisting of Dowker complexes then compute persistent homology. Persistent classes correspond to a form of negative statistical association among the variables. Suppose we have two groups of binary variables each displaying negative association, manifested in nontrivial concurrence homology in dimensions $p$ and in one group and $q$ in the other \emph{when the groups of variables are considered individually.} Suppose, however, that the two \emph{groups} of variables are statistically independent of each other. Now combine the two groups of variables and suppose the sample size is large. Then representative cycles, one from each group of variables, will combine to produce a cycle in dimension $p+q+1$. This is a chain level phenomenon, but we show it has a signature in homology. Looking for this signature can be used to study the dependence among groups of variables.

math.ST

Singularity of Data Analytic Operations

Statistical data by their very nature are indeterminate in the sense that if one repeats the process of collecting the data the new data set will be different from the original. But two data sets generated in the same way should ``tell the same story''. Therefore, a statistical method, a map $\Phi$ taking a data set $x$ to a point in some space $\mathsf{F}$, should be stable at $x$: Small perturbations in $x$ should result in a small change in $\Phi(x)$. Otherwise, $\Phi$ is useless at $x$ or -- and this is important -- near $x$. So one doesn't want $\Phi$ to have "singularities," data sets $x$ such that the the limit of $\Phi(y)$ as $y$ approaches $x$ doesn't exist. (The same issue arises elsewhere in applied math.) We prove that broad classes of statistical methods have topological obstructions to continuity: They must have singularities. We derive broadly applicable lower bounds on the Hausdorff dimension, even Hausdorff measure, of the set of singularities of data maps. General results concerning severity of singularities are proved. For illustration, we show our results apply to plane fitting, measuring location of data on spheres, and to linear classification. This is not a "final" version, merely another attempt.

math.ST

Describing High-Order Statistical Dependence Using "Concurrence Topology", with Application to Functional MRI Brain Data

For multivariate data, dependence beyond pair-wise can be important. This is true, for example, in using functional MRI (fMRI) data to investigate brain functional connectivity. When one has more than a few variables, however, the number of simple summaries of even third-order dependence can be unmanageably large. "Concurrence topology" is an apparently new nonparametric method for describing high-order dependence among up to dozens of dichotomous variables (e.g., seventh-order dependence in 32 variables). This method generally produces summaries of $p^{th}$-order dependence of manageable size no matter how big $p$ is. (But computing time can be lengthy.) For time series, this method can be applied in both the time and Fourier domains. Write each observation as a vector of 0's and 1's. A "concurrence" is a group of variables all "1" in the same observation. The collection of concurrences can be represented as a sequence of shapes ("filtration"). Holes in the filtration indicate weak or negative association among the variables. The pattern of the holes in the filtration can be analyzed using computational topology. This method is demonstrated on dichotomized fMRI data. The dataset includes subjects diagnosed with ADHD and healthy controls. In an exploratory analysis numerous group differences in the topology of the filtrations are found.

stat.ME

On the Approximation of a Function Continuous off a Closed Set by One Continuous Off a Polyhedron

Let $P$ be a finite simplicial comple with underlying space (union of simplices in $P$) $|P|$. Let $Q$ be a subcomplex of $P$. Let $a \geq 0$. Then there exists $K < \infty$, \emph{depending only on $a$ and $Q$,} with the following property. Let $\mathcal{S} \subset |P|$ be closed and suppose $Φ$ is a continuous map of $|P| \setminus \mathcal{S}$ into some topological space $\mathcal{F}$. Suppose $\dim (\tilde{\mathcal{S}} \cap |Q|) \leq a$, where "$\dim$" = Hausdorff dimension. Then there exists $\tilde{\mathcal{S}} \subset |P|$ such that $\tilde{\mathcal{S}} \cap |Q|$ is the underlying space of a subcomplex of $Q$ and there is a continuous map $\tildeΦ$ of $|P| \setminus \tilde{\mathcal{S}}$ into $\mathcal{F}$ such that $\mathcal{H}^{a} \bigl(\tilde{\mathcal{S}} \cap |Q| \bigr) \leq K \mathcal{H}^{a} \bigl(\mathcal{S} \cap |Q| \bigr)$, where $\mathcal{H}^{a}$ denotes $a$-dimensional Hausdorff measure; if $x \in \tilde{\mathcal{S}}$ then $x$ belongs to a simplex in $P$ intersecting $\mathcal{S}$; if $x \in |P| \setminus \mathcal{S}$, $x \in σ\in P$, and $σ$ does not intersect any simplex in $Q$ whose simplicial interior intersects $\mathcal{S}$, then $\tildeΦ(x)$ is defined and equals $= Φ(x)$; if $σ\in P$ then $\tildeΦ(σ\setminus \tilde{\mathcal{S}}) \subset Φ(σ\setminus \mathcal{S})$; and if $\mathcal{F}$ is a metric space and $Φ$ is locally Lipschitz on $|P| \setminus \mathcal{S}$ then $\tildeΦ$ is locally Lipschitz on $|P| \setminus \tilde{\mathcal{S}}$ Moreover, $P$ can be replaced by an arbitrarily fine subdivision without changing $K$.

math.GN

An Algorithm for Unconstrained Quadratically Penalized Convex Optimization

A descent algorithm, "Quasi-Quadratic Minimization with Memory" (QQMM), is proposed for unconstrained minimization of the sum, $F$, of a non-negative convex function, $V$, and a quadratic form. Such problems come up in regularized estimation in machine learning and statistics. In addition to values of $F$, QQMM requires the (sub)gradient of $V$. Two features of QQMM help keep low the number of evaluations of the objective function it needs. First, QQMM provides good control over stopping the iterative search. This feature makes QQMM well adapted to statistical problems because in such problems the objective function is based on random data and therefore stopping early is sensible. Secondly, QQMM uses a complex method for determining trial minimizers of $F$. After a description of the problem and algorithm a simulation study comparing QQMM to the popular BFGS optimization algorithm is described. The simulation study and other experiments suggest that QQMM is generally substantially faster than BFGS in the problem domain for which it was designed. A QQMM-BFGS hybrid is also generally substantially faster than BFGS but does better than QQMM when QQMM is very slow.

stat.CO