SearcharxivSearch

arXiv subjects

C. E. Priebe

Publications and source records attributed to C. E. Priebe.

5 recordsLinked to original sources

Vertex nomination schemes for membership prediction

Suppose that a graph is realized from a stochastic block model where one of the blocks is of interest, but many or all of the vertices' block labels are unobserved. The task is to order the vertices with unobserved block labels into a ``nomination list'' such that, with high probability, vertices from the interesting block are concentrated near the list's beginning. We propose several vertex nomination schemes. Our basic - but principled - setting and development yields a best nomination scheme (which is a Bayes-Optimal analogue), and also a likelihood maximization nomination scheme that is practical to implement when there are a thousand vertices, and which is empirically near-optimal when the number of vertices is small enough to allow comparison to the best nomination scheme. We then illustrate the robustness of the likelihood maximization nomination scheme to the modeling challenges inherent in real data, using examples which include a social network involving human trafficking, the Enron Graph, a worm brain connectome and a political blog network.

stat.ML

A New Family of Random Graphs for Testing Spatial Segregation

We discuss a graph-based approach for testing spatial point patterns. This approach falls under the category of data-random graphs, which have been introduced and used for statistical pattern recognition in recent years. Our goal is to test complete spatial randomness against segregation and association between two or more classes of points. To attain this goal, we use a particular type of parametrized random digraph called proximity catch digraph (PCD) which is based based on relative positions of the data points from various classes. The statistic we employ is the relative density of the PCD. When scaled properly, the relative density of the PCD is a $U$-statistic. We derive the asymptotic distribution of the relative density, using the standard central limit theory of $U$-statistics. The finite sample performance of the test statistic is evaluated by Monte Carlo simulations, and the asymptotic performance is assessed via Pitman's asymptotic efficiency, thereby yielding the optimal parameters for testing. Furthermore, the methodology discussed in this article is also valid for data in multiple dimensions.

stat.ME

On the Distribution of the Domination Number of a New Family of Parametrized Random Digraphs

We derive the asymptotic distribution of the domination number of a new family of random digraph called proximity catch digraph (PCD), which has application to statistical testing of spatial point patterns and to pattern recognition. The PCD we use is a parametrized digraph based on two sets of points on the plane, where sample size and locations of the elements of one is held fixed, while the sample size of the other whose elements are randomly distributed over a region of interest goes to infinity. PCDs are constructed based on the relative allocation of the random set of points with respect to the Delaunay triangulation of the other set whose size and locations are fixed. We introduce various auxiliary tools and concepts for the derivation of the asymptotic distribution. We investigate these concepts in one Delaunay triangle on the plane, and then extend them to the multiple triangle case. The methods are illustrated for planar data, but are applicable in higher dimensions also.

math.CO

Relative Density of the Random $r$-Factor Proximity Catch Digraph for Testing Spatial Patterns of Segregation and Association

Statistical pattern classification methods based on data-random graphs were introduced recently. In this approach, a random directed graph is constructed from the data using the relative positions of the data points from various classes. Different random graphs result from different definitions of the proximity region associated with each data point and different graph statistics can be employed for data reduction. The approach used in this article is based on a parameterized family of proximity maps determining an associated family of data-random digraphs. The relative arc density of the digraph is used as the summary statistic, providing an alternative to the domination number employed previously. An important advantage of the relative arc density is that, properly re-scaled, it is a $U$-statistic, facilitating analytic study of its asymptotic distribution using standard $U$-statistic central limit theory. The approach is illustrated with an application to the testing of spatial patterns of segregation and association. Knowledge of the asymptotic distribution allows evaluation of the Pitman and Hodges-Lehmann asymptotic efficacies, and selection of the proximity map parameter to optimize efficiency. Furthermore the approach presented here also has the advantage of validity for data in any dimension.

stat.ME

The Use of Domination Number of a Random Proximity Catch Digraph for Testing Spatial Patterns of Segregation and Association

Priebe et al. (2001) introduced the class cover catch digraphs and computed the distribution of the domination number of such digraphs for one dimensional data. In higher dimensions these calculations are extremely difficult due to the geometry of the proximity regions; and only upper-bounds are available. In this article, we introduce a new type of data-random proximity map and the associated (di)graph in $\mathbb R^d$. We find the asymptotic distribution of the domination number and use it for testing spatial point patterns of segregation and association.

stat.ME