SearcharxivSearch

arXiv subjects

Marianna Bolla

Publications and source records attributed to Marianna Bolla.

15 recordsLinked to original sources

The non-backtracking random walk and its usage for vertex clustering

In case of sparse graphs, relation between the real eigenvalues of the non-backtracking matrix and those of the non-backtracking transition probability matrix is considered with respect to vertex clustering. For this purpose, the random walk along the non-backtracking graph is considered, the vertices of which are the bioriented edges, and the adjacency relation depends on whether the random walk goes through the oriented edges with the rule of ``not going back in the next step''. This is encoded in the non-backtracking matrix that is the adjacency matrix of the non-backtracking graph. The structural real eigenvalues of the transition probability matrix are related to the constant multiples of the non-backtracking one, the concordance of which indicates the existence of a sparse stochastic block model behind the graph. ``Inflation--deflation'' techniques are also developed for clustering the vertices of the original graph together with real world applications.

math.CO

Structure and Noise in Dense and Sparse Random Graphs: Percolated Stochastic Block Model via the EM Algorithm and Belief Propagation with Non-Backtracking Spectra

In this survey paper it is illustrated how spectral clustering methods for unweighted graphs are adapted to the dense and sparse regimes. Whereas Laplacian and modularity based spectral clustering is apt to dense graphs, recent results show that for sparse ones, the non-backtracking spectrum is the best candidate to find assortative clusters of nodes. Here belief propagation in the sparse stochastic block model is derived with arbitrarily given model parameters that results in a non-linear system of equations; with linear approximation, the spectrum of the non-backtracking matrix is able to specify the number $k$ of clusters. Then the model parameters themselves can be estimated by the EM algorithm. Bond percolation in the assortative model is considered in the following two senses: the within- and between-cluster edge probabilities decrease with the number of nodes and edges coming into existence in this way are retained with probability $β$. As a consequence, the optimal $k$ is the number of the structural real eigenvalues (greater than $\sqrt{c}$, where $c$ is the average degree) of the non-backtracking matrix of the graph. Assuming, these eigenvalues $μ_1 >\dots > μ_k$ are distinct, the multiple phase transitions obtained for $β$ are $β_i =\frac{c}{μ_i^2}$; further, at $β_i$ the number of detectable clusters is $i$, for $i=1,\dots ,k$. Inflation-deflation techniques are also discussed to classify the nodes themselves, which can be the base of the sparse spectral clustering. Simulation results, as well as real life examples are presented.

math.CO

Causal Vector Autoregression Enhanced with Covariance and Order Selection

A causal vector autoregressive (CVAR) model is introduced for weakly stationary multivariate processes, combining a recursive directed graphical model for the contemporaneous components and a vector autoregressive model longitudinally. Block Cholesky decomposition with varying block sizes is used to solve the model equations and estimate the path coefficients along a directed acyclic graph (DAG). If the DAG is decomposable, i.e. the zeros form a reducible zero pattern (RZP) in its adjacency matrix, then covariance selection is applied that assigns zeros to the corresponding path coefficients. Real life applications are also considered, where for the optimal order $p\ge 1$ of the fitted CVAR$(p)$ model, order selection is performed with various information criteria.

stat.ME

Regularity based spectral clustering and mapping the Fiedler-carpet

Spectral clustering is discussed from many perspectives, by extending it to rectangular arrays and discrepancy minimization too. Near optimal clusters are obtained with singular value decomposition and with the weighted $k$-means algorithm. In case of rectangular arrays, this means enhancing the method of correspondence analysis with clustering, and in case of edge-weighted graphs, a normalized Laplacian based clustering. In the latter case it is proved that a spectral gap between the $(k-1)$th and $k$th smallest positive eigenvalues of the normalized Laplacian matrix gives rise to a sudden decrease of the inner cluster variances when the number of clusters of the vertex representatives is $2^{k-1}$, but only the first $k-1$ eigenvectors, constituting the so-called Fiedler-carpet, are used in the representation. Application to directed migration graphs is also discussed.

math.CO

Generalized quasirandom properties of expanding graph sequencesedding

We consider special multiclass spectral, discrepancy, degree, and codegree properties of expanding graph sequences. As we can prove equivalences and implications between them and the definition of the generalized quasirandomness of Lovász--Sós (2008), they can be regarded as generalized quasirandom properties akin to the equivalent quasirandom properties of the seminal Chung--Graham--Wilson paper (1989) in the one-class scenario. Since these properties are valid for certain deterministic graph sequences, irrespective of stochastic models, the partial implications also justify for law-dimensional embedding of large-scale graphs and for discrepancy minimizing spectral clustering.

math.CO

Estimating parameters of a directed weighted graph model with beta-distributed edge-weights

We introduce a directed, weighted random graph model, where the edge-weights are independent and beta-distributed with parameters depending on their endpoints. We will show that the row- and column-sums of the transformed edge-weight matrix are sufficient statistics for the parameters, and use the theory of exponential families to prove that the ML estimate of the parameters exists and is unique. Then an algorithm to find this estimate is introduced together with convergence proof that uses properties of the digamma function. Simulation results and applications are also presented.

math.ST

Properties of the multiway discrepancy

Some properties of the multiway discrepanc of rectangular matrices of nonnegative entries are discussed. We are able to prove the continuity of this discrepancy, as well as some statements about the multiway discrepancy of some special matrices and graphs. We also conjecture that the k-way discrepancy is monotonic in k.

math.CO

Estimating parameters of a multipartite loglinear graph model via the EM algorithm

We will amalgamate the Rash model (for rectangular binary tables) and the newly introduced $α$-$β$ models (for random undirected graphs) in the framework of a semiparametric probabilistic graph model. Our purpose is to give a partition of the vertices of an observed graph so that the generated subgraphs and bipartite graphs obey these models, where their strongly connected parameters give multiscale evaluation of the vertices at the same time. In this way, a heterogeneous version of the stochastic block model is built via mixtures of loglinear models and the parameters are estimated with a special EM iteration. In the context of social networks, the clusters can be identified with social groups and the parameters with attitudes of people of one group towards people of the other, which attitudes depend on the cluster memberships. The algorithm is applied to randomly generated and real-word data.

stat.ME

Relating multiway discrepancy and singular values of graphs and contingency tables

The $k$-way discrepancy $\disc_k (\C)$ of a rectangular array $\C$ of nonnegative entries is the minimum of the maxima of the within- and between-cluster discrepancies that can be obtained by simultaneous $k$-clusterings (proper partitions) of its rows and columns. In the main theorem, irrespective of the size of $\C$, we give the following estimate for the $k$th largest non-trivial singular value of the normalized table: $s_k \le 9\disc_{k } (\C ) (k+2 -9k\ln \disc_{k } (\C ))$, provided $\disc_{k } (\C ) <1$ and $k\le \rk (\C )$. This statement is the converse of Theorem 7 of Bolla \cite{Bolla14}, and the proof uses some lemmas and ideas of Butler \cite{Butler}, where only the $k=1$ case is treated, in which case our upper bound is the tighter. The result naturally extends to the singular values of the normalized adjacency matrix of a weighted undirected or directed graph.

math.CO

When the largest eigenvalue of the modularity and normalized modularity matrix is zero

In July 2012, at the Conference on Applications of Graph Spectra in Computer Science, Barcelona, D. Stevanovic posed the following open problem: which graphs have the zero as the largest eigenvalue of their modularity matrix? The conjecture was that only the complete and complete multipartite graphs. They indeed have this property, but are they the only ones? In this paper, we will give an affirmative answer to this question and prove a bit more: both the modularity and the normalized modularity matrix of a graph is negative semidefinite if and only if the graph is complete or complete multipartite.

math.SP

Modularity spectra, eigen-subspaces, and structure of weighted graphs

The role of the normalized modularity matrix in finding homogeneous cuts will be presented. We also discuss the testability of the structural eigenvalues and that of the subspace spanned by the corresponding eigenvectors of this matrix. In the presence of a spectral gap between the k-1 largest absolute value eigenvalues and the remainder of the spectrum, this in turn implies the testability of the sum of the inner variances of the k clusters that are obtained by applying the k-means algorithm for the appropriately chosen vertex representatives.

math.ST

SVD, discrepancy, and regular structure of contingency tables

We will use the factors obtained by correspondence analysis to find biclustering of a contingency table such that the row-column cluster pairs are regular, i.e., they have small discrepancy. In our main theorem, the constant of the so-called volume-regularity is related to the SVD of the normalized contingency table. Our result is applicable to two-way cuts when both the rows and columns are divided into the same number of clusters, thus extending partly the result of Butler estimating the discrepancy of a contingency table by the second largest singular value of the normalized table (one-cluster, rectangular case), and partly a former result of the author for estimating the constant of volume-regularity by the structural eigenvalues and the distances of the corresponding eigen-subspaces of the normalized modularity matrix of an edge-weighted graph (several clusters, symmetric case).

math.ST

Beyond the Expanders

Expander graphs are widely used in communication problems and construction of error correcting codes. In such graphs, information gets through very quickly. Typically, it is not true for social or biological networks, though we may find a partition of the vertices such that the induced subgraphs on them and the bipartite subgraphs between any pair of them exhibit regular behavior of information flow within or between the subsets. Implications between spectral and regularity properties are discussed.

math.CO

Testability of minimum balanced multiway cut densities

Testable weighted graph parameters and equivalent notions of testability are investigated based on papers of Laszlo Lovasz and coauthors. We prove that certain balanced minimum multiway cut densities are testable. Using this fact, quadratic programming techniques are applied to approximate some of these quantities. The problem is related to cluster analysis and statistical physics. Convergence of special noisy graph sequences is also discussed.

math.PR

Singular value decomposition of large random matrices (for two-way classification of microarrays)

Asymptotic behavior of the singular value decomposition (SVD) of blown up matrices and normalized blown up contingency tables exposed to Wigner-noise is investigated.It is proved that such an m\times n matrix almost surely has a constant number of large singular values (of order \sqrt{mn}), while the rest of the singular values are of order \sqrt{m+n} as m,n\to\infty. Concentration results of Alon et al. for the eigenvalues of large symmetric random matrices are adapted to the rectangular case, and on this basis, almost sure results for the singular values as well as for the corresponding isotropic subspaces are proved. An algorithm, applicable to two-way classification of microarrays, is also given that finds the underlying block structure.

math.PR