SearcharxivSearch

arXiv subjects

Behzad Aalipur

Publications and source records attributed to Behzad Aalipur.

3 recordsLinked to original sources

Exact Recovery by Neighborhood Smoothing in Directed Stochastic Block Models

We study exact community recovery in sparse directed stochastic block models using neighborhood smoothing of connection-probability profiles. The proposed method clusters vertices according to their estimated outgoing connection-probability profiles. For each vertex, its complete outgoing profile is estimated by averaging the adjacency rows of empirically similar vertices, after which \(K\)-means is applied to the estimated profiles. An analogous procedure based on incoming connection-probability profiles is obtained by applying the same construction to the transposed adjacency matrix. We establish a finite-sample uniform row-wise error bound for the asymmetric smoothed estimator and derive consistency in the normalized two-to-infinity norm. We then show that exact recovery follows when the minimum separation between distinct population profiles dominates the row-wise estimation error. The result permits a vanishing sparsity factor, a general asymmetric block-probability matrix, and a number of communities that may diverge with the network size. We are unaware of a previous exact-recovery theorem that simultaneously covers these features for a directed stochastic block model. Numerical studies and an application to a directed neuronal connectome illustrate the practical behavior and limitations of the profile-based clustering approach.

stat.ML

Leave-One-Out Neighborhood Smoothing for Graphons: Berry--Esseen Bounds, Confidence Intervals, and Edge-Wise No-Leakage Tuning

We investigate entrywise uncertainty quantification in graphon models, where existing methods can achieve strong estimation guarantees but often lack tractable distributional theory for inference on individual edge probabilities. We address this difficulty by introducing a leave-one-out (LOO) neighborhood smoother that constructs each neighborhood without using the corresponding target-column edges. Conditional on the selected neighborhood and the latent probability matrix, the observations being averaged are therefore independent Bernoulli variables whose success probabilities may differ. This separation yields finite-sample concentration bounds and a Berry--Esseen normal-approximation bound for each fixed target pair, with a numerical constant independent of that pair. Under a hybrid blockwise model with smooth blocks, blocks having identical probability profiles, and quantitative two-hop identifiability, we obtain an explicit bias bound that separates geometric localization error from stochastic error in the empirical two-hop distance. The resulting conditional risk upper bound balances at $h_n\asymp n^{2/3}$, whereas centered Gaussian inference requires a smaller neighborhood size. Bias control and Gaussian confidence intervals for $P_{ij}$ require the smoothing vertex to lie in an unobserved trimmed interior region; the observable empirical Bernstein interval instead targets the selected-neighborhood mean because bias-aware coverage of $P_{ij}$ requires an unknown model-dependent constant. Finally, the same leave-one-out structure yields an edge-wise no-leakage cross-validation identity: the held-out edge is not used to construct its own predictor.

stat.ME

Distributional Limits for Eigenvalues of Graphon Kernel Matrices

We study the fluctuation behavior of individual eigenvalues of kernel matrices arising from dense graphon-based random graphs. Under minimal integrability and boundedness assumptions on the graphon, we establish distributional limits for simple, well-separated eigenvalues of the associated integral operator. A sharp probabilistic dichotomy emerges: in the non-degenerate regime, the properly normalized empirical eigenvalue satisfies a central limit theorem with an explicit variance, whereas in the degenerate regime the leading stochastic term vanishes and the centered eigenvalue converges to a weighted chi-square law determined by the operator spectrum. The analysis requires no smoothness or Lipschitz conditions on the kernel. Prior work under comparable assumptions established only operator convergence and eigenspace consistency; the present results characterize the full distributional behavior of individual eigenvalues, extending fluctuation theory beyond the reach of classical operator-level arguments. The proofs combine second-order perturbation expansions, concentration bounds for kernel matrices, and Hoeffding decompositions for symmetric statistics, revealing that at the $\sqrt{n}$ scale the dominant randomness arises from latent-position sampling rather than Bernoulli edge noise.

math.PR