SearcharxivSearch

arXiv subjects

Sofia C. Olhede

Publications and source records attributed to Sofia C. Olhede.

At least 19 recordsLinked to original sources

nethist: An R package for Nonparametric Graphon Estimation via Network Histograms

Understanding the generative mechanism of real-world networks is crucial for analyzing connection patterns and making inference from network data. Graphons are widely used to model such mechanisms. Network histogram methods are nonparametric approaches based on blockmodel approximations that provide an intuitive view of network connection structures. However, there is a lack of software packages that construct network histograms. We introduce the R package nethist, which implements network histogram-type graphon estimators within a unified interface. This package is applicable to both single-layer and multilayer networks, and includes graphical summaries for examining both local and global network structures. By providing comprehensive network analysis tools, nethist facilitates understanding of complex systems of interrelated vertices.

stat.CO

Joint Estimation of Sparse Multilayer Networks via Graph Limits

Network datasets in modern applications often involve multiple types of interactions occurring over a shared set of individuals. Characterizing the generating mechanisms of these interactions can be enhanced by joint modelling, as shared vertices allow layers to help explain the structure of other layers. We model multiplex observations using graph limits, called a scaled set of graphons, and develop a nonparametric joint estimator based on blockmodel approximations, termed the multi-network histogram. This nonparametric framework captures each layer's varying sparsity and connection structure, accounting for heterogeneity via shared latent variables across all layers. We establish the theoretical properties of the multi-network histogram, providing an upper bound for the weighted mean integrated squared error and deriving the optimal bandwidth that minimizes this error. By leveraging information across layers, this joint modelling achieves a reduction in error and a smaller optimal bandwidth, which enables high-resolution estimation even in sparser layers. Its usefulness is demonstrated through simulation studies and an application to socioeconomic networks in an Indian village.

stat.ME

Irregularly and incompletely sampled random fields in the Earth sciences: Analysis and synthesis of parameterized covariance models

We study how sampling geometry contributes to uncertainty in modeling spatial geophysical observations as sampled random fields characterized by stationary, isotropic, parametric covariance functions. We incorporate the signature of discrete spatial sampling patterns into an asymptotically unbiased spectral maximum-likelihood estimation method along with analytical uncertainty calculation. We illustrate the broad applicability of our modeling through synthetic and real data examples with sampling patterns that include irregularly bounded contiguous region(s) of interest, structured sweeps of instrumental measurements, and missing observations dispersed across the domain of a field, which spur behaviors from the estimator. We find through asymptotic studies that allocating samples following a growing-domain strategy rather than a densifying, infill scheme best reduces estimator bias and (co)variance, whether the field has been sampled regularly or not. As our modeling assumptions, too, shape how (well) an observed random field can be characterized, we study the effect of covariance parameters assumed a priori. We demonstrate the desirable behavior of the general Matern class and show how to interrogate goodness-of-fit criteria to detect departures from the null hypothesis of Gaussianity, stationarity, and isotropy.

stat.ME

Graphon estimation beyond binary edges: inference for decorated graphs with applications to multiplex and weighted networks

We introduce the first doubly non-parametric estimation method for decorated graphons, a generalisation of graphons that encodes edge weights, edge types, and other edge-level attributes in large networks. Graphons describe the limiting behaviour of large unlabelled networks through a symmetric measurable function governing the probability of edge formation, but the standard framework is restricted to binary edge information. Decorated graphons lift this restriction, yet no inference procedure has previously been available for them. The proposed estimator extends classical graphon estimation techniques to this enriched setting. We derive rates of convergence and show that, for compactly supported decorations, these rates agree with known non-parametric rates for estimating real-valued functions. Monte Carlo experiments confirm that the theoretical rates are attained in finite samples, and applications to synthetic and empirical networks show improved fit relative to binary-edge baselines. The method extends graphon-based inference to multiplex networks and attributed graphs simultaneously.

stat.ME

Decorated graphons for temporal network estimation

We propose a unified nonparametric framework for modeling time-evolving networks using decorated graphons (also known as probability-graphons): symmetric functions that assign to each node pair a probability distribution over binary edge time series. This generalizes the static decorated-graphon construction to dynamic graphs while preserving node exchangeability and allowing temporal dynamics such as memory and periodicity. Models in which edges evolve independently given the latent variables, such as autoregressive and Markov edge processes, arise as special cases. We develop a two-stage estimation procedure that separates temporal modeling from network structure. Because the network stage requires only mild regularity conditions on the edge-process estimator, a broad class of temporal edge models can be used in the first stage. We establish nonparametric convergence rates in both block-model and Hölder-smooth regimes, and make explicit how the rate depends on the number of observed time steps and on the quality of the edge-level estimation. We illustrate the method on simulated data and a hospital contact network, recovering latent community structure and time-varying interaction patterns. The framework gives a nonparametric baseline for dynamic network analysis with explicit convergence guarantees.

stat.ME

The partial K function

The K function and its related statistics have been an enduring tool in the analysis of spatial point processes, providing an easy to compute and interpret summary statistic for characterising the interactions between points of one type, or between two different types of points. In this paper, we introduce a partial K function, enabling us to account for some of the effects of the other point types when analysing point-point interactions. The partial K function we introduce reduces to the usual K function when the other points are independent of the points of interest and has a similar interpretation. Using examples, we demonstrate how the partial K function can unpick dependence between point types that would otherwise be hidden in the usual K function. We also discuss important bias correction steps and hyperparameter selection. In addition, we introduce an extension to account for other spatial covariates, and demonstrate the methodology on the Lansing Woods dataset.

stat.ME

Maximum-likelihood estimation of the Matérn covariance structure of isotropic spatial random fields on finite, sampled grids

We present a statistically and computationally efficient spectral-domain maximum-likelihood procedure to solve for the structure of Gaussian spatial random fields within the Matern covariance hyperclass. For univariate, stationary, and isotropic fields, the three controlling parameters are the process variance, smoothness, and range. The debiased Whittle likelihood maximization explicitly treats discretization and edge effects for finite sampled regions in parameter estimation and uncertainty quantification. As even the best parameter estimate may not be good enough, we provide a test for whether the model specification itself warrants rejection. Our results are practical and relevant for the study of a variety of geophysical fields, and for spatial interpolation, out-of-sample extension, kriging, machine learning, and feature detection of geological data. We present procedural details and high-level results on real-world examples.

stat.ME

Spectral estimation for spatial point processes and random fields

Spatial variables can be observed in many different forms, such as regularly sampled random fields (lattice data), point processes, and randomly sampled spatial processes. Joint analysis of such collections of observations is clearly desirable, but complicated by the lack of an easily implementable analysis framework. We fill this gap by providing a multitaper analysis framework using coupled discrete and continuous data tapers, combined with the discrete Fourier transform for inference. Using this set of tools is important, as it forms the backbone for practical spectral analysis. In higher dimensions it is important not to be constrained to Cartesian product domains, and so we develop the methodology for spectral analysis using irregular domain data tapers, and the tapered discrete Fourier transform. We discuss its fast implementation, and the asymptotic as well as large finite domain properties. Estimators of partial association between different spatial processes are provided as are principled methods to determine their significance, and we demonstrate their practical utility on a large-scale ecological dataset.

stat.ME

A network and machine learning approach to detect Value Added Tax fraud

Value Added Tax (VAT) fraud erodes public revenue and puts legitimate businesses at a disadvantaged position thereby impacting inequality. Identifying and combating VAT fraud before it occurs is therefore important for welfare. This paper proposes flexible machine learning algorithms which detect fraudulent transactions, utilising the information provided by the complex VAT network structure of a large dimension. VAT fraud detection is implemented through a combination of a suitably constructed Laplacian matrix with classification algorithms that rely on scalable machine learning techniques. The method is implemented on the universe of Bulgarian VAT data and detects around 50 percent of the VAT fraud, outperforming well-known techniques that ignore the information provided by the network of VAT transactions. Importantly, the proposed methods are automated, and can be implemented following the taxpayers submission of their VAT returns. This allows tax revenue authorities to prevent large losses of tax revenues through performing early identification of fraud between business-to-business transactions within the VAT system.

physics.soc-ph

Entropy of Exchangeable Random Graphs

Quantifying the complexity of large graphs requires measures that extend beyond predefined structural features and scale efficiently with graph size. This work adopts a generative perspective, modeling large networks as exchangeable graphs to quantify the information content of their generating mechanisms via graphon entropy. As a graph property, graphon entropy is invariant under isomorphisms, making it an effective measure of complexity; however, it is not directly computable. To address this, we introduce a suite of graphon entropy estimators, including a nonparametric estimator for broad applicability and specialized versions for structured graphons arising from well-studied random graph models such as Erdős-Rényi, Chung-Lu, and stochastic block models. We establish their large-sample properties, deriving convergence rates and Central Limit Theorems. Simulations illustrate how the nonparametric graphon entropy estimator captures structural variations in graphs, while real-world applications demonstrate its role in characterizing evolving network dynamics.

cs.IT

Quantifying Multivariate Graph Dependencies: Theory and Estimation for Multiplex Graphs

Multiplex graphs, characterised by their layered structure, exhibit informative interdependencies within layers that are crucial for understanding complex network dynamics. Quantifying the interaction and shared information among these layers is challenging due to the non-Euclidean structure of graphs. Our paper introduces a comprehensive theory of multivariate information measures for multiplex graphs. We introduce graphon mutual information for pairs of graphs and expand this to graphon interaction information for three or more graphs, including their conditional variants. We then define graphon total correlation and graphon dual total correlation, along with their conditional forms, and introduce graphon $O-$information. We discuss and quantify the concepts of synergy and redundancy in graphs for the first time, introduce consistent nonparametric estimators for these multivariate graphon information--theoretic measures, and provide their convergence rates. We also conduct a simulation study to illustrate our theoretical findings and demonstrate the relationship between the introduced measures, multiplex graph structure, and higher--order interdependecies. Real-world applications further show the utility of our estimators in revealing shared information and dependence structures in real-world multiplex graphs. This work not only answers fundamental questions about information sharing across multiple graphs but also sets the stage for advanced pattern analysis in complex networks.

math.ST

Nonparametric inference of higher order interaction patterns in networks

We propose a method for obtaining parsimonious decompositions of networks into higher order interactions which can take the form of arbitrary motifs.The method is based on a class of analytically solvable generative models, where vertices are connected via explicit copies of motifs, which in combination with non-parametric priors allow us to infer higher order interactions from dyadic graph data without any prior knowledge on the types or frequencies of such interactions. Crucially, we also consider 'degree--corrected' models that correctly reflect the degree distribution of the network and consequently prove to be a better fit for many real world--networks compared to non-degree corrected models. We test the presented approach on simulated data for which we recover the set of underlying higher order interactions to a high degree of accuracy. For empirical networks the method identifies concise sets of atomic subgraphs from within thousands of candidates that cover a large fraction of edges and include higher order interactions of known structural and functional significance. The method not only produces an explicit higher order representation of the network but also a fit of the network to analytically tractable models opening new avenues for the systematic study of higher order network structures.

cs.SI

Hybrid of node and link communities for graphon estimation

Networks serve as a tool used to examine the large-scale connectivity patterns in complex systems. Modelling their generative mechanism nonparametrically is often based on step-functions, such as the stochastic block models. These models are capable of addressing two prominent topics in network science: link prediction and community detection. However, such methods often have a resolution limit, making it difficult to separate small-scale structures from noise. To arrive at a smoother representation of the network's generative mechanism, we explicitly trade variance for bias by smoothing blocks of edges based on stochastic equivalence. As such, we propose a different estimation method using a new model, which we call the stochastic shape model. Typically, analysis methods are based on modelling node or link communities. In contrast, we take a hybrid approach, bridging the two notions of community. Consequently, we obtain a more parsimonious representation, enabling a more interpretable and multiscale summary of the network structure. By considering multiple resolutions, we trade bias and variance to ensure that our estimator is rate-optimal. We also examine the performance of our model through simulations and applications to real network data.

stat.ME

Visualizing the Wavenumber Content of a Point Pattern

Spatial point patterns are a commonly recorded form of data in ecology, medicine, astronomy, criminology, epidemiology and many other application fields. One way to understand their second order dependence structure is via their spectral density function. However, unlike time series analysis, for point patterns such approaches are currently underutilized. In part, this is because the interpretation of the spectral representation of the underlying point processes is challenging. In this paper, we demonstrate how to band-pass filter point patterns, thus enabling us to explore the spectral representation of point patterns in space by isolating the signal corresponding to certain sets of wavenumbers.

stat.AP

What is the Fourier Transform of a Spatial Point Process?

This paper determines how to define a discretely implemented Fourier transform when analysing an observed spatial point process. To develop this transform we answer four questions; first what is the natural definition of a Fourier transform, and what are its spectral moments, second we calculate fourth order moments of the Fourier transform using Campbell's theorem. Third we determine how to implement tapering, an important component for spectral analysis of other stochastic processes. Fourth we answer the question of how to produce an isotropic representation of the Fourier transform of the process. This determines the basic spectral properties of an observed spatial point process.

stat.ME

Networks with Correlated Edge Processes

This article proposes methods to model nonstationary temporal graph processes. This corresponds to modelling the observation of edge variables (relationships between objects) indicating interactions between pairs of nodes (or objects) exhibiting dependence (correlation) and evolution in time over interactions. This article thus blends (integer) time series models with flexible static network models to produce models of temporal graph data, and statistical fitting procedures for time-varying interaction data. We illustrate the power of our proposed fitting method by analysing a hospital contact network, and this shows the high dimensional data challenge of modelling and inferring correlation between a large number of variables.

stat.ME

The Debiased Spatial Whittle Likelihood

We provide a computationally and statistically efficient method for estimating the parameters of a stochastic covariance model observed on a regular spatial grid in any number of dimensions. Our proposed method, which we call the Debiased Spatial Whittle likelihood, makes important corrections to the well-known Whittle likelihood to account for large sources of bias caused by boundary effects and aliasing. We generalise the approach to flexibly allow for significant volumes of missing data including those with lower-dimensional substructure, and for irregular sampling boundaries. We build a theoretical framework under relatively weak assumptions which ensures consistency and asymptotic normality in numerous practical settings including missing data and non-Gaussian processes. We also extend our consistency results to multivariate processes. We provide detailed implementation guidelines which ensure the estimation procedure can be conducted in O(n log n) operations, where n is the number of points of the encapsulating rectangular grid, thus keeping the computational scalability of Fourier and Whittle-based methods for large data sets. We validate our procedure over a range of simulated and real-world settings, and compare with state-of-the-art alternatives, demonstrating the enduring practical appeal of Fourier-based methods, provided they are corrected by the procedures developed in this paper.

stat.ME

Edge coherence in multiplex networks

This paper introduces a nonparametric framework for the setting where multiple networks are observed on the same set of nodes, also known as multiplex networks. Our objective is to provide a simple parameterization which explicitly captures linear dependence between the different layers of networks. For non-Euclidean observations, such as shapes and graphs, the notion of "linear" must be defined appropriately. Taking inspiration from the representation of stochastic processes and the analogy of the multivariate spectral representation of a stochastic process with joint exchangeability of Bernoulli arrays, we introduce the notion of edge coherence as a measure of linear dependence in the graph limit space. Edge coherence is defined for pairs of edges from any two network layers and is the key novel parameter. We illustrate the utility of our approach by eliciting simple models such as a correlated stochastic blockmodel and a correlated inhomogeneous graph limit model.

stat.ME