SearcharxivSearch

arXiv subjects

Tiandong Wang

Publications and source records attributed to Tiandong Wang.

At least 19 recordsLinked to original sources

Pooling Mobility Obscures Epidemic Invasion Routes

Epidemic models often pool air travel, commuting, and other mobility layers into a single weighted network. Pooling keeps the total imported infections into a region but discards the transport mode and route that delivered them, the information a mode-specific intervention needs. We make this precise for directed multilayer flows with heavy-tailed variation. Pooling acts as an asymmetric filter. The magnitude of an extreme importation, its radial tail index, is an exact invariant of sampling and nonnegative aggregation, whereas its composition across layers and routes, the angular part, is reweighted by layer-specific sampling and can be irrecoverably merged. We give a necessary and sufficient condition for recovering route composition, attach a surveillance decision cost to the loss, and show that the first established route need not be the busiest. A controlled metapopulation study calibrates the cost, and two contrasting reconstructions show where it bites. In US pandemic influenza, air adds fitted information beyond commuting, and a layer-resolved ranking targets more air-import burden than a pooled one. In the early Italian COVID-19 wave, commuting improves onset reconstruction while air attribution stays unresolved. Pooling can therefore support total-importation surveillance but cannot in general identify the mode or route behind an importation.

math.ST

Hub Neighbor-Degree Diagnostics for Sparse Random Graphs

Networks with nearly identical degree distributions can place their hubs in sharply different neighborhoods. We develop a model diagnostic based on the mean degree of the neighbors of a degree-$k$ vertex. Under rank-one inhomogeneous random graphs, this statistic has degree-invariant centering and $k^{-1/2}$ fluctuations. Under non-rank-one kernels, posterior uncertainty about the root type can instead determine both centering and scale. Under linear preferential attachment, the statistic grows as $(m+\delta)\log k$. We turn these model-specific limits into goodness-of-fit tests for specified sparse-graph nulls and a weighted log-degree slope test for residual hub-neighborhood trends. Simulations evaluate null calibration, degree-distribution misspecification, and power against degree-matched preferential-attachment alternatives. Applications to high-school contact and arXiv coauthorship networks show that the method separates level misspecification from disassortative and positive residual trends. Reddit interaction networks provide a further appendix example.

math.ST

Efficient Topic Model Estimation under Heavy-Tailed Document Lengths

Early inquiries into the statistical properties of natural language found that words tend to occur with power-law frequencies. This observation, closely associated with Zipf's law, has spurred many investigations into why this power-law pattern emerges with such regularity. Rarely, however, has this property of text been leveraged in statistical inference. In this paper, we demonstrate that the Latent Dirichlet Allocation (LDA) model can accommodate power-law word frequencies. In particular, when the document length distribution is regularly varying, the word frequency distribution admits a hierarchy of power laws across documents and topics. We further leverage this finding to develop an efficient tensor decomposition algorithm for estimating the topic matrix via the moments of normalized extreme word frequencies. Applying our algorithm to the twenty newsgroups corpus reveals that the extreme-value methodology exhibits robustness to certain choices made in the pre-processing of the data. This work furthers the recent interest in adapting machine learning methods to the study of multivariate extremes.

stat.ME

Spatial Dependence in Directed Preferential-Attachment Networks

Spatially embedded directed networks, such as airline networks, often exhibit simultaneous high activity at nearby nodes. Preferential attachment (PA) explains hub dominance. We extend it to spatial co-movement through a directed PA model whose out- and in-node weights follow temporally persistent Gaussian-process lognormal fields. Under sublinear PA, out-degree proportions converge to explicit normalized powered weights, whereas self-loop exclusion yields a coupled in-degree limit. We derive a strictly concave inverse that recovers the in-weights from terminal degree proportions. For ordered network histories, we develop a minorization-maximization (MM) weight estimator and profile likelihood for the PA exponent; temporal pre-whitening and a spatial quasi-likelihood estimate the latent covariance. Simulations verify transmission of distance-decaying dependence and show how random segment volume creates a distance-independent common mode in raw degrees. An analysis of U.S. domestic flights (2015-2019) separates network-wide volume variation from a short-range spatial component. An observed-volume reconstruction reproduces the raw-degree common mode, and the fitted field yields an exploratory co-exceedance transition scale of roughly 150 km. A per-carrier analysis of European air traffic also reveals the same decomposition.

stat.ME

A Multilayer Probit Network Model for Community Detection with Dependent Edges and Layers

Community detection in multilayer networks, which aims to identify groups of nodes exhibiting similar connectivity patterns across multiple network layers, has attracted considerable attention in recent years. Most existing methods are based on the assumption that different layers are either independent or follow specific dependence structures, and edges within the same layer are independent. In this article, we propose a novel method for community detection in multilayer networks that accounts for a broad range of inter-layer and intra-layer dependence structures. The proposed method integrates the multilayer stochastic block model for community detection with a multivariate probit model to capture the structures of inter-layer dependence, which also allows intra-layer dependence. To facilitate parameter estimation, we develop a constrained pairwise likelihood method coupled with an efficient alternating updating algorithm. The asymptotic properties of the proposed method are also established, with a focus on examining the influence of inter-layer and intra-layer dependences on the accuracy of both parameter estimation and community detection. The theoretical results are supported by extensive numerical experiments on both simulated networks and a real-world multilayer trade network.

stat.ME

Structural Causal Models for Extremes: an Approach Based on Exponent Measures

We introduce a new formulation of structural causal models for extremes, called the extremal structural causal model (eSCM). Unlike conventional structural causal models, where randomness is governed by a probability distribution, eSCMs use an exponent measure, an infinite-mass law that naturally arises in the analysis of multivariate extremes. Central to this framework are activation variables, which abstract the single-big-jump principle, along with additional randomization that enriches the class of eSCM laws. This formulation encompasses all possible laws of directed graphical models under the recently introduced notion of extremal conditional independence. We also identify an inherent asymmetry in eSCMs under natural assumptions, enabling the identifiability of causal directions, a central challenge in causal inference. Finally, we propose a method that utilizes this causal asymmetry and demonstrate its effectiveness in both simulated and real datasets.

math.ST

Classification of Extremal Dependence in Financial Markets via Bootstrap Inference

Accurately identifying the extremal dependence structure in multivariate heavy-tailed data is a fundamental yet challenging task, particularly in financial applications. Following a recently proposed bootstrap-based testing procedure, we apply the methodology to absolute log returns of U.S. S&P 500 and Chinese A-share stocks over a time period well before the U.S. election in 2024. The procedure reveals more isolated clustering of dependent assets in the U.S. economy compared with China which exhibits different characteristics and a more interconnected pattern of extremal dependence. Cross-market analysis identifies strong extremal linkages in sectors such as materials, consumer staples and consumer discretionary, highlighting the effectiveness of the testing procedure for large-scale empirical applications.

math.ST

Multilayer Network Regression with Eigenvector Centrality and Community Structure

In the analysis of complex networks, centrality measures and community structures play pivotal roles. For multilayer networks, a critical challenge lies in effectively integrating information across diverse layers while accounting for the dependence structures both within and between layers. We propose an innovative two-stage regression model for multilayer networks, combining eigenvector centrality and network community structure within fourth-order tensor-like multilayer networks. We develop new community-based centrality measures, integrated into a regression framework. To address the inherent noise in network data, we conduct separate analyses of centrality measures with and without measurement errors and establish consistency for the least squares estimates in the regression model. The proposed methodology is applied to the world input-output dataset, investigating how input-output network data among different countries and industries influence the gross output of each industry.

stat.ME

Systemic Risk Management via Maximum Independent Set in Extremal Dependence Networks

The failure of key financial institutions may accelerate risk contagion due to their interconnections within the system. In this paper, we propose a robust portfolio strategy to mitigate systemic risks during extreme events. We use the stock returns of key financial institutions as an indicator of their performance, apply extreme value theory to assess the extremal dependence among stocks of financial institutions, and construct a network model based on a threshold approach that captures extremal dependence. Our analysis reveals different dependence structures in the Chinese and U.S. financial systems. By applying the maximum independent set (MIS) from graph theory, we identify a subset of institutions with minimal extremal dependence, facilitating the construction of diversified portfolios resilient to risk contagion. We also compare the performance of our proposed portfolios with that of the market portfolios in the two economies.

q-fin.PM

Measuring Interlayer Dependence of Large Degrees in Multilayer Inhomogeneous Random Graphs

Accurately capturing interlayer dependence is essential for understanding the structure of complex multilayer networks. We propose an upper tail dependence estimator specifically designed for multilayer networks, leveraging multilayer inhomogeneous random graphs and multivariate regular variation to model extremal dependence. We establish the consistency of the estimator and demonstrate its practical effectiveness through real-data analysis of Reddit. Our findings reveal how financial market dynamics influence user interactions in the BitcoinMarkets subreddit and how seasonal trends shape engagement in sports-related subreddits. This work provides a rigorous and practical tool for quantifying extremal dependence across network layers, offering valuable insights into risk propagation and interaction patterns in multilayer systems.

stat.AP

On tail inference in scale-free inhomogeneous random graphs

Both empirical and theoretical investigations of scale-free network models have found that large degrees in a network exert an outsized impact on its structure. However, the tools used to infer the tail behavior of degree distributions in scale-free networks often lack a strong theoretical foundation. In this paper, we introduce a new framework for analyzing the asymptotic distribution of estimators for degree tail indices in scale-free inhomogeneous random graphs. Our framework leverages the relationship between the large weights and large degrees of Norros-Reittu and Chung-Lu random graphs. In particular, we determine a rate for the number of nodes $k(n) \rightarrow \infty$ such that for all $i = 1, \dots, k(n)$, the node with the $i$-th largest weight will have the $i$-th largest degree with high probability. Such alignment of upper-order statistics is then employed to establish the asymptotic normality of three different tail index estimators based on the upper degrees. These results suggest potential applications of the framework to threshold selection and goodness-of-fit testing in scale-free networks, issues that have long challenged the network science community.

math.ST

Mitigating Extremal Risks: A Network-Based Portfolio Strategy

In financial markets marked by inherent volatility, extreme events can result in substantial investor losses. This paper proposes a portfolio strategy designed to mitigate extremal risks. By applying extreme value theory, we evaluate the extremal dependence between stocks and develop a network model reflecting these dependencies. We use a threshold-based approach to construct this complex network and analyze its structural properties. To improve risk diversification, we utilize the concept of the maximum independent set from graph theory to develop suitable portfolio strategies. Since finding the maximum independent set in a given graph is NP-hard, we further partition the network using either sector-based or community-based approaches. Additionally, we use value at risk and expected shortfall as specific risk measures and compare the performance of the proposed portfolios with that of the market portfolio.

q-fin.PM

Likelihood-based Inference for Random Networks with Changepoints

Generative, temporal network models play an important role in analyzing the dependence structure and evolution patterns of complex networks. Due to the complicated nature of real network data, it is often naive to assume that the underlying data-generative mechanism itself is invariant with time. Such observation leads to the study of changepoints or sudden shifts in the distributional structure of the evolving network. In this paper, we propose a likelihood-based methodology to detect changepoints in undirected, affine preferential attachment networks, and establish a hypothesis testing framework to detect a single changepoint, together with a consistent estimator for the changepoint. Such results require establishing consistency and asymptotic normality of the MLE under the changepoint regime, which suffers from long range dependence. The methodology is then extended to the multiple changepoint setting via both a sliding window method and a more computationally efficient score statistic. We also compare the proposed methodology with previously developed non-parametric estimators of the changepoint via simulation, and the methods developed herein are applied to modeling the popularity of a topic in a Twitter network over time.

stat.ME

Simultaneous Identification of Sparse Structures and Communities in Heterogeneous Graphical Models

Exploring and detecting community structures hold significant importance in genetics, social sciences, neuroscience, and finance. Especially in graphical models, community detection can encourage the exploration of sets of variables with group-like properties. In this paper, within the framework of Gaussian graphical models, we introduce a novel decomposition of the underlying graphical structure into a sparse part and low-rank diagonal blocks (non-overlapped communities). We illustrate the significance of this decomposition through two modeling perspectives and propose a three-stage estimation procedure with a fast and efficient algorithm for the identification of the sparse structure and communities. Also on the theoretical front, we establish conditions for local identifiability and extend the traditional irrepresentability condition to an adaptive form by constructing an effective norm, which ensures the consistency of model selection for the adaptive $\ell_1$ penalized estimator in the second stage. Moreover, we also provide the clustering error bound for the K-means procedure in the third stage. Extensive numerical experiments are conducted to demonstrate the superiority of the proposed method over existing approaches in estimating graph structures. Furthermore, we apply our method to the stock return data, revealing its capability to accurately identify non-overlapped community structures.

stat.ML

Emergence of Multivariate Extremes in Multilayer Inhomogeneous Random Graphs

In this paper, we propose a multilayer inhomogeneous random graph model (MIRG), whose layers may consist of both single-edge and multi-edge graphs. In the single layer case, it has been shown that the regular variation of the weight distribution underlying the inhomogeneous random graph implies the regular variation of the typical degree distribution. We extend this correspondence to the multilayer case by showing that the multivariate regular variation of the weight distribution implies the multivariate regular variation of the asymptotic degree distribution. Furthermore, in certain circumstances, the extremal dependence structure present in the weight distribution will be adopted by the asymptotic degree distribution. By considering the asymptotic degree distribution, a wider class of Chung-Lu and Norros-Reittu graphs may be incorporated into the MIRG layers. Additionally, we prove consistency of the Hill estimator when applied to degrees of the MIRG that have a tail index greater than 1. Simulation results indicate that, in practice, hidden regular variation may be consistently detected from an observed MIRG.

math.PR

2RV+HRV and Testing for Strong VS Full Dependence

Preferential attachment models of network growth are bivariate heavy tailed models for in- and out-degree with limit measures which either concentrate on a ray of positive slope from the origin or on all of the positive quadrant depending on whether the model includes reciprocity or not. Concentration on the ray is called full dependence. If there were a reliable way to distinguish full dependence from not-full, we would have guidance about which model to choose. This motivates investigating tests that distinguish between (i) full dependence; (ii) strong dependence (support of the limit measure is a proper subcone of the positive quadrant); (iii) weak dependence (limit measure concentrates on positive quadrant). We give two test statistics, analyze their asymptotically normal behavior under full and not-full dependence, and discuss applicability using bootstrap methods applied to simulated and real data.

math.ST

Generating General Preferential Attachment Networks with R Package wdnet

Preferential attachment (PA) network models have a wide range of applications in various scientific disciplines. Efficient generation of large-scale PA networks helps uncover their structural properties and facilitate the development of associated analytical methodologies. Existing software packages only provide limited functions for this purpose with restricted configurations and efficiency. We present a generic, user-friendly implementation of weighted, directed PA network generation with R package wdnet. The core algorithm is based on an efficient binary tree approach. The package further allows adding multiple edges at a time, heterogeneous reciprocal edges, and user-specified preference functions. The engine under the hood is implemented in C++. Usages of the package are illustrated with detailed explanation. A benchmark study shows that wdnet is efficient for generating general PA networks not available in other packages. In restricted settings that can be handled by existing packages, wdnet provides comparable efficiency.

stat.CO

Modeling Random Networks with Heterogeneous Reciprocity

Reciprocity, or the tendency of individuals to mirror behavior, is a key measure that describes information exchange in a social network. Users in social networks tend to engage in different levels of reciprocal behavior. Differences in such behavior may indicate the existence of communities that reciprocate links at varying rates. In this paper, we develop methodology to model the diverse reciprocal behavior in growing social networks. In particular, we present a preferential attachment model with heterogeneous reciprocity that imitates the attraction users have for popular users, plus the heterogeneous nature by which they reciprocate links. We compare Bayesian and frequentist model fitting techniques for large networks, as well as computationally efficient variational alternatives. Cases where the number of communities are known and unknown are both considered. We apply the presented methods to the analysis of a Facebook wallpost network where users have non-uniform reciprocal behavior patterns. The fitted model captures the heavy-tailed nature of the empirical degree distributions in the Facebook data and identifies multiple groups of users that differ in their tendency to reply to and receive responses to wallposts.

stat.ML