Searcharxiv⌕ Search

arXiv subjects

Ting Yan

Publications and source records attributed to Ting Yan.

At least 37 records · Page 2Linked to original sources

Chip-scale sensor for spectroscopic metrology

Miniaturized spectrometers hold great promise for in situ, in vitro, and even in vivo sensing applications. However, their size reduction imposes vital performance constraints in meeting the rigorous demands of spectroscopy, including fine resolution, high accuracy, and ultra-wide observation window. The prevailing view in the community holds that miniaturized spectrometers are most suitable for the coarse identification of signature peaks. In this paper, we present an integrated reconstructive spectrometer that enables near-infrared (NIR) spectroscopic metrology, and demonstrate a fully packaged sensor with auxiliary electronics. Such a sensor operates over a 520 nm bandwidth together with a resolution of less than 8 pm, which translates into a record-breaking bandwidth-to-resolution ratio of over 65,000. The classification of different types of solid substances and the concentration measurement of aqueous and organic solutions are performed, all achieving approximately 100% accuracy. Notably, the detection limit of our sensor matches that of the commercial benchtop counterparts, which is as low as 0.1% (i.e. 100 mg/dL) for identifying the concentration of glucose solution.

physics.optics↗

A two-way heterogeneity model for dynamic networks

Dynamic network data analysis requires joint modelling individual snapshots and time dynamics. This paper proposes a new two-way heterogeneity model towards this goal. The new model equips each node of the network with two heterogeneity parameters, one to characterize the propensity of forming ties with other nodes and the other to differentiate the tendency of retaining existing ties over time. Though the negative log-likelihood function is non-convex, it is locally convex in a neighbourhood of the true value of the parameter vector. By using a novel method of moments estimator as the initial value, the consistent local maximum likelihood estimator (MLE) can be obtained by a gradient descent algorithm. To establish the upper bound for the estimation error of the MLE, we derive a new uniform deviation bound, which is of independent interest. The usefulness of the model and the associated theory are further supported by extensive simulation and the analysis of some real network data sets.

stat.ME↗

Differentially private analysis of networks with covariates via a generalized $β$-model

How to achieve the tradeoff between privacy and utility is one of fundamental problems in private data analysis.In this paper, we give a rigourous differential privacy analysis of networks in the appearance of covariates via a generalized $β$-model, which has an $n$-dimensional degree parameter $β$ and a $p$-dimensional homophily parameter $γ$.Under $(k_n, ε_n)$-edge differential privacy, we use the popular Laplace mechanism to release the network statistics.The method of moments is used to estimate the unknown model parameters. We establish the conditions guaranteeing consistency of the differentially private estimators $\widehatβ$ and $\widehatγ$ as the number of nodes $n$ goes to infinity, which reveal an interesting tradeoff between a privacy parameter and model parameters. The consistency is shown by applying a two-stage Newton's method to obtain the upper bound of the error between $(\widehatβ,\widehatγ)$ and its true value $(β, γ)$ in terms of the $\ell_\infty$ distance, which has a convergence rate of rough order $1/n^{1/2}$ for $\widehatβ$ and $1/n$ for $\widehatγ$, respectively. Further, we derive the asymptotic normalities of $\widehatβ$ and $\widehatγ$, whose asymptotic variances are the same as those of the non-private estimators under some conditions. Our paper sheds light on how to explore asymptotic theory under differential privacy in a principled manner; these principled methods should be applicable to a class of network models with covariates beyond the generalized $β$-model. Numerical studies and a real data analysis demonstrate our theoretical findings.

stat.ME↗

Time-varying $β$-model for dynamic directed networks

We extend the well-known $β$-model for directed graphs to dynamic network setting, where we observe snapshots of adjacency matrices at different time points. We propose a kernel-smoothed likelihood approach for estimating $2n$ time-varying parameters in a network with $n$ nodes, from $N$ snapshots. We establish consistency and asymptotic normality properties of our kernel-smoothed estimators as either $n$ or $N$ diverges. Our results contrast their counterparts in single-network analyses, where $n\to\infty$ is invariantly required in asymptotic studies. We conduct comprehensive simulation studies that confirm our theory's prediction and illustrate the performance of our method from various angles. We apply our method to an email data set and obtain meaningful results.

stat.ME↗

A degree-corrected Cox model for dynamic networks

Continuous time network data have been successfully modeled by multivariate counting processes, in which the intensity function is characterized by covariate information. However, degree heterogeneity has not been incorporated into the model which may lead to large biases for the estimation of homophily effects. In this paper, we propose a degree-corrected Cox network model to simultaneously analyze the dynamic degree heterogeneity and homophily effects for continuous time directed network data. Since each node has individual-specific in- and out-degree effects in the model, the dimension of the time-varying parameter vector grows with the number of nodes, which makes the estimation problem non-standard. We develop a local estimating equations approach to estimate unknown time-varying parameters, and establish consistency and asymptotic normality of the proposed estimators by using the powerful martingale process theories. We further propose test statistics to test for trend and degree heterogeneity in dynamic networks. Simulation studies are provided to assess the finite sample performance of the proposed method and a real data analysis is used to illustrate its practical utility.

math.ST↗

Wilks' theorems in the $β$-model

Likelihood ratio tests and the Wilks theorems have been pivotal in statistics but have rarely been explored in network models with an increasing dimension. We are concerned here with likelihood ratio tests in the $β$-model for undirected graphs. For two growing dimensional null hypotheses including a specified null $H_0: β_i=β_i^0$ for $i=1,\ldots, r$ and a homogenous null $H_0: β_1=\cdots=β_r$, we reveal high dimensional Wilks' phenomena that the normalized log-likelihood ratio statistic, $[2\{\ell(\widehat{\boldsymbolβ}) - \ell(\widehat{\boldsymbolβ}^0)\} - r]/(2r)^{1/2}$, converges in distribution to the standard normal distribution as $r$ goes to infinity. Here, $\ell( \boldsymbolβ)$ is the log-likelihood function on the vector parameter $\boldsymbolβ=(β_1, \ldots, β_n)^\top$, $\widehat{\boldsymbolβ}$ is its maximum likelihood estimator (MLE) under the full parameter space, and $\widehat{\boldsymbolβ}^0$ is the restricted MLE under the null parameter space. For the corresponding fixed dimensional null $H_0: β_i=β_i^0$ for $i=1,\ldots, r$ and the homogenous null $H_0: β_1=\cdots=β_r$ with a fixed $r$, we establish Wilks type of results that $2\{\ell(\widehat{\boldsymbolβ}) - \ell(\widehat{\boldsymbolβ}^0)\}$ converges in distribution to a Chi-square distribution with respective $r$ and $r-1$ degrees of freedom, as the total number of parameters, $n$, goes to infinity. The Wilks type of results are further extended into a closely related Bradley--Terry model for paired comparisons, where we discover a different phenomenon that the log-likelihood ratio statistic under the fixed dimensional specified null asymptotically follows neither a Chi-square nor a rescaled Chi-square distribution. Simulation studies and an application to NBA data illustrate the theoretical results.

math.ST↗

Asymptotic theory in network models with covariates and a growing number of node parameters

We propose a general model that jointly characterizes degree heterogeneity and homophily in weighted, undirected networks. We present a moment estimation method using node degrees and homophily statistics. We establish consistency and asymptotic normality of our estimator using novel analysis. We apply our general framework to three applications, including both exponential family and non-exponential family models. Comprehensive numerical studies and a data example also demonstrate the usefulness of our method.

math.ST↗

Wilks' theorems in some exponential random graph models

We are concerned here with the likelihood ratio statistics in two exponential random graph models -- the $β$-model and the Bradley-Terry model, in which the degree sequence on an undirected graph and the out-degree sequence on a weighted directed graph are the exclusively sufficient statistics in the exponential-family distributions on graphs, respectively. We prove the Wilks type of theorems for some fixed and growing dimensional hypothesis testing problems. More specifically, under two fixed dimensional null hypotheses $H_0: β_i=β_i^0$ for $i=1,\ldots, r$ and $H_0: β_1=\ldots=β_r$, we show that $2[\ell(\widehat{\boldsymbolβ}) - \ell(\widehat{\boldsymbolβ}^0)]$ converges in distribution to a Chi-square distribution with the respective degrees of freedoms, $r$ and $r-1$, as the dimension $n$ of the full parameter space goes to infinity. Here, $\ell(\boldsymbolβ)$ is the log-likelihood function on the parameter $\boldsymbolβ$, $\widehat{\boldsymbolβ}$ is the MLE under the full parameter space, and $\widehat{\boldsymbolβ}^0$ is the restricted MLE under the null parameter space. For two increasing dimensional null hypotheses $H_0: β_i = β_i^0$ for $i=1, \ldots, n$ and $H_0: β_1=\ldots=β_r$ with $r/n \ge c$, we show that the normalized log-likelihood ratio statistics, $(2[\ell(\widehat{\boldsymbolβ}) - \ell(\boldsymbolβ^0)] -n)/(2n)^{1/2}$ and $(2[\ell(\widehat{\boldsymbolβ}) - \ell(\widehat{\boldsymbolβ}^0)] -r)/(2r)^{1/2}$, both converge in distribution to the standard normal distribution. Simulation studies and an application to NBA data illustrate the theoretical results.

math.ST↗

Asymptotic Theory for Differentially Private Generalized $β$-models with Parameters Increasing

Modelling edge weights play a crucial role in the analysis of network data, which reveals the extent of relationships among individuals. Due to the diversity of weight information, sharing these data has become a complicated challenge in a privacy-preserving way. In this paper, we consider the case of the non-denoising process to achieve the trade-off between privacy and weight information in the generalized $β$-model. Under the edge differential privacy with a discrete Laplace mechanism, the Z-estimators from estimating equations for the model parameters are shown to be consistent and asymptotically normally distributed. The simulations and a real data example are given to further support the theoretical results.

math.ST↗

Directed Networks with a Differentially Private Bi-degree Sequence

Although a lot of approaches are developed to release network data with a differentially privacy guarantee, inference using noisy data in many network models is still unknown or not properly explored. In this paper, we release the bi-degree sequences of directed networks using the Laplace mechanism and use the $p_0$ model for inferring the degree parameters. The $p_0$ model is an exponential random graph model with the bi-degree sequence as its exclusively sufficient statistic. We show that the estimator of the parameter without the denoised process is asymptotically consistent and normally distributed. This is contrast sharply with some known results that valid inference such as the existence and consistency of the estimator needs the denoised process. Along the way, a new phenomenon is revealed in which an additional variance factor appears in the asymptotic variance of the estimator when the noise becomes large. Further, we propose an efficient algorithm for finding the closet point lying in the set of all graphical bi-degree sequences under the global $L_1$ optimization problem. Numerical studies demonstrate our theoretical findings.

stat.ME↗

A Unified Framework for Inference in Network Models with Degree Heterogeneity and Homophily

The degree heterogeneity and homophily are two typical features in network data. In this paper, we formulate a general model for undirected networks with these two features and present the moment estimation for inferring the degree and homophily parameters. The binary or nonbinary network edges are simultaneously considered. We establish a unified theoretical framework under which the consistency of the moment estimator holds as the size of networks goes to infinity. We also derive the asymptotic representation of the moment estimator that can be used to characterize its limiting distribution. The asymptotic representation of the moment estimator of the homophily parameter contains a bias term. Two applications are provided to illustrate the theoretical result. Numerical studies and a real data analysis demonstrate our theoretical findings.

stat.ME↗

Corrected Bayesian information criterion for stochastic block models

Estimating the number of communities is one of the fundamental problems in community detection. We re-examine the Bayesian paradigm for stochastic block models and propose a "corrected Bayesian information criterion",to determine the number of communities and show that the proposed estimator is consistent under mild conditions. The proposed criterion improves those used in Wang and Bickel (2016) and Saldana et al. (2017) which tend to underestimate and overestimate the number of communities, respectively. Along the way, we establish the Wilks theorem for stochastic block models. Moreover, we show that, to obtain the consistency of model selection for stochastic block models, we need a so-called "consistency condition". We also provide sufficient conditions for both homogenous networks and non-homogenous networks. The results are further extended to degree corrected stochastic block models. Numerical studies demonstrate our theoretical results.

stat.ME↗

Using Maximum Entry-Wise Deviation to Test the Goodness-of-Fit for Stochastic Block Models

The stochastic block model is widely used for detecting community structures in network data. How to test the goodness-of-fit of the model is one of the fundamental problems and has gained growing interests in recent years. In this article, we propose a novel goodness-of-fit test based on the maximum entry of the centered and re-scaled adjacency matrix for the stochastic block model. One noticeable advantage of the proposed test is that the number of communities can be allowed to grow linearly with the number of nodes ignoring a logarithmic factor. We prove that the null distribution of the test statistic converges in distribution to a Gumbel distribution, and we show that both the number of communities and the membership vector can be tested via the proposed method. Further, we show that the proposed test has asymptotic power guarantee against a class of alternatives. We also demonstrate that the proposed method can be extended to the degree-corrected stochastic block model. Both simulation studies and real-world data examples indicate that the proposed method works well.

stat.ME↗

Approximating the inverse of a diagonally dominant matrix with positive elements

For an $n\times n$ diagonally dominant matrix $T=(t_{i,j})_{n\times n}$ with positive elements satisfying certain bounding conditions, we propose to use a diagonal matrix $S=(s_{i,j})_{n\times n}$ to approximate the inverse of $T$, where $s_{i,j}=δ_{i,j}/t_{i,i}$ and $δ_{i,j}$ is the Kronecker delta function. We derive an explicitly upper bound on the approximation error, which is in the magnitude of $O(n^{-2})$. It shows that $S$ is a very good approximation to $T^{-1}$.

math.NA↗

A Probit Network Model with Arbitrary Dependence

In this paper, we adopt a latent variable method to formulate a network model with arbitrarily dependent structure. We assume that the latent variables follow a multivariate normal distribution and a link between two nodes forms if the sum of the corresponding node parameters exceeds the latent variable. The dependent structure among edges is induced by the covariance matrix of the latent variables. The marginal distribution of an edge is a probit function. We refer this model to as the \emph{Probit Network Model}. We show that the moment estimator of the node parameter is consistent. To the best of our knowledge, this is the first time to derive consistency result in a single observed network with globally dependent structures. We extend the model to allow node covariate information.

stat.ME↗

Statistical Inference in a Directed Network Model with Covariates

Networks are often characterized by node heterogeneity for which nodes exhibit different degrees of interaction and link homophily for which nodes sharing common features tend to associate with each other. In this paper, we propose a new directed network model to capture the former via node-specific parametrization and the latter by incorporating covariates. In particular, this model quantifies the extent of heterogeneity in terms of outgoingness and incomingness of each node by different parameters, thus allowing the number of heterogeneity parameters to be twice the number of nodes. We study the maximum likelihood estimation of the model and establish the uniform consistency and asymptotic normality of the resulting estimators. Numerical studies demonstrate our theoretical findings and a data analysis confirms the usefulness of our model.

stat.ME↗

Affiliation networks with an increasing degree sequence

Affiliation network is one kind of two-mode social network with two different sets of nodes (namely, a set of actors and a set of social events) and edges representing the affiliation of the actors with the social events. Although a number of statistical models are proposed to analyze affiliation networks, the asymptotic behaviors of the estimator are still unknown or have not been properly explored. In this paper, we study an affiliation model with the degree sequence as the exclusively natural sufficient statistic in the exponential family distributions. We establish the uniform consistency and asymptotic normality of the maximum likelihood estimator when the numbers of actors and events both go to infinity. Simulation studies and a real data example demonstrate our theoretical results.

stat.ME↗

Asymptotic generalized bivariate extreme with random index

In many biological, agricultural, military activity problems and in some quality control problems, it is almost impossible to have a fixed sample size, because some observations are always lost for various reasons. Therefore, the sample size itself is considered frequently to be a random variable (rv). The class of limit distribution functions (df's) of the random bivariate extreme generalized order statistics (GOS) from independent and identically distributed RV's are fully characterized. When the random sample size is assumed to be independent of the basic variables and its df is assumed to converge weakly to a non-degenerate limit, the necessary and sufficient conditions for the weak convergence of the random bivariate extreme GOS are obtained. Furthermore, when the interrelation of the random size and the basic rv's is not restricted, sufficient conditions for the convergence and the forms of the limit df's are deduced. Illustrative examples are given which lend further support to our theoretical results.

math.ST↗