Searcharxiv⌕ Search

arXiv subjects

Zheng Yu

Publications and source records attributed to Zheng Yu.

42 records · Page 3Linked to original sources

Generalized Leverage Score Sampling for Neural Networks

Leverage score sampling is a powerful technique that originates from theoretical computer science, which can be used to speed up a large number of fundamental questions, e.g. linear regression, linear programming, semi-definite programming, cutting plane method, graph sparsification, maximum matching and max-flow. Recently, it has been shown that leverage score sampling helps to accelerate kernel methods [Avron, Kapralov, Musco, Musco, Velingker and Zandieh 17]. In this work, we generalize the results in [Avron, Kapralov, Musco, Musco, Velingker and Zandieh 17] to a broader class of kernels. We further bring the leverage score sampling into the field of deep learning theory. $\bullet$ We show the connection between the initialization for neural network training and approximating the neural tangent kernel with random features. $\bullet$ We prove the equivalence between regularized neural network and neural tangent kernel ridge regression under the initialization of both classical random Gaussian and leverage score sampling.

cs.LG↗

Recurrent Dirichlet Belief Networks for Interpretable Dynamic Relational Data Modelling

The Dirichlet Belief Network~(DirBN) has been recently proposed as a promising approach in learning interpretable deep latent representations for objects. In this work, we leverage its interpretable modelling architecture and propose a deep dynamic probabilistic framework -- the Recurrent Dirichlet Belief Network~(Recurrent-DBN) -- to study interpretable hidden structures from dynamic relational data. The proposed Recurrent-DBN has the following merits: (1) it infers interpretable and organised hierarchical latent structures for objects within and across time steps; (2) it enables recurrent long-term temporal dependence modelling, which outperforms the one-order Markov descriptions in most of the dynamic probabilistic frameworks. In addition, we develop a new inference strategy, which first upward-and-backward propagates latent counts and then downward-and-forward samples variables, to enable efficient Gibbs sampling for the Recurrent-DBN. We apply the Recurrent-DBN to dynamic relational data problems. The extensive experiment results on real-world data validate the advantages of the Recurrent-DBN over the state-of-the-art models in interpretable latent structure discovery and improved link prediction performance.

cs.LG↗

Fragmentation Coagulation Based Mixed Membership Stochastic Blockmodel

The Mixed-Membership Stochastic Blockmodel~(MMSB) is proposed as one of the state-of-the-art Bayesian relational methods suitable for learning the complex hidden structure underlying the network data. However, the current formulation of MMSB suffers from the following two issues: (1), the prior information~(e.g. entities' community structural information) can not be well embedded in the modelling; (2), community evolution can not be well described in the literature. Therefore, we propose a non-parametric fragmentation coagulation based Mixed Membership Stochastic Blockmodel (fcMMSB). Our model performs entity-based clustering to capture the community information for entities and linkage-based clustering to derive the group information for links simultaneously. Besides, the proposed model infers the network structure and models community evolution, manifested by appearances and disappearances of communities, using the discrete fragmentation coagulation process (DFCP). By integrating the community structure with the group compatibility matrix we derive a generalized version of MMSB. An efficient Gibbs sampling scheme with Polya Gamma (PG) approach is implemented for posterior inference. We validate our model on synthetic and real world data.

stat.ML↗

Scalable Lattice Influence Maximization

Influence maximization is the task of finding k seed nodes in a social network such that the expected number of activated nodes in the network (under certain influence propagation model), referred to as the influence spread, is maximized. Lattice influence maximization (LIM) generalizes influence maximization such that, instead of selecting k seed nodes, one selects a vector x = (x_1, ..., x_d) from a discrete space X called a lattice, where x_j corresponds to the j-th marketing strategy and x represents a marketing strategy mix. Each strategy mix x has probability h_u(x) to activate a node u as a seed.LIM is the task of finding a strategy mix under the constraint x_1+...+x_d <= k such that its influence spread is maximized. We adapt the reverse influence sampling (RIS) approach and design scalable algorithms for LIM. We first design the IMM-PRR algorithm based on partial reverse-reachable sets as a general solution for LIM, and improve IMM-PRR for a large family of models where each strategy independently activates seed nodes. We then propose an alternative algorithm IMM-VSN based on virtual strategy nodes, for the family of models with independent strategy activations. We prove that both IMM-PRR and IMM-VSN guarantees 1-e-εapproximation for small ε> 0. Empirically, through extensive tests we demonstrate that IMM-VSN runs faster than IMM-PRR and much faster than other baseline algorithms while providing the same level of influence spread. We conclude that IMM-VSN is the best one for models with independent strategy activations, while IMM-PRR works for general modes without this assumption. Finally, we extend LIM to the partitioned budget case where strategies are partitioned into groups, each of which has a separate budget, and show that a minor variation of our algorithms would achieve 1/2 -εapproximation ratio with the same time complexity.

cs.SI↗

Optical properties of dense lithium in electride phases by first-principles calculations

The metal-semiconductor-metal transition in dense lithium is considered as an archetype of interplay between interstitial electron localization and delocalization induced by compression, which leads to exotic electride phases. In this work, the dynamic dielectric response and optical properties of the high-pressure electride phases of cI16, oC40 and oC24 in lithium spanning a wide pressure range from 40 to 200 GPa by first-principles calculations are reported. Both interband and intraband contribution to the dielectric function are deliberately treated with the linear response theory. One intraband and two interband plasmons in cI16 at 70 GPa induced by a structural distortion at 2.1, 4.1, and 7.7 eV are discovered, which make the reflectivity of this weak metallic phase abnormally lower than the insulating phase oC40 at the corresponding frequencies. More strikingly, oC24 as a reentrant metallic phase with higher conductivity becomes more transparent than oC40 in infrared and visible light range due to its unique electronic structure around Fermi surface. An intriguing reflectivity anisotropy in both oC40 and oC24 is predicted, with the former being strong enough for experimental detection within the spectrum up to 10 eV. The important role of interstitial localized electrons is highlighted, revealing diversity and rich physics in electrides.

cond-mat.mtrl-sci↗

Decentralized RLS with Data-Adaptive Censoring for Regressions over Large-Scale Networks

The deluge of networked data motivates the development of algorithms for computation- and communication-efficient information processing. In this context, three data-adaptive censoring strategies are introduced to considerably reduce the computation and communication overhead of decentralized recursive least-squares (D-RLS) solvers. The first relies on alternating minimization and the stochastic Newton iteration to minimize a network-wide cost, which discards observations with small innovations. In the resultant algorithm, each node performs local data-adaptive censoring to reduce computations, while exchanging its local estimate with neighbors so as to consent on a network-wide solution. The communication cost is further reduced by the second strategy, which prevents a node from transmitting its local estimate to neighbors when the innovation it induces to incoming data is minimal. In the third strategy, not only transmitting, but also receiving estimates from neighbors is prohibited when data-adaptive censoring is in effect. For all strategies, a simple criterion is provided for selecting the threshold of innovation to reach a prescribed average data reduction. The novel censoring-based (C)D-RLS algorithms are proved convergent to the optimal argument in the mean-square deviation sense. Numerical experiments validate the effectiveness of the proposed algorithms in reducing computation and communication overhead.

eess.SY↗