SearcharxivSearch

arXiv subjects

Mingwei Dai

Publications and source records attributed to Mingwei Dai.

10 recordsLinked to original sources

Towards determination of the strong coupling $\alpha_s(m_Z)$ from four-flavor lattice QCD using the continuous $\beta$-function method

The precise value of the strong coupling $\alpha_s(m_{Z})$ at the $Z$-boson mass $m_{Z}$ is essential for high-energy phenomenology and precision tests of quantum chromodynamics (QCD). We present the status of a program targeting a $\sim 0.3\%$ determination of $\alpha_s(m_{Z})$ using the renormalization group $\beta$-function in the infinite volume gradient flow scheme based on lattice QCD simulations of degenerate four-flavor highly improved staggered quark (HISQ) ensembles. In particular, we analyze both tree-level cutoff effects and finite-mass effects. We also outline the next steps of the analysis, including the infinite-volume and continuum extrapolations required for a precise determination of $\alpha_s(m_Z)$.

hep-lat

Dynamical Dark Energy from Lattice Quantum Gravity

We study the behavior of the vacuum in Euclidean dynamical triangulations (EDT). Algorithmic improvements and better lattice spacing determinations allow us to test the properties of the emergent de Sitter geometries of our simulations to higher precision than previously possible. Although the agreement with de Sitter is good, the improved precision reveals deviations that can be interpreted as non-trivial vacuum dynamics, well-described by a cosmological constant that runs with scale. The simulations show that the dominant running is quadratic and that the scale can be identified with the Hubble rate. Several key cross-checks support this picture, including consistent results across multiple lattice spacings and the fact that the null energy condition is not violated. The parameters of the running are fully determined by simulations, enabling predictions when extrapolated to the scales relevant for our universe. This leads to a model for dark energy that is compatible with current observations, but which predicts deviations from $\Lambda$CDM at the ${\cal O}(10^{-3})$ level in cosmological observables that could be tested with future improvements in precision measurements.

hep-lat

An improved algorithm for dynamical triangulations and simulations of finer lattices

We introduce a new algorithm for the simulation of Euclidean dynamical triangulations that mimics the Metropolis-Hastings algorithm, but where all proposed moves are accepted. This rejection-free algorithm allows for the factorization of local and global terms in the action, a condition needed for efficient simulation of theories with global terms, while still maintaining detailed balance. We test our algorithm on the $2d$ Ising model, and against results for EDT obtained with standard Metropolis. Our new algorithm allows us to simulate EDT at finer lattice spacings than previously possible, and we find geometries that resemble semiclassical Euclidean de Sitter space in agreement with earlier results at coarser lattices. The agreement between lattice data and the classical de Sitter solution continues to get better as the lattice spacing decreases.

hep-lat

Nyström Regularization for Time Series Forecasting

This paper focuses on learning rate analysis of Nyström regularization with sequential sub-sampling for $τ$-mixing time series. Using a recently developed Banach-valued Bernstein inequality for $τ$-mixing sequences and an integral operator approach based on second-order decomposition, we succeed in deriving almost optimal learning rates of Nyström regularization with sequential sub-sampling for $τ$-mixing time series. A series of numerical experiments are carried out to verify our theoretical results, showing the excellent learning performance of Nyström regularization with sequential sub-sampling in learning massive time series data. All these results extend the applicable range of Nyström regularization from i.i.d. samples to non-i.i.d. sequences.

cs.LG

Newtonian Binding from Lattice Quantum Gravity

We study scalar fields propagating on Euclidean dynamical triangulations (EDT). In this work we study the interaction of two scalar particles, and we show that in the appropriate limit we recover an interaction compatible with Newton's gravitational potential in four dimensions. Working in the quenched approximation, we calculate the binding energy of a two-particle bound state, and we study its dependence on the constituent particle mass in the non-relativistic limit. We find a binding energy compatible with what one expects for the ground state energy by solving the Schrödinger equation for Newton's potential. Agreement with this expectation is obtained in the infinite-volume, continuum limit of the lattice calculation, providing non-trivial evidence that EDT is in fact a theory of gravity in four dimensions. Furthermore, this result allows us to determine the lattice spacing within an EDT calculation for the first time, and we find that the various lattice spacings are smaller than the Planck length, suggesting that we can achieve a separation of scales and that there is no obstacle to taking a continuum limit. This lends further support to the asymptotic safety scenario for gravity.

hep-lat

BOLT-SSI: A Statistical Approach to Screening Interaction Effects for Ultra-High Dimensional Data

Detecting interaction effects among predictors on the response variable is a crucial step in various applications. In this paper, we first propose a simple method for sure screening interactions (SSI). Although its computation complexity is $O(p^2n)$, SSI works well for problems of moderate dimensionality (e.g., $p=10^3\sim10^4$), without the heredity assumption. To ultra-high dimensional problems (e.g., $p = 10^6$), motivated by discretization associated Boolean representation and operations and the contingency table for discrete variables, we propose a fast algorithm, named "BOLT-SSI". The statistical theory has been established for SSI and BOLT-SSI, guaranteeing their sure screening property. The performance of SSI and BOLT-SSI are evaluated by comprehensive simulation and real case studies. Numerical results demonstrate that SSI and BOLT-SSI can often outperform their competitors in terms of computational efficiency and statistical accuracy. The proposed method can be applied for fully detecting interactions with more than 300,000 predictors. Based on this study, we believe that there is a great need to rethink the relationship between statistical accuracy and computational efficiency. We have shown that the computational performance of a statistical method can often be greatly improved by exploring the advantages of computational architecture with a tolerable loss of statistical accuracy.

stat.ME

Joint Analysis of Individual-level and Summary-level GWAS Data by Leveraging Pleiotropy

A large number of recent genome-wide association studies (GWASs) for complex phenotypes confirm the early conjecture for polygenicity, suggesting the presence of large number of variants with only tiny or moderate effects. However, due to the limited sample size of a single GWAS, many associated genetic variants are too weak to achieve the genome-wide significance. These undiscovered variants further limit the prediction capability of GWAS. Restricted access to the individual-level data and the increasing availability of the published GWAS results motivate the development of methods integrating both the individual-level and summary-level data. How to build the connection between the individual-level and summary-level data determines the efficiency of using the existing abundant summary-level resources with limited individual-level data, and this issue inspires more efforts in the existing area. In this study, we propose a novel statistical approach, LEP, which provides a novel way of modeling the connection between the individual-level data and summary-level data. LEP integrates both types of data by \underline{LE}veraing \underline{P}leiotropy to increase the statistical power of risk variants identification and the accuracy of risk prediction. The algorithm for parameter estimation is developed to handle genome-wide-scale data. Through comprehensive simulation studies, we demonstrated the advantages of LEP over the existing methods. We further applied LEP to perform integrative analysis of Crohn's disease from WTCCC and summary statistics from GWAS of some other diseases, such as Type 1 diabetes, Ulcerative colitis and Primary biliary cirrhosis. LEP was able to significantly increase the statistical power of identifying risk variants and improve the risk prediction accuracy from 63.39\% ($\pm$ 0.58\%) to 68.33\% ($\pm$ 0.32\%) using about 195,000 variants.

q-bio.GN

BIVAS: A scalable Bayesian method for bi-level variable selection with applications

In this paper, we consider a Bayesian bi-level variable selection problem in high-dimensional regressions. In many practical situations, it is natural to assign group membership to each predictor. Examples include that genetic variants can be grouped at the gene level and a covariate from different tasks naturally forms a group. Thus, it is of interest to select important groups as well as important members from those groups. The existing Markov Chain Monte Carlo (MCMC) methods are often computationally intensive and not scalable to large data sets. To address this problem, we consider variational inference for bi-level variable selection (BIVAS). In contrast to the commonly used mean-field approximation, we propose a hierarchical factorization to approximate the posterior distribution, by utilizing the structure of bi-level variable selection. Moreover, we develop a computationally efficient and fully parallelizable algorithm based on this variational approximation. We further extend the developed method to model data sets from multi-task learning. The comprehensive numerical results from both simulation studies and real data analysis demonstrate the advantages of BIVAS for variable selection, parameter estimation and computational efficiency over existing methods. The method is implemented in R package `bivas' available at https://github.com/mxcai/bivas.

stat.AP

LPG: a four-groups probabilistic approach to leveraging pleiotropy in genome-wide association studies

To date, genome-wide association studies (GWAS) have successfully identified tens of thousands of genetic variants among a variety of traits/diseases, shedding a light on the genetic architecture of complex diseases. Polygenicity of complex diseases, which refers to the phenomenon that a vast number of risk variants collectively contribute to the heritability of complex diseases with modest individual effects, have been widely accepted. This imposes a major challenge towards fully characterizing the genetic bases of complex diseases. An immediate implication of polygenicity is that a much larger sample size is required to detect risk variants with weak/moderate effects. Meanwhile, accumulating evidence suggests that different complex diseases can share genetic risk variants, a phenomenon known as pleiotropy. In this study, we propose a statistical framework for Leveraging Pleiotropic effects in large-scale GWAS data (LPG). LPG utilizes a variational Bayesian expectation-maximization (VBEM) algorithm, making it computationally efficient and scalable for genome-wide scale analysis. To demon- strate the advantage of LPG over existing methods that do not leverage pleiotropy, we conducted extensive simulation studies and also applied LPG to analyze three au- toimmune disorders (Crohn's disease, rheumatoid arthritis and Type 1 diabetes). The results indicate that LPG can improve the power of prioritization of risk variants and accuracy of risk prediction by leveraging pleiotropy. The software is available at http- s://github.com/Shufeyangyi2015310117/LPG.

stat.ME

LSMM: A statistical approach to integrating functional annotations with genome-wide association studies

Thousands of risk variants underlying complex phenotypes (quantitative traits and diseases) have been identified in genome-wide association studies (GWAS). However, there are still two major challenges towards deepening our understanding of the genetic architectures of complex phenotypes. First, the majority of GWAS hits are in the non-coding region and their biological interpretation is still unclear. Second, accumulating evidence from GWAS suggests the polygenicity of complex traits, i.e., a complex trait is often affected by many variants with small or moderate effects, whereas a large proportion of risk variants with small effects remains unknown. The availability of functional annotation data enables us to address the above challenges. In this study, we propose a latent sparse mixed model (LSMM) to integrate functional annotations with GWAS data. Not only does it increase statistical power of the identification of risk variants, but also offers more biological insights by detecting relevant functional annotations. To allow LSMM scalable to millions of variants and hundreds of functional annotations, we developed an efficient variational expectation-maximization (EM) algorithm for model parameter estimation and statistical inference. We first conducted comprehensive simulation studies to evaluate the performance of LSMM. Then we applied it to analyze 30 GWAS of complex phenotypes integrated with 9 genic category annotations and 127 tissue-specific functional annotations from the Roadmap project. The results demonstrate that our method possesses more statistical power over conventional methods, and can help researchers achieve deeper understanding of genetic architecture of these complex phenotypes.

stat.ME