Searcharxiv⌕ Search

arXiv subjects

Lin Lin

Publications and source records attributed to Lin Lin.

At least 289 records · Page 16Linked to original sources

SCDM-k: Localized orbitals for solids via selected columns of the density matrix

The recently developed selected columns of the density matrix (SCDM) method [J. Chem. Theory Comput. 11, 1463, 2015] is a simple, robust, efficient and highly parallelizable method for constructing localized orbitals from a set of delocalized Kohn-Sham orbitals for insulators and semiconductors with $Γ$ point sampling of the Brillouin zone. In this work we generalize the SCDM method to Kohn-Sham density functional theory calculations with k-point sampling of the Brillouin zone, which is needed for more general electronic structure calculations for solids. We demonstrate that our new method, called SCDM-k, is by construction gauge independent and is a natural way to describe localized orbitals. SCDM-k computes localized orbitals without the use of an optimization procedure, and thus does not suffer from the possibility of being trapped in a local minimum. Furthermore, the computational complexity of using SCDM-k to construct orthogonal and localized orbitals scales as O(N log N ) where N is the total number of k-points in the Brillouin zone. SCDM-k is therefore efficient even when a large number of k-points are used for Brillouin zone sampling. We demonstrate the numerical performance of SCDM-k using systems with model potentials in two and three dimensions.

physics.comp-ph↗

Chebyshev polynomial filtered subspace iteration in the Discontinuous Galerkin method for large-scale electronic structure calculations

The Discontinuous Galerkin (DG) electronic structure method employs an adaptive local basis (ALB) set to solve the Kohn-Sham equations of density functional theory (DFT) in a discontinuous Galerkin framework. The adaptive local basis is generated on-the-fly to capture the local material physics, and can systematically attain chemical accuracy with only a few tens of degrees of freedom per atom. A central issue for large-scale calculations, however, is the computation of the electron density (and subsequently, ground state properties) from the discretized Hamiltonian in an efficient and scalable manner. We show in this work how Chebyshev polynomial filtered subspace iteration (CheFSI) can be used to address this issue and push the envelope in large-scale materials simulations in a discontinuous Galerkin framework. We describe how the subspace filtering steps can be performed in an efficient and scalable manner using a two-dimensional parallelization scheme, thanks to the orthogonality of the DG basis set and block-sparse structure of the DG Hamiltonian matrix. The on-the-fly nature of the ALBs requires additional care in carrying out the subspace iterations. We demonstrate the parallel scalability of the DG-CheFSI approach in calculations of large-scale two-dimensional graphene sheets and bulk three-dimensional lithium-ion electrolyte systems. Employing 55,296 computational cores, the time per self-consistent field iteration for a sample of the bulk 3D electrolyte containing 8,586 atoms is 90 seconds, and the time for a graphene sheet containing 11,520 atoms is 75 seconds.

physics.comp-ph↗

Magnetar-like X-ray Bursts from a Rotation Powered Pulsar, PSR J1119-6127

Two energetic hard X-ray bursts have recently triggered the Fermi and Swift space observatories from the rotation powered pulsar, PSR J1119-6127. We have performed in depth spectral and temporal analyses of these two events. Our extensive searches in both observatory data for lower luminosity bursts uncovered 10 additional events from the source. We report here on the timing and energetics of the 12 bursts from PSR J1119-6127 during its burst active phase of 2016 July 26 and 28. We also found a spectral softer X-ray flux enhancement in a post burst episode, which shows evidence of cooling. We discuss here the implications of these results on the nature of this unusual high-field radio pulsar, which firmly place it within the typical magnetar population.

astro-ph.HE↗

Electron Bunch Train Excited Higher-Order Modes in a Superconducting RF Cavity

Higher-order mode (HOM) based intra-cavity beam diagnostics has been proved effectively and conveniently in superconducting radio-frequency (SRF) accelerators. Our recent research shows that the beam harmonics in the bunch train excited HOM spectrum, which have much higher signal-to-noise ratio than the intrinsic HOM peaks, may also be useful for beam diagnostics. In this paper, we will present our study on bunch train excited HOMs, including the theoretic model and recent experiments carried out based on the DC-SRF photoinjector and SRF linac at Peking University.

physics.acc-ph↗

Hidden Markov Models on Variable Blocks with a Modal Clustering Algorithm and Applications

Motivated by high-throughput single-cell cytometry data with applications to vaccine development and immunological research, we consider statistical clustering in large-scale data that contain multiple rare clusters. We propose a new hierarchical mixture model, namely Hidden Markov Model on Variable Blocks (HMM-VB), and a new mode search algorithm called Modal Baum-Welch (MBW) for efficient clustering. Exploiting the widely accepted chain-like dependence among groups of variables in the cytometry data, we propose to treat the hierarchy of variable groups as a figurative time line and employ a HMM-type model, namely HMM-VB. We also propose to use mode-based clustering, aka modal clustering, and overcome the exponential computational complexity by MBW. In a series of experiments on simulated data HMM-VB and MBW have better performance than existing methods. We also apply our method to identify rare cell subsets in cytometry data and examine its strengths and limitations.

stat.ME↗

Decay estimates of discretized Green's functions for Schrödinger type operators

For a sparse non-singular matrix $A$, generally $A^{-1}$ is a dense matrix. However, for a class of matrices, $A^{-1}$ can be a matrix with off-diagonal decay properties, i.e. $\lvert A^{-1}_{ij}\rvert$ decays fast to $0$ with respect to the increase of a properly defined distance between $i$ and $j$. Here we consider the off-diagonal decay properties of discretized Green's functions for Schrödinger type operators. We provide decay estimates for discretized Green's functions obtained from the finite difference discretization, and from a variant of the pseudo-spectral discretization. The asymptotic decay rate in our estimate is independent of the domain size and of the discretization parameter. We verify the decay estimate with numerical results for one-dimensional Schrödinger type operators.

math.NA↗

Adaptively compressed polarizability operator for accelerating large scale \textit{ab initio} phonon calculations

Phonon calculations based on first principle electronic structure theory, such as the Kohn-Sham density functional theory, have wide applications in physics, chemistry and material science. The computational cost of first principle phonon calculations typically scales steeply as $\mathcal{O}(N_e^4)$, where $N_e$ is the number of electrons in the system. In this work, we develop a new method to reduce the computational complexity of computing the full dynamical matrix, and hence the phonon spectrum, to $\mathcal{O}(N_e^3)$. The key concept for achieving this is to compress the polarizability operator adaptively with respect to the perturbation of the potential due to the change of the atomic configuration. Such adaptively compressed polarizability operator (ACP) allows accurate computation of the phonon spectrum. The reduction of complexity only weakly depends on the size of the band gap, and our method is applicable to insulators as well as semiconductors with small band gaps. We demonstrate the effectiveness of our method using one-dimensional and two-dimensional model problems.

math.NA↗

Low Rank Approximation in $G_0W_0$ Approximation

The single particle energies obtained in a Kohn--Sham density functional theory (DFT) calculation are generally known to be poor approximations to electron excitation energies that are measured in transport, tunneling and spectroscopic experiments such as photo-emission spectroscopy. The correction to these energies can be obtained from the poles of a single particle Green's function derived from a many-body perturbation theory. From a computational perspective, the accuracy and efficiency of such an approach depends on how a self energy term that properly accounts for dynamic screening of electrons is approximated. The $G_0W_0$ approximation is a widely used technique in which the self energy is expressed as the convolution of a non-interacting Green's function ($G_0$) and a screened Coulomb interaction ($W_0$) in the frequency domain. The computational cost associated with such a convolution is high due to the high complexity of evaluating $W_0$ at multiple frequencies. In this paper, we discuss how the cost of $G_0W_0$ calculation can be reduced by constructing a low rank approximation to the frequency dependent part of $W_0$. In particular, we examine the effect of such a low rank approximation on the accuracy of the $G_0W_0$ approximation. We also discuss how the numerical convolution of $G_0$ and $W_0$ can be evaluated efficiently and accurately by using a contour deformation technique with an appropriate choice of the contour.

math.NA↗

A Left-Looking Selected Inversion Algorithm and Task Parallelism on Shared Memory Systems

Given a sparse matrix $A$, the selected inversion algorithm is an efficient method for computing certain selected elements of $A^{-1}$. These selected elements correspond to all or some nonzero elements of the LU factors of $A$. In many ways, the type of matrix updates performed in the selected inversion algorithm is similar to that performed in the LU factorization, although the sequence of operation is different. In the context of LU factorization, it is known that the left-looking and right-looking algorithms exhibit different memory access and data communication patterns, and hence different behavior on shared memory and distributed memory parallel machines. Corresponding to right-looking and left-looking LU factorization, selected inversion algorithm can be organized as a left-looking and a right-looking algorithm. The parallel right-looking version of the algorithm has been developed in [1]. The sequence of operations performed in this version of the selected inversion algorithm is similar to those performed in a left-looking LU factorization algorithm. In this paper, we describe the left-looking variant of the selected inversion algorithm, and based on task parallel method, present an efficient implementation of the algorithm for shared memory machines. We demonstrate that with the task scheduling features provided by OpenMP 4.0, the left-looking selected inversion algorithm can scale well both on the Intel Haswell multicore architecture and on the Intel Knights Corner (KNC) manycore architecture. Compared to the right-looking selected inversion algorithm, the left-looking formulation facilitates pipelining of work along different branches of the elimination tree, and can be a promising candidate for future development of massively parallel selected inversion algorithms on heterogeneous architecture.

cs.MS↗

A posteriori error estimates for discontinuous Galerkin methods using non-polynomial basis functions. Part II: Eigenvalue problems

We present the first systematic work for deriving a posteriori error estimates for general non-polynomial basis functions in an interior penalty discontinuous Galerkin (DG) formulation for solving eigenvalue problems associated with second order linear operators. Eigenvalue problems of such types play important roles in scientific and engineering applications, particularly in theoretical chemistry, solid state physics and material science. Based on the framework developed in [{\it L. Lin, B. Stamm, http://dx.doi.org/10.1051/m2an/2015069}] for second order PDEs, we develop residual type upper and lower bound error estimates for measuring the a posteriori error for eigenvalue problems. The main merit of our method is that the method is parameter-free, in the sense that all but one solution-dependent constants appearing in the upper and lower bound estimates are explicitly computable by solving local and independent eigenvalue problems, and the only non-computable constant can be reasonably approximated by a computable one without affecting the overall effectiveness of the estimates in practice. Compared to the PDE case, we find that a posteriori error estimators for eigenvalue problems must neglect certain terms, which involves explicitly the exact eigenvalues or eigenfunctions that are not accessible in numerical simulations. We define such terms carefully, and justify numerically that the neglected terms are indeed numerically high order terms compared to the computable estimators. Numerical results for a variety of problems in 1D and 2D demonstrate that both the upper bound and lower bound are effective for measuring the error of eigenvalues and eigenfunctions.

math.NA↗

Adaptively Compressed Exchange Operator

The Fock exchange operator plays a central role in modern quantum chemistry. The large computational cost associated with the Fock exchange operator hinders Hartree-Fock calculations and Kohn-Sham density functional theory calculations with hybrid exchange-correlation functionals, even for systems consisting of hundreds of atoms. We develop the adaptively compressed exchange operator (ACE) formulation, which greatly reduces the computational cost associated with the Fock exchange operator without loss of accuracy. The ACE formulation does not depend on the size of the band gap, and thus can be applied to insulating, semiconducting as well as metallic systems. In an iterative framework for solving Hartree-Fock-like systems, the ACE formulation only requires moderate modification of the code, and can be potentially beneficial for all electronic structure software packages involving exchange calculations. Numerical results indicate that the ACE formulation can become advantageous even for small systems with tens of atoms. In particular, the cost of each self-consistent field iteration for the electron density in the ACE formulation is only marginally larger than that of the generalized gradient approximation (GGA) calculation, and thus offers orders of magnitude speedup for Hartree-Fock-like calculations.

physics.comp-ph↗

Localized spectrum slicing

Given a sparse Hermitian matrix $A$ and a real number $μ$, we construct a set of sparse vectors, each approximately spanned only by eigenvectors of $A$ corresponding to eigenvalues near $μ$. This set of vectors spans the column space of a localized spectrum slicing (LSS) operator, and is called an LSS basis set. The sparsity of the LSS basis set is related to the decay properties of matrix Gaussian functions. We present a divide-and-conquer strategy with controllable error to construct the LSS basis set. This is a purely algebraic process using only submatrices of $A$, and can therefore be applied to general sparse Hermitian matrices. The LSS basis set leads to sparse projected matrices with reduced sizes, which allows the projected problems to be solved efficiently with techniques using sparse linear algebra. As an example, we demonstrate that the LSS basis set can be used to solve interior eigenvalue problems for a discretized second order partial differential operator in one-dimensional and two-dimensional domains, as well as for a matrix of general sparsity pattern.

math.NA↗

Randomized estimation of spectral densities of large matrices made accurate

For a large Hermitian matrix $A\in \mathbb{C}^{N\times N}$, it is often the case that the only affordable operation is matrix-vector multiplication. In such case, randomized method is a powerful way to estimate the spectral density (or density of states) of $A$. However, randomized methods developed so far for estimating spectral densities only extract information from different random vectors independently, and the accuracy is therefore inherently limited to $\mathcal{O}(1/\sqrt{N_{v}})$ where $N_{v}$ is the number of random vectors. In this paper we demonstrate that the "$\mathcal{O}(1/\sqrt{N_{v}})$ barrier" can be overcome by taking advantage of the correlated information of random vectors when properly filtered by polynomials of $A$. Our method uses the fact that the estimation of the spectral density essentially requires the computation of the trace of a series of matrix functions that are numerically low rank. By repeatedly applying $A$ to the same set of random vectors and taking different linear combination of the results, we can sweep through the entire spectrum of $A$ by building such low rank decomposition at different parts of the spectrum. Under some assumptions, we demonstrate that a robust and efficient implementation of such spectrum sweeping method can compute the spectral density accurately with $\mathcal{O}(N^2)$ computational cost and $\mathcal{O}(N)$ memory cost. Numerical results indicate that the new method can significantly outperform existing randomized methods in terms of accuracy. As an application, we demonstrate a way to accurately compute a trace of a smooth matrix function, by carefully balancing the smoothness of the integrand and the regularized density of states using a deconvolution procedure.

math.NA↗

Experimental observation of magnetoelectricity in spin ice Dy$_2$Ti$_2$O$_7$

The intrinsic noncollinear spin patterns in rare-earth pyrochlore are physically interesting, hosting many emergent properties, e.g. spin ice and monopole-type excitation. Recently, the magnetic monopole excitation of spin ice systems was predicted to be magnetoelectric active, while rare experimental works have directly confirmed this scenario. In this work, we performed systematic experimental investigation on the magnetoelectricity of Dy$_2$Ti$_2$O$_7$ by probing the ferroelectricity, spin dynamics, and dielectric behaviors. Two ferroelectric transitions at $T_{c1}$=25 K and $T_{c2}$=13 K have been observed. Remarkable magnetoelectric coupling is identified below the lower transition temperature, with a significant suppression of the electric polarization upon applied magnetic field. It is surprised that the lower ferroelectric transition temperature just coincides with the Ising-spin paramagnetic transition point, below which the quasi-particle-like monopoles are populated, indicating implicit correlation between electric dipoles and spin moments. The possible magnetoelectric mechanisms have also been discussed although a decent theory remains unavailable up to date. Our results will stimulate more investigations to explore multiferroicity in these spin ice systems and other frustrated magnets.

cond-mat.mtrl-sci↗

Adaptive local basis set for Kohn-Sham density functional theory in a discontinuous Galerkin framework II: Force, vibration, and molecular dynamics calculations

Recently, we have proposed the adaptive local basis set for electronic structure calculations based on Kohn-Sham density functional theory in a pseudopotential framework. The adaptive local basis set is efficient and systematically improvable for total energy calculations. In this paper, we present the calculation of atomic forces, which can be used for a range of applications such as geometry optimization and molecular dynamics simulation. We demonstrate that, under mild assumptions, the computation of atomic forces can scale nearly linearly with the number of atoms in the system using the adaptive local basis set. We quantify the accuracy of the Hellmann-Feynman forces for a range of physical systems, benchmarked against converged planewave calculations, and find that the adaptive local basis set is efficient for both force and energy calculations, requiring at most a few tens of basis functions per atom to attain accuracy required in practice. Since the adaptive local basis set has implicit dependence on atomic positions, Pulay forces are in general nonzero. However, we find that the Pulay force is numerically small and systematically decreasing with increasing basis completeness, so that the Hellmann-Feynman force is sufficient for basis sizes of a few tens of basis functions per atom. We verify the accuracy of the computed forces in static calculations of quasi-1D and 3D disordered Si systems, vibration calculation of a quasi-1D Si system, and molecular dynamics calculations of H$_2$ and liquid Al-Si alloy systems, where we find excellent agreement with independent benchmark results in literature.

physics.comp-ph↗

Plasma cleaning: A new possible treatment for niobium superconducting cavity after nitrogen doping

Nitrogen doping treatment with the subsequent electropolishing (EP) of the niobium superconducting cavity can significantly increase the cavity's quality factor up to a factor of 3. But the process of the EP removal may reintroduce hydrogen in the cavity surface, which may influence the cavity's radio frequency performance. Plasma cleaning study on niobium samples with gas mixtures of argon and oxgen intended to remove contaminations (hydrocarbons and micronicdust particles) from cavity surface to avoid field emission, was performed in Peking University. The niobium samples have been analyzed using the time of flight secondary ion mass spectrometry (TOF-SIMS) to measure the depth profiles of H, C, O, F, P and Nb. The measuring results show that the plasma cleaning with gas mixtures of argon and oxgen and conditions of about 20Pa and 100W can remarkably reduce the contents of impurity elements in the depth of about 30 nm without introducing hydrogen in the cavity surface. So plasma cleaning has been proposed to be a new possible treatment for niobium superconducting cavity after nitrogen doping.

physics.acc-ph↗

DGDFT: A Massively Parallel Method for Large Scale Density Functional Theory Calculations

We describe a massively parallel implementation of the recently developed discontinuous Galerkin density functional theory (DGDFT) [J. Comput. Phys. 2012, 231, 2140] method, for efficient large-scale Kohn-Sham DFT based electronic structure calculations. The DGDFT method uses adaptive local basis (ALB) functions generated on-the-fly during the self-consistent field (SCF) iteration to represent the solution to the Kohn-Sham equations. The use of the ALB set provides a systematic way to improve the accuracy of the approximation. It minimizes the number of degrees of freedom required to represent the solution to the Kohn-Sham problem for a desired level of accuracy. In particular, DGDFT can reach the planewave accuracy with far fewer numbers of degrees of freedom. By using the pole expansion and selected inversion (PEXSI) technique to compute electron density, energy and atomic forces, we can make the computational complexity of DGDFT scale at most quadratically with respect to the number of electrons for both insulating and metallic systems. We show that DGDFT can achieve 80% parallel efficiency on 128,000 high performance computing cores when it is used to study the electronic structure of two-dimensional (2D) phosphorene systems with 3,500-14,000 atoms. This high parallel efficiency results from a two-level parallelization scheme that we will describe in detail.

physics.comp-ph↗

A posteriori error estimates for discontinuous Galerkin methods using non-polynomial basis functions. Part I: Second order linear PDE

We present the first systematic work for deriving a posteriori error estimates for general non-polynomial basis functions in an interior penalty discontinuous Galerkin (DG) formulation for solving second order linear PDEs. Our residual type upper and lower bound error estimates measure the error in the energy norm. The main merit of our method is that the method is parameter-free, in the sense that all but one solution-dependent constants appearing in the upper and lower bound estimates are explicitly computable by solving local eigenvalue problems, and the only non-computable constant can be reasonably approximated by a computable one without affecting the overall effectiveness of the estimates in practice. As a side product of our formulation, the penalty parameter in the interior penalty formulation can be automatically determined as well. We develop an efficient numerical procedure to compute the error estimators. Numerical results for a variety of problems in 1D and 2D demonstrate that both the upper bound and lower bound are effective.

math.NA↗