SearcharxivSearch

arXiv subjects

Matt Challacombe

Publications and source records attributed to Matt Challacombe.

At least 19 recordsLinked to original sources

A N-Body Solver for Free Mesh Interpolation

Factorization of the Gaussian RBF kernel is developed for free-mesh interpolation in the flat, polynomial limit corresponding to Taylor expansion and the Vandermonde basis of geometric moments. With this spectral approximation, a top-down octree-scoping of an interpolant is found by recursively decomposing the residual, similar to the work of Driscoll and Heryudono (2007), except that in the current approach the grid is decoupled from the low rank approximation, allowing partial separation of sampling errors (the mesh) from representation errors (the polynomial order). Then, it is possible to demonstrate roughly 5 orders of magnitude improvement in free-mesh interpolation errors for the three-dimensional Franke function, relative to previous benchmarks. As in related work on $N$-body methods for factorization by square root iteration (Challacombe 2015), some emphasis is placed on resolution of the identity.

math.NA

A $N$-Body Solver for Square Root Iteration

We develop the Sparse Approximate Matrix Multiply ($\tt SpAMM$) $n$-body solver for first order Newton Schulz iteration of the matrix square root and inverse square root. The solver performs recursive two-sided metric queries on a modified Cauchy-Schwarz criterion, culling negligible sub-volumes of the product-tensor for problems with structured decay in the sub-space metric. These sub-structures are shown to bound the relative error in the matrix-matrix product, and in favorable cases, to enjoy a reduced computational complexity governed by dimensionality reduction of the product volume. A main contribution is demonstration of a new, algebraic locality that develops under contractive identity iteration, with collapse of the metric-subspace onto the identity's plane diagonal, resulting in a stronger $\tt SpAMM$ bound. Also, we carry out a first order {Fr\'{e}chet} analyses for single and dual channel instances of the square root iteration, and look at bifurcations due to ill-conditioning and a too aggressive $\tt SpAMM$ approximation. Then, we show that extreme $\tt SpAMM$ approximation and contractive identity iteration can be achieved for ill-conditioned systems through regularization, and we demonstrate the potential for acceleration with a scoping, product representation of the inverse factor.

math.NA

Solvers for $\mathcal{O} (N)$ Electronic Structure in the Strong Scaling Limit

We present a hybrid OpenMP/Charm++ framework for solving the $\mathcal{O} (N)$ Self-Consistent-Field eigenvalue problem with parallelism in the strong scaling regime, $P\gg{N}$, where $P$ is the number of cores, and $N$ a measure of system size, i.e. the number of matrix rows/columns, basis functions, atoms, molecules, etc. This result is achieved with a nested approach to Spectral Projection and the Sparse Approximate Matrix Multiply [Bock and Challacombe, SIAM J.~Sci.~Comput. 35 C72, 2013], and involves a recursive, task-parallel algorithm, often employed by generalized $N$-Body solvers, to occlusion and culling of negligible products in the case of matrices with decay. Employing classic technologies associated with generalized $N$-Body solvers, including over-decomposition, recursive task parallelism, orderings that preserve locality, and persistence-based load balancing, we obtain scaling beyond hundreds of cores per molecule for small water clusters ([H${}_2$O]${}_N$, $N \in \{ 30, 90, 150 \}$, $P/N \approx \{ 819, 273, 164 \}$) and find support for an increasingly strong scalability with increasing system size $N$.

math.NA

An N-Body Solution to the Problem of Fock Exchange

We report an N-Body approach to computing the Fock exchange matrix with and without permutational symmetry. The method achieves an O(N lg N) computational complexity through an embedded metric-query, allowing hierarchical application of direct SCF criteria. The advantages of permutational symmetry are found to be 4-fold for small systems, but decreasing with increasing system size and/or more permissive neglect criteria. This work sets the stage for: (1) the introduction of range queries in multi-level multipole schemes for rank reduction, and (2) recursive task parallelism.

cs.DS

An Optimized Sparse Approximate Matrix Multiply for Matrices with Decay

We present an optimized single-precision implementation of the Sparse Approximate Matrix Multiply (\SpAMM{}) [M. Challacombe and N. Bock, arXiv {\bf 1011.3534} (2010)], a fast algorithm for matrix-matrix multiplication for matrices with decay that achieves an $\mathcal{O} (n \log n)$ computational complexity with respect to matrix dimension $n$. We find that the max norm of the error achieved with a \SpAMM{} tolerance below $2 \times 10^{-8}$ is lower than that of the single-precision {\tt SGEMM} for dense quantum chemical matrices, while outperforming {\tt SGEMM} with a cross-over already for small matrices ($n \sim 1000$). Relative to naive implementations of \SpAMM{} using Intel's Math Kernel Library ({\tt MKL}) or AMD's Core Math Library ({\tt ACML}), our optimized version is found to be significantly faster. Detailed performance comparisons are made for quantum chemical matrices with differently structured sub-blocks. Finally, we discuss the potential of improved hardware prefetch to yield 2--3x speedups.

math.NA

Fast Multiplication of Matrices with Decay

A fast algorithm for the approximate multiplication of matrices with decay is introduced; the Sparse Approximate Matrix Multiply (SpAMM) reduces complexity in the product space, a different approach from current methods that economize within the matrix space through truncation or rank reduction. Matrix truncation (element dropping) is compared to SpAMM for quantum chemical matrices with approximate exponential and algebraic decay. For matched errors in the electronic total energy, SpAMM is found to require fewer to far fewer floating point operations relative to dropping. The challenges and opportunities afforded by this new approach are discussed, including the potential for high performance implementations.

cs.DS

Linear Scaling Solution of the Time-Dependent Self-Consistent-Field Equations

A new approach to solving the Time-Dependent Self-Consistent-Field equations is developed based on the double quotient formulation of Tsiper [J. Phys. B, 34 L401 (2001)]. Dual channel, quasi-independent non-linear optimization of these quotients is found to yield convergence rates approaching those of the best case (single channel) Tamm-Dancoff approximation. This formulation is variational with respect to matrix truncation, admitting linear scaling solution of the matrix-eigenvalue problem, which is demonstrated for bulk excitons in the polyphenylene vinylene oligimer and the (4,3) carbon nanotube segment.

quant-ph

Cache oblivious storage and access heuristics for blocked matrix-matrix multiplication

We investigate effects of ordering in blocked matrix--matrix multiplication. We find that submatrices do not have to be stored contiguously in memory to achieve near optimal performance. Instead it is the choice of execution order of the submatrix multiplications that leads to a speedup of up to four times for small block sizes. This is in contrast to results for single matrix elements showing that contiguous memory allocation quickly becomes irrelevant as the blocksize increases.

cs.DS

Molecular-orbital-free algorithm for excited states in time-dependent perturbation theory

A non-linear conjugate gradient optimization scheme is used to obtain excitation energies within the Random Phase Approximation (RPA). The solutions to the RPA eigenvalue equation are located through a variational characterization using a modified Thouless functional, which is based upon an asymmetric Rayleigh quotient, in an orthogonalized atomic orbital representation. In this way, the computational bottleneck of calculating molecular orbitals is avoided. The variational space is reduced to the physically-relevant transitions by projections. The feasibility of an RPA implementation scaling linearly with system size, N, is investigated by monitoring convergence behavior with respect to the quality of initial guess and sensitivity to noise under thresholding, both for well- and ill-conditioned problems. The molecular- orbital-free algorithm is found to be robust and computationally efficient providing a first step toward a large-scale, reduced complexity calculation of time-dependent optical properties and linear response. The algorithm is extensible to other forms of time-dependent perturbation theory including, but not limited to, time-dependent Density Functional theory.

cond-mat.mtrl-sci

Time-reversible Born-Oppenheimer molecular dynamics

We present a time-reversible Born-Oppenheimer molecular dynamics scheme, based on self-consistent Hartree-Fock or density functional theory, where both the nuclear and the electronic degrees of freedom are propagated in time. We show how a time-reversible adiabatic propagation of the electronic degrees of freedom is possible despite the non-linearity and incompleteness of the self-consistent field procedure. Time-reversal symmetry excludes a systematic long-term energy drift for a microcanonical ensemble and the number of self-consistency cycles can be kept low (often only 2-4 cycles per nuclear time step) thanks to a good initial guess given by the adiabatic propagation of the electronic degrees of freedom. The time-reversible Born-Oppenheimer molecular dynamics scheme therefore combines a low computational cost with a physically correct time-reversible representation of the dynamics, which preserves a detailed balance between propagation forwards and backwards in time.

cond-mat.mtrl-sci

Geometry Optimization of Crystals by the Quasi-Independent Curvilinear Coordinate Approximation

The quasi-independent curvilinear coordinate approximation (QUICCA) method [K. Németh and M. Challacombe, J. Chem. Phys. {\bf 121}, 2877, (2004)] is extended to the optimization of crystal structures. We demonstrate that QUICCA is valid under periodic boundary conditions, enabling simultaneous relaxation of the lattice and atomic coordinates, as illustrated by tight optimization of polyethylene, hexagonal boron-nitride, a (10,0) carbon-nanotube, hexagonal ice, quartz and sulfur at the $Γ$-point RPBE/STO-3G level of theory.

physics.chem-ph

Non-Orthogonal Density Matrix Perturbation Theory

Density matrix perturbation theory [Phys. Rev. Lett. Vol. 92, 193001 (2004)] provides an efficient framework for the linear scaling computation of response properties [Phys. Rev. Lett. Vol. 92, 193002 (2004)]. In this article, we generalize density matrix perturbation theory to include properties computed with a perturbation dependent non-orthogonal basis. Such properties include analytic derivatives of the energy with respect to nuclear displacement, as well as magnetic response computed with a field dependent basis. The non-orthogonal density matrix perturbation theory is developed in the context of recursive purification methods, which are briefly reviewed.

physics.comp-ph

Higher Order Response in ${\cal O}(N)$ by Perturbed Projection

Perturbed projection for linear scaling solution of the coupled-perturbed self-consistent-field equations [Weber, Niklasson and Challacombe, Phys. Rev.\ Lett. {\bf 92}, 193002 (2004)] is extended to the computation of higher order static response properties. Although generally applicable, perturbed projection is developed here in the context of the self-consistent first and second electric hyperpolarizabilities of three dimensional water clusters at the Hartree-Fock level of theory. Non-orthogonal, density matrix analogues of Wigner's $2n+1$ rule are given up to fourth order. Linear scaling and locality of the higher order response densities under perturbation by a global electric field are demonstrated.

cond-mat.mtrl-sci

Linear scaling computation of the Fock matrix. IX. Parallel computation of the Coulomb matrix

We present parallelization of a quantum-chemical tree-code [J. Chem. Phys. {\bf 106}, 5526 (1997)] for linear scaling computation of the Coulomb matrix. Equal time partition [J. Chem. Phys. {\bf 118}, 9128 (2003)] is used to load balance computation of the Coulomb matrix. Equal time partition is a measurement based algorithm for domain decomposition that exploits small variation of the density between self-consistent-field cycles to achieve load balance. Efficiency of the equal time partition is illustrated by several tests involving both finite and periodic systems. It is found that equal time partition is able to deliver 91 -- 98 % efficiency with 128 processors in the most time consuming part of the Coulomb matrix calculation. The current parallel quantum chemical tree code is able to deliver 63 -- 81% overall efficiency on 128 processors with fine grained parallelism (less than two heavy atoms per processor).

cond-mat.other

Linear scaling computation of the Fock matrix VII. Periodic Density Functional Theory at the $Γ$-point

Linear scaling quantum chemical methods for Density Functional Theory are extended to the condensed phase at the $Γ$-point. For the two-electron Coulomb matrix, this is achieved with a tree-code algorithm for fast Coulomb summation [J. Chem. Phys. {\bf 106}, 5526 (1997)], together with multipole representation of the crystal field [J. Chem. Phys. {\bf 107}, 10131 (1997)]. A periodic version of the hierarchical cubature algorithm [J. Chem. Phys. {\bf 113}, 10037 (2000)], which builds a telescoping adaptive grid for numerical integration of the exchange-correlation matrix, is shown to be efficient when the problem is posed as integration over the unit cell. Commonalities between the Coulomb and exchange-correlation algorithms are discussed, with an emphasis on achieving linear scaling through the use of modern data structures. With these developments, convergence of the $Γ$-point supercell approximation to the ${\bf k}$-space integration limit is demonstrated for MgO and NaCl. Linear scaling construction of the Fockian and control of error is demonstrated for RBLYP/6-21G* diamond up to 512 atoms.

cond-mat.mtrl-sci

Linear scaling computation of the Fock matrix. VIII. Periodic boundaries for exact exchange at the $Γ$-point

A translationally invariant formulation of the Hartree-Fock (HF) $Γ$-point approximation is presented. This formulation is achieved through introduction of the Minimum Image Convention (MIC) at the level of primitive two-electron integrals, and implemented in a periodic version of the ONX algorithm [J. Chem. Phys, {\bf 106} 9708 (1997)] for linear scaling computation of the exchange matrix. Convergence of the HF-MIC $Γ$-point model to the HF ${\bf k}$-space limit is demonstrated for fully periodic magnesium oxide, ice and diamond. Computation of the diamond lattice constant using the HF-MIC model together with the hybrid PBE0 density functional [Theochem, {\bf 493} 145 (1999)] yields $a_0=3.569$Åwith the 6-21G* basis set and a $3\times3\times3$ supercell. Linear scaling computation of the HF-MIC exchange matrix is demonstrated for diamond and ice in the condensed phase

cond-mat.mtrl-sci

A New View on Geometry Optimization: the Quasi-Independent Curvilinear Coordinate Approximation

This article presents a new and efficient alternative to well established algorithms for molecular geometry optimization. The new approach exploits the approximate decoupling of molecular energetics in a curvilinear internal coordinate system, allowing separation of the 3N-dimensional optimization problem into an O(N) set of quasi-independent one-dimensional problems. Each uncoupled optimization is developed by a weighted least squares fit of energy gradients in the internal coordinate system followed by extrapolation. In construction of the weights, only an implicit dependence on topologically connected internal coordinates is present. This new approach is competitive with the best internal coordinate geometry optimization algorithms in the literature and works well for large biological problems with complicated hydrogen bond networks and ligand binding motifs.

cond-mat.mtrl-sci

Ab initio linear scaling response theory: Electric polarizability by perturbed projection

A linear scaling method for calculation of the static {\em ab inito} response within self-consistent field theory is developed and applied to calculation of the static electric polarizability. The method is based on density matrix perturbation theory [Niklasson and Challacombe, cond-mat/0311591], obtaining response functions directly via a perturbative approach to spectral projection. The accuracy and efficiency of the linear scaling method is demonstrated for a series of three-dimensional water clusters at the RHF/6-31G** level of theory. Locality of the response under a global electric field perturbation is numerically demonstrated by approximate exponential decay of derivative density matrix elements.

cond-mat.mtrl-sci