SearcharxivSearch

arXiv subjects

Tilo Wettig

Publications and source records attributed to Tilo Wettig.

At least 19 recordsLinked to original sources

Order-separated tensor-network method for QCD in the strong-coupling expansion

We introduce the order-separated Grassmann higher-order tensor renormalization group (OS-GHOTRG) method for QCD with staggered quarks in the strong-coupling expansion. The method allows us to determine the expansion coefficients of the partition function, from which we can obtain the strong-coupling expansions of thermodynamical observables. We use the method in two dimensions to compute the free energy, the particle-number density, and the chiral condensate as a function of the chemical potential up to third order in the inverse coupling $\beta$. Although near the phase transition the expansion is only a good approximation to the full theory at small $\beta$, we show that the range of applicability can be greatly extended by fits to judiciously chosen transition functions.

hep-lat

A novel gauge-equivariant neural-network architecture for preconditioners in lattice QCD

Lattice QCD simulations are computationally expensive, with the solution of the Dirac equation being the major computational bottleneck of many calculations. We introduce a novel gauge-equivariant neural-network architecture for preconditioning the Dirac equation in the regime where critical slowing down occurs. We study the behavior of this preconditioner as a function of topological charge and lattice volume and show that it mitigates critical slowing down. We also show that this preconditioner transfers to unseen gauge configurations without any retraining, therefore enabling applications not possible with competing methods.

hep-lat

Tensor-network formulation of QCD in the strong-coupling expansion

We present a tensor-network formulation for the strong-coupling expansion of QCD with staggered quarks at nonzero chemical potential, for arbitrary number of dimensions, colors, and flavors. We integrate out the gauge and quark degrees of freedom and rewrite the partition function as the complete trace of a tensor network. This network consists of local tensors that contain a numerical and a Grassmann part. We truncate the initial tensor at a fixed order in the inverse coupling $\beta$ and compute analytical results for the partition function, the free energy, and the chiral condensate on a $2\times2$ lattice up to order $\beta^4$. In a follow-up paper we will introduce an enhanced tensor-network method, order-separated GHOTRG, to explicitly compute the expansion coefficients of the partition function for larger lattices. To demonstrate its potential, first results obtained with this new method are already presented here.

hep-lat

Combined track finding with GNN & CKF

The application of Graph Neural Networks (GNN) in track reconstruction is a promising approach to cope with the challenges arising at the High-Luminosity upgrade of the Large Hadron Collider (HL-LHC). GNNs show good track-finding performance in high-multiplicity scenarios and are naturally parallelizable on heterogeneous compute architectures. Typical high-energy-physics detectors have high resolution in the innermost layers to support vertex reconstruction but lower resolution in the outer parts. GNNs mainly rely on 3D space-point information, which can cause reduced track-finding performance in the outer regions. In this contribution, we present a novel combination of GNN-based track finding with the classical Combinatorial Kalman Filter (CKF) algorithm to circumvent this issue: The GNN resolves the track candidates in the inner pixel region, where 3D space points can represent measurements very well. These candidates are then picked up by the CKF in the outer regions, where the CKF performs well even for 1D measurements. Using the ACTS infrastructure, we present a proof of concept based on truth tracking in the pixels as well as a dedicated GNN pipeline trained on $t\bar{t}$ events with pile-up 200 in the OpenDataDetector.

hep-ex

Taming numerical imprecision by adapting the KL divergence to negative probabilities

The Kullback-Leibler (KL) divergence is frequently used in data science. For discrete distributions on large state spaces, approximations of probability vectors may result in a few small negative entries, rendering the KL divergence undefined. We address this problem by introducing a parameterized family of substitute divergence measures, the shifted KL (sKL) divergence measures. Our approach is generic and does not increase the computational overhead. We show that the sKL divergence shares important theoretical properties with the KL divergence and discuss how its shift parameters should be chosen. If Gaussian noise is added to a probability vector, we prove that the average sKL divergence converges to the KL divergence for small enough noise. We also show that our method solves the problem of negative entries in an application from computational oncology, the optimization of Mutual Hazard Networks for cancer progression using tensor-train approximations.

stat.CO

Gauge-equivariant pooling layers for preconditioners in lattice QCD

We demonstrate that gauge-equivariant pooling and unpooling layers can perform as well as traditional restriction and prolongation layers in multigrid preconditioner models for lattice QCD. These layers introduce a gauge degree of freedom on the coarse grid, allowing for the use of explicitly gauge-equivariant layers on the coarse grid. We investigate the construction of coarse-grid gauge fields and study their efficiency in the preconditioner model. We show that a combined multigrid neural network using a Galerkin construction for the coarse-grid gauge field eliminates critical slowing down.

hep-lat

Provenance for Lattice QCD workflows

We present a provenance model for the generic workflow of numerical Lattice Quantum Chromodynamics (QCD) calculations, which constitute an important component of particle physics research. These calculations are carried out on the largest supercomputers worldwide with data in the multi-PetaByte range being generated and analyzed. In the Lattice QCD community, a custom metadata standard (QCDml) that includes certain provenance information already exists for one part of the workflow, the so-called generation of configurations. In this paper, we follow the W3C PROV standard and formulate a provenance model that includes both the generation part and the so-called measurement part of the Lattice QCD workflow. We demonstrate the applicability of this model and show how the model can be used to answer some provenance-related research questions. However, many important provenance questions in the Lattice QCD community require extensions of this provenance model. To this end, we propose a multi-layered provenance approach that combines prospective and retrospective elements.

hep-lat

Gauge-equivariant neural networks as preconditioners in lattice QCD

We demonstrate that a state-of-the art multi-grid preconditioner can be learned efficiently by gauge-equivariant neural networks. We show that the models require minimal re-training on different gauge configurations of the same gauge ensemble and to a large extent remain efficient under modest modifications of ensemble parameters. We also demonstrate that important paradigms such as communication avoidance are straightforward to implement in this framework.

hep-lat

MRHS multigrid solver for Wilson-clover fermions

We describe our implementation of a multigrid solver for Wilson-clover fermions, which increases parallelism by solving for multiple right-hand sides (MRHS) simultaneously. The solver is based on Grid and thus runs on all computing architectures supported by the Grid framework. We present detailed benchmarks of the relevant kernels, such as hopping and clover term on the various multigrid levels, intergrid operators, and reductions. The benchmarks were performed on the JUWELS Booster system at Jülich Supercomputing Centre, which is based on Nvidia A100 GPUs. For example, solving a $24^3\times128$ lattice on 16 GPUs, the overall speedup obtained solely from MRHS is about 10x.

hep-lat

Approximation formula for complex spacing ratios in the Ginibre ensemble

Recently, Sá, Ribeiro and Prosen introduced complex spacing ratios to analyze eigenvalue correlations in non-Hermitian systems. At present there are no analytical results for the probability distribution of these ratios in the limit of large system size. We derive an approximation formula for the Ginibre universality class of random matrix theory which converges exponentially fast to the limit of infinite matrix size. We also give results for moments of the distribution in this limit.

cond-mat.stat-mech

Differentiated uniformization: A new method for inferring Markov chains on combinatorial state spaces including stochastic epidemic models

Motivation: We consider continuous-time Markov chains that describe the stochastic evolution of a dynamical system by a transition-rate matrix $Q$ which depends on a parameter $θ$. Computing the probability distribution over states at time $t$ requires the matrix exponential $\exp(tQ)$, and inferring $θ$ from data requires its derivative $\partial\exp\!(tQ)/\partialθ$. Both are challenging to compute when the state space and hence the size of $Q$ is huge. This can happen when the state space consists of all combinations of the values of several interacting discrete variables. Often it is even impossible to store $Q$. However, when $Q$ can be written as a sum of tensor products, computing $\exp(tQ)$ becomes feasible by the uniformization method, which does not require explicit storage of $Q$. Results: Here we provide an analogous algorithm for computing $\partial\exp\!(tQ)/\partialθ$, the differentiated uniformization method. We demonstrate our algorithm for the stochastic SIR model of epidemic spread, for which we show that $Q$ can be written as a sum of tensor products. We estimate monthly infection and recovery rates during the first wave of the COVID-19 pandemic in Austria and quantify their uncertainty in a full Bayesian analysis. Availability: Implementation and data are available at https://github.com/spang-lab/TenSIR.

stat.ML

Grid on QPACE 4

In 2020 we deployed QPACE 4, which features 64 Fujitsu A64FX model FX700 processors interconnected by InfiniBand EDR. QPACE 4 runs an open-source software stack. For Lattice QCD simulations we ported the Grid LQCD framework to support the ARM Scalable Vector Extension (SVE). In this contribution we discuss our SVE port of Grid, the status of SVE compilers and the performance of Grid. We also present the benefits of an alternative data layout of complex numbers for the Domain Wall operator.

hep-lat

Complex spacing ratios of the non-Hermitian Dirac operator in universality classes AI$^\dagger$ and AII$^\dagger$

We consider non-Hermitian Dirac operators in QCD-like theories coupled to a chiral U(1) potential or an imaginary chiral chemical potential. We show that in the continuum they fall into the recently discovered universality classes AI$^\dagger$ or AII$^\dagger$ of random matrix theory if the fermions transform in pseudoreal or real representations of the gauge group, respectively. For staggered fermions on the lattice this correspondence is reversed. We verify our predictions by computing spacing ratios of complex eigenvalues, whose distribution is universal without the need for unfolding.

hep-lat

Machine learning for surface prediction in ACTS

We present an ongoing R&D activity for machine-learning-assisted navigation through detectors to be used for track reconstruction. We investigate different approaches of training neural networks for surface prediction and compare their results. This work is carried out in the context of the ACTS tracking toolkit.

physics.ins-det

ECM modeling and performance tuning of SpMV and Lattice QCD on A64FX

The A64FX CPU is arguably the most powerful Arm-based processor design to date. Although it is a traditional cache-based multicore processor, its peak performance and memory bandwidth rival accelerator devices. A good understanding of its performance features is of paramount importance for developers who wish to leverage its full potential. We present an architectural analysis of the A64FX used in the Fujitsu FX1000 supercomputer at a level of detail that allows for the construction of Execution-Cache-Memory (ECM) performance models for steady-state loops. In the process we identify architectural peculiarities that point to viable generic optimization strategies. After validating the model using simple streaming loops we apply the insight gained to sparse matrix-vector multiplication (SpMV) and the domain wall (DW) kernel from quantum chromodynamics (QCD). For SpMV we show why the CRS matrix storage format is not a good practical choice on this architecture and how the SELL-C-sigma format can achieve bandwidth saturation. For the DW kernel we provide a cache-reuse analysis and show how an appropriate choice of data layout for complex arrays can realize memory-bandwidth saturation in this case as well. A comparison with state-of-the-art high-end Intel Cascade Lake AP and Nvidia V100 systems puts the capabilities of the A64FX into perspective. We also explore the potential for power optimizations using the tuning knobs provided by the Fugaku system, achieving energy savings of about 31% for SpMV and 18% for DW.

cs.PF

New universality classes of the non-Hermitian Dirac operator in QCD-like theories

In non-Hermitian random matrix theory there are three universality classes for local spectral correlations: the Ginibre class and the nonstandard classes $\mathrm{AI}^\dagger$ and $\mathrm{AII}^\dagger$. We show that the continuum Dirac operator in two-color QCD coupled to a chiral $\mathrm{U}(1)$ gauge field or an imaginary chiral chemical potential falls in class $\mathrm{AI}^\dagger$ ($\mathrm{AII}^\dagger$) for fermions in pseudoreal (real) representations of $\mathrm{SU}(2)$. We introduce the corresponding chiral random matrix theories and verify our predictions in lattice simulations with staggered fermions, for which the correspondence between representation and universality class is reversed. Specifically, we compute the complex eigenvalue spacing ratios introduced recently. We also derive novel spectral sum rules.

hep-th

Performance Modeling of Streaming Kernels and Sparse Matrix-Vector Multiplication on A64FX

The A64FX CPU powers the current number one supercomputer on the Top500 list. Although it is a traditional cache-based multicore processor, its peak performance and memory bandwidth rival accelerator devices. Generating efficient code for such a new architecture requires a good understanding of its performance features. Using these features, we construct the Execution-Cache-Memory (ECM) performance model for the A64FX processor in the FX700 supercomputer and validate it using streaming loops. We also identify architectural peculiarities and derive optimization hints. Applying the ECM model to sparse matrix-vector multiplication (SpMV), we motivate why the CRS matrix storage format is inappropriate and how the SELL-C-sigma format with suitable code optimizations can achieve bandwidth saturation for SpMV.

cs.PF

Low-rank tensor methods for Markov chains with applications to tumor progression models

Continuous-time Markov chains describing interacting processes exhibit a state space that grows exponentially in the number of processes. This state-space explosion renders the computation or storage of the time-marginal distribution, which is defined as the solution of a certain linear system, infeasible using classical methods. We consider Markov chains whose transition rates are separable functions, which allows for an efficient low-rank tensor representation of the operator of this linear system. Typically, the right-hand side also has low-rank structure, and thus we can reduce the cost for computation and storage from exponential to linear. Previously known iterative methods also allow for low-rank approximations of the solution but are unable to guarantee that its entries sum up to one as required for a probability distribution. We derive a convergent iterative method using low-rank formats satisfying this condition. We also perform numerical experiments illustrating that the marginal distribution is well approximated with low rank.

math.NA