Searcharxiv⌕ Search

arXiv subjects

Jörn Stöhler

Publications and source records attributed to Jörn Stöhler.

2 recordsLinked to original sources

Efficient all-electron Bethe-Salpeter implementation using crystal symmetries

We describe an all-electron implementation of the Bethe-Salpeter equation (BSE) for the calculation of optical absorption spectra in the full-potential linearized augmented-plane-wave (FLAPW) method. So far, FLAPW implementations have resorted to a simple plane-wave basis for the bare and screened Coulomb potentials, thereby forgoing the all-electron description to some extent. In contrast, we expand the interaction potentials in the all-electron mixed basis. As in most implementations, the BSE is solved by the diagonalization of a two-particle Hamiltonian matrix, whose dimension is proportional to the number of $\mathbf{k}$ points. Due to the large number of $\mathbf{k}$ points required to converge the BSE, the resulting matrix becomes large even for small unit cells. We describe a method that exploits the crystal symmetries to accelerate the construction and diagonalization of the two-particle Hamiltonian. In particular, we employ group theoretical tools to bring the Hamiltonian into block-diagonal form. Furthermore, it is shown that often only one of the blocks needs to be taken into account for the optical absorption spectrum leading to a considerable speedup of the diagonalization step. The code allows for the inclusion of spin-orbit coupling and is parallelized with the possibility of storing the Hamiltonian in distributed memory over many nodes, keeping the memory demands low. To validate our implementation, we show optical absorption spectra and report exciton binding energies for bulk Si, LiF, and MoS$_2$. By exploiting the crystal symmetries, we can reduce the dimension of the Hamiltonian matrix of Si by a factor of five, resulting in a 125-fold speedup in its diagonalization. The calculated exciton binding energies of 22~meV and 76~meV for Si and MoS$_2$ are closer to experimental values than in previous BSE studies.

cond-mat.mtrl-sci↗

The Local Interaction Basis: Identifying Computationally-Relevant and Sparsely Interacting Features in Neural Networks

Mechanistic interpretability aims to understand the behavior of neural networks by reverse-engineering their internal computations. However, current methods struggle to find clear interpretations of neural network activations because a decomposition of activations into computational features is missing. Individual neurons or model components do not cleanly correspond to distinct features or functions. We present a novel interpretability method that aims to overcome this limitation by transforming the activations of the network into a new basis - the Local Interaction Basis (LIB). LIB aims to identify computational features by removing irrelevant activations and interactions. Our method drops irrelevant activation directions and aligns the basis with the singular vectors of the Jacobian matrix between adjacent layers. It also scales features based on their importance for downstream computation, producing an interaction graph that shows all computationally-relevant features and interactions in a model. We evaluate the effectiveness of LIB on modular addition and CIFAR-10 models, finding that it identifies more computationally-relevant features that interact more sparsely, compared to principal component analysis. However, LIB does not yield substantial improvements in interpretability or interaction sparsity when applied to language models. We conclude that LIB is a promising theory-driven approach for analyzing neural networks, but in its current form is not applicable to large language models.

cs.LG↗