SearcharxivSearch

arXiv subjects

Tim Fischer

Publications and source records attributed to Tim Fischer.

24 records · Page 2Linked to original sources

TCN-CUTIE: A 1036 TOp/s/W, 2.72 uJ/Inference, 12.2 mW All-Digital Ternary Accelerator in 22 nm FDX Technology

Tiny Machine Learning (TinyML) applications impose uJ/Inference constraints, with a maximum power consumption of tens of mW. It is extremely challenging to meet these requirements at a reasonable accuracy level. This work addresses the challenge with a flexible, fully digital Ternary Neural Network (TNN) accelerator in a RISC-V-based System-on-Chip (SoC). Besides supporting Ternary Convolutional Neural Networks, we introduce extensions to the accelerator design that enable the processing of time-dilated Temporal Convolutional Neural Networks (TCNs). The design achieves 5.5 uJ/Inference, 12.2 mW, 8000 Inferences/sec at 0.5 V for a Dynamic Vision Sensor (DVS) based TCN, and an accuracy of 94.5 % and 2.72 uJ/Inference, 12.2 mW, 3200 Inferences/sec at 0.5 V for a non-trivial 9-layer, 96 channels-per-layer convolutional network with CIFAR-10 accuracy of 86 %. The peak energy efficiency is 1036 TOp/s/W, outperforming the state-of-the-art silicon-proven TinyML quantized accelerators by 1.67x while achieving competitive accuracy.

cs.AR

MiniFloat-NN and ExSdotp: An ISA Extension and a Modular Open Hardware Unit for Low-Precision Training on RISC-V cores

Low-precision formats have recently driven major breakthroughs in neural network (NN) training and inference by reducing the memory footprint of the NN models and improving the energy efficiency of the underlying hardware architectures. Narrow integer data types have been vastly investigated for NN inference and have successfully been pushed to the extreme of ternary and binary representations. In contrast, most training-oriented platforms use at least 16-bit floating-point (FP) formats. Lower-precision data types such as 8-bit FP formats and mixed-precision techniques have only recently been explored in hardware implementations. We present MiniFloat-NN, a RISC-V instruction set architecture extension for low-precision NN training, providing support for two 8-bit and two 16-bit FP formats and expanding operations. The extension includes sum-of-dot-product instructions that accumulate the result in a larger format and three-term additions in two variations: expanding and non-expanding. We implement an ExSdotp unit to efficiently support in hardware both instruction types. The fused nature of the ExSdotp module prevents precision losses generated by the non-associativity of two consecutive FP additions while saving around 30% of the area and critical path compared to a cascade of two expanding fused multiply-add units. We replicate the ExSdotp module in a SIMD wrapper and integrate it into an open-source floating-point unit, which, coupled to an open-source RISC-V core, lays the foundation for future scalable architectures targeting low-precision and mixed-precision NN training. A cluster containing eight extended cores sharing a scratchpad memory, implemented in 12 nm FinFET technology, achieves up to 575 GFLOPS/W when computing FP8-to-FP16 GEMMs at 0.8 V, 1.26 GHz.

cs.AR

Automatic detection of lesion load change in Multiple Sclerosis using convolutional neural networks with segmentation confidence

The detection of new or enlarged white-matter lesions in multiple sclerosis is a vital task in the monitoring of patients undergoing disease-modifying treatment for multiple sclerosis. However, the definition of 'new or enlarged' is not fixed, and it is known that lesion-counting is highly subjective, with high degree of inter- and intra-rater variability. Automated methods for lesion quantification hold the potential to make the detection of new and enlarged lesions consistent and repeatable. However, the majority of lesion segmentation algorithms are not evaluated for their ability to separate progressive from stable patients, despite this being a pressing clinical use-case. In this paper we show that change in volumetric measurements of lesion load alone is not a good method for performing this separation, even for highly performing segmentation methods. Instead, we propose a method for identifying lesion changes of high certainty, and establish on a dataset of longitudinal multiple sclerosis cases that this method is able to separate progressive from stable timepoints with a very high level of discrimination (AUC = 0.99), while changes in lesion volume are much less able to perform this separation (AUC = 0.71). Validation of the method on a second external dataset confirms that the method is able to generalize beyond the setting in which it was trained, achieving an accuracy of 83% in separating stable and progressive timepoints. Both lesion volume and count have previously been shown to be strong predictors of disease course across a population. However, we demonstrate that for individual patients, changes in these measures are not an adequate means of establishing no evidence of disease activity. Meanwhile, directly detecting tissue which changes, with high confidence, from non-lesion to lesion is a feasible methodology for identifying radiologically active patients.

cs.CV

Microscopic model for Bose-Einstein condensation and quasiparticle decay

Sufficiently dimerized quantum antiferromagnets display elementary S=1 excitations, triplon quasiparticles, protected by a gap at low energies. At higher energies, the triplons may decay into two or more triplons. A strong enough magnetic field induces Bose-Einstein condensation of triplons. For both phenomena the compound IPA-CuCl3 is an excellent model system. Nevertheless no quantitative model was determined so far despite numerous studies. Recent theoretical progress allows us to analyse data of inelastic neutron scattering (INS) and of magnetic susceptibility to determine the four magnetic couplings J1=-2.3meV, J2=1.2meV, J3=2.9meV and J4=-0.3meV. These couplings determine IPA-CuCl3 as system of coupled asymmetric S=1/2 Heisenberg ladders quantitatively. The magnetic field dependence of the lowest modes in the condensed phase as well as the temperature dependence of the gap without magnetic field corroborate this microscopic model.

cond-mat.str-el

Truncation errors in self-similar continuous unitary transformations

Effects of truncation in self-similar continuous unitary transformations (S-CUT) are estimated rigorously. We find a formal description via an inhomogeneous flow equation. In this way, we are able to quantify truncation errors within the framework of the S-CUT and obtain rigorous error bounds for the ground state energy and the highest excited level. These bounds can be lowered exploiting symmetries of the Hamiltonian. We illustrate our approach with results for a toy model of two interacting hard-core bosons and the dimerized S=1/2 Heisenberg chain.

cond-mat.str-el

Adapted continuous unitary transformation to treat systems with quasiparticles of finite lifetime

An improved generator for continuous unitary transformations is introduced to describe systems with unstable quasiparticles. Its general properties are derived and discussed. To illustrate this approach we investigate the asymmetric antiferromagnetic spin-1/2 Heisenberg ladder which allows for spontaneous triplon decay. We present results for the low energy spectrum and the momentum resolved spectral density of this system. In particular, we show the resonance behavior of the decaying triplon explicitly.

cond-mat.str-el