SearcharxivSearch

arXiv · 2305.19455

Implementation of a framework for deploying AI inference engines in FPGAs

Abstract

The LCLS2 Free Electron Laser FEL will generate xray pulses to beamline experiments at up to 1Mhz These experimentals will require new ultrahigh rate UHR detectors that can operate at rates above 100 kHz and generate data throughputs upwards of 1 TBs a data velocity which requires prohibitively large investments in storage infrastructure Machine Learning has demonstrated the potential to digest large datasets to extract relevant insights however current implementations show latencies that are too high for realtime data reduction objectives SLAC has endeavored on the creation of a software framework which translates MLs structures for deployment on Field Programmable Gate Arrays FPGAs deployed at the Edge of the data chain close to the instrumentation This framework leverages Xilinxs HLS framework presenting an API modeled after the open source Keras interface to the TensorFlow library This SLAC Neural Network Library SNL framework is designed with a streaming data approach optimizing the data flow between layers while minimizing the buffer data buffering requirements The goal is to ensure the highest possible framerate while keeping the maximum latency constrained to the needs of the experiment Our framework is designed to ensure the RTL implementation of the network layers supporting full redeployment of weights and biases without requiring resynthesis after training The ability to reduce the precision of the implemented networks through quantization is necessary to optimize the use of both DSP and memory resources in the FPGA We currently have a preliminary version of the toolset and are experimenting with both general purpose example networks and networks being designed for specific LCLS2 experiments.

Explore related subjects

Keep this discovery

BibTeXRIS

Ryan Herbst, Ryan Coffee, Nathan Fronk, Kukhee Kim, Kuktae Kim, Larry Ruckman, J. J. Russell. 2023-05-30. Implementation of a framework for deploying AI inference engines in FPGAs. https://arxiv.org/abs/2305.19455

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

High-Speed Semi-FE Readout Module for ATLAS MDT at HL-LHC: Design and Production-Level Characterization

The High-Luminosity upgrade of the Large Hadron Collider (HL-LHC) introduces increased demands on the ATLAS Muon Spectrometer, particularly in terms of data throughput, timing distribution and system reliability. The Phase-II Chamber Service Module (CSM) is a key component of the upgraded Monitored Drift Tube (MDT) trigger and readout system, providing a high-speed interface between the front-end electronics and the backend systems. This paper describes the design and implementation of the Phase-II CSM, together with its validation. The results show that the CSM supports two independent optical uplinks, each operating at a line rate of 10.24 Gbps, together with clock distribution and slow control in the expected operating environment. Integration with small-diameter MDT (sMDT) chambers and tests with the prototype L0MDT trigger system are also presented. The CSM boards are now in production and will be used for installation and integration during the upcoming LHC Long Shutdown.

physics.ins-det

Spectral Discrimination of Deposited Gamma-Ray Energies in a Simulated CeBr$_3$ Scintillator

We show that wavelength measurements of individual detected optical photons may provide additional information about gamma-ray energy deposited in a CeBr$_3$ crystal when the detected-photon-count distributions overlap for nearby gamma-ray energies. Monoenergetic 662 and 629 keV gammas are used in a Geant4 simulation of a $25\times25\times20~\mathrm{mm^3}$ CeBr$_3$ crystal. Assuming a light yield of $6.0\times10^4$ photons/MeV, a wavelength-independent photon-detection efficiency of 30%, and a wavelength resolution of $\sigma_{\lambda}=40$ nm, we find that the fraction of photons reconstructed above 385 nm gives an event-level separation of $\sim$ 2 standard deviations between the 662 and 629 keV event populations selected within the same $\sim$ 1%-wide detected-photon-count interval. No timing or reconstructed interaction-position information is used. The result demonstrates, within the present simulation model, that event-dependent optical spectra can retain energy information beyond an undifferentiated photon count.

physics.ins-det

Operation of a negative ion gas time projection chamber without electronegative fill gases

The high fidelity reconstruction of particle tracks in micropatterned gaseous time projection chambers renders this technology ideal for future rare-event searches, including direction-sensitive dark matter experiments. Large drift distances are typically required for such experiments, so that the overall spatial resolution is limited by diffusion. Negative ion drift exhibits lower diffusion than electron drift and is thus an attractive option for realising a large-scale detector. The use of electronegative gases to create negative ions introduces technical challenges, most notably a reduction in gain when compared to conventional gas mixtures. In this study, we demonstrate a new method for negative ion generation via dissociative electron attachment using the conventional molecular fill gas CF$_4$. Our optical measurements of negative ion drift indicate electron attachment lengths of $<$1 mm and comparable gain to electron avalanches. The individual negative ion avalanches were also time-resolved, allowing the number of ions reaching the readout to be counted. We measure an improved energy resolution by single ion counting, relative to an integrated electron avalanche signal measured under identical gain conditions.

physics.ins-det