SearcharxivSearch

arXiv subjects

Ryan Herbst

Publications and source records attributed to Ryan Herbst.

At least 19 recordsLinked to original sources

HeteroViT: A Versatile Single-Layer Vision Transformer Concept, Co-Designed for Distributed Real-Time Data Reduction on Scientific Detectors

Next-generation X-ray detectors generate data faster than any system can affordably store or process. LCLS-II, the upgraded Linac Coherent Light Source at SLAC, produces data on the order of terabytes per second, with raw-data transfer and storage projected to be prohibitively costly, even though much of the data is not scientifically useful. This concept paper focuses on two major points. The first is versatility: a deliberately tiny, single-layer Vision Transformer (ViT) is enough to serve distinct scientific quick-evaluation tasks. We demonstrate this on two very different problems: (a) a supervised hit/miss/maybe classification on the CSPAD dataset, made to resemble ePixUHR-like detector frames, and (b) a self-supervised latent space for rare-event detection in X-ray diffraction spanning two learning paradigms, two output types, and two detector modalities, with one small backbone. The second is hardware co-design: because the ViT's blocks are structurally uniform, the model maps cleanly onto the heterogeneous hardware already present in the LCLS detector pipeline (ASIC -> FPGA -> GPU) under a simple rule one ASIC is one token so the data is reduced progressively at each stage and a keep/discard decision is produced in real time at the edge. The two claims reinforce each other: versatility is precisely what justifies freezing the front-end in silicon, since a reusable front-end is only worth committing to hardware if it serves many tasks. We are explicit that this is a concept supported by early software analysis, not a hardware demonstration. The natural and primary next phase is the hardware implementation of this distributed pipeline. The decisive evidence still owed an end-to-end latency budget, ASIC feasibility of the in-sensor embedding, and the false-negative behavior that matters for a data veto defines that program. HeteroViT is our first step toward it.

physics.ins-det

Discrete Wavelet Transform for Serial X-ray Crystallography Image Segmentation

Upcoming LCLS-II/II-HE operation at repetition rates approaching 1MHz demands on-detector data reduction to manage the resulting data volumes. We present a 2D discrete wavelet transform (DWT) pre-processing algorithm that segments background scatter from crystal diffraction in serial crystallography images, enabling early data analysis and, when combined with peak finding, lossy compression by transmitting only the identified diffraction peaks. The method zeroes the approximation (LL) coefficients of a multi-level Haar wavelet decomposition and reconstructs from detail subbands only, exploiting the natural separation of smooth background and sharp Bragg peaks in the wavelet domain. Evaluated on 100 simulated nanoBragg frames with known ground truth, the pipeline achieves $F1 \approx 0.96$ at four decomposition levels ($J = 4$), substantially outperforming the established peakfinder8 algorithm ($F1 \approx 0.37$) in both precision ($P \approx 1.00$ vs.\ $0.94$) and recall ($R \approx 0.92$ vs.\ $0.24$). A comparison of 12 wavelet families confirms that Haar is optimal for Bragg-peak detection due to its minimal filter support. Downstream crystallographic analysis performed on real ePix10kA data shows that CC* and $R_\mathrm{split}$ converge at $J = 4$ and track the unprocessed baseline through the practical resolution limit. Under added noise exceeding $\sim$50 ADU, the current pipeline's precision degrades significantly more than that of the pf8 algorithm, exposing a limitation of the proposed strategy. We also demonstrate an FPGA implementation of the DWT filters on an Alveo U200 at 200MHz, with a projected resource footprint compatible with integration into the upcoming ePixUHR firmware and a path to on-detector ASIC implementation in SparkPix detector family.

physics.ins-det

FPGA-Accelerated Real-Time Diagnostics at DIII-D Using the SLAC Neural Network Library for ML Inference

In this work, we demonstrate the deployment of a hardware-accelerated machine learning (ML) inference system integrated into a real-time processing at the DIII-D tokamak fusion reactor. The team has successfully deployed an AMD/Xilinx KCU1500 field-programmable gate array (FPGA) into the realtime Plasma Control System (PCS) nodes that receives the live Beam Emission Spectroscopy (BES) signal used for Edge Localized Mode (ELM) forecasting. The FPGA hosts a dense neural network using the SLAC Neural Network Library (SNL) that has been trained to infer the likelihood of disruptive ELM conditions. This likelihood then feeds a separate plasma controller that uses Resonant Magnetic Perturbation coils to suppress the predicted disruptive condition. The SNL allows for on-the-fly updates of the neural network weights and biases without requiring full hardware resynthesis for the FPGA. Judicious design of the neural-network architecture can further allow for the hot-swapping of multiple classification tasks to be executed on the single FPGA, significantly enhancing the real-time adaptability of the system for context-aware control strategies that respond in real-time to evolving reactor conditions. These adaptive weights naturally support continuous model refinement and seamless task switching during live experimental operation. This use case is chosen as a high rate signal processing example that can serve as a template for general ML-based reactor diagnostic processing for active reactor control systems. We see this as an essential development for achieving reactor relevant operation in future continuous operation fusion devices.

physics.plasm-ph

Machine Learning on Heterogeneous, Edge, and Quantum Hardware for Particle Physics (ML-HEQUPP)

The next generation of particle physics experiments will face a new era of challenges in data acquisition, due to unprecedented data rates and volumes along with extreme environments and operational constraints. Harnessing this data for scientific discovery demands real-time inference and decision-making, intelligent data reduction, and efficient processing architectures beyond current capabilities. Crucial to the success of this experimental paradigm are several emerging technologies, such as artificial intelligence and machine learning (AI/ML), silicon microelectronics, and the advent of quantum algorithms and processing. Their intersection includes areas of research such as low-power and low-latency devices for edge computing, heterogeneous accelerator systems, reconfigurable hardware, novel codesign and synthesis strategies, readout for cryogenic or high-radiation environments, and analog computing. This white paper presents a community-driven vision to identify and prioritize research and development opportunities in hardware-based ML systems and corresponding physics applications, contributing towards a successful transition to the new data frontier of fundamental science.

physics.ins-det

FPGA-Accelerated Real-Time Beam Emission Spectroscopy Diagnostics at DIII-D Using the SLAC Neural Network Library for ML Inference

Achieving reliable real-time control of tokamak plasmas is essential for sustaining high-performance operation in next-generation fusion reactors. A major challenge is the accurate and timely prediction of edge-localized modes (ELMs), especially in high-confinement regimes such as wide-pedestal quiescent H-mode. We present a hardware-accelerated machine learning (ML) inference system integrated into the RTSTAB processing node of the DIII-D real-time diagnostic and control infrastructure. The system uses an AMD/Xilinx KCU1500 FPGA to enable ultra low latency plasma state classification and ELM forecasting. Input features come from real-time Beam Emission Spectroscopy (BES), and the ML model is implemented as a dense neural network using the SLAC Neural Network Library (SNL). A key capability is SNL dynamic parameter loading, which allows on-the-fly updates of neural network weights and biases without hardware resynthesis. This enables multiple classification tasks on a single FPGA design and supports adaptive control strategies that respond to evolving plasma conditions. By decoupling inference from fixed-weight configurations, the system supports continuous model refinement and seamless task switching during live operation. The SNL-based inference engine is fully integrated with the FPGA in the DIII-D RTSTAB Plasma Control System (PCS), improving ELM avoidance, confinement, and operational stability. These results show the feasibility of embedding dynamically reconfigurable FPGA-based ML inference into real-time fusion diagnostic pipelines, providing a scalable and resilient path toward intelligent and autonomous plasma control in future magnetic confinement fusion devices.

physics.plasm-ph

High Precision RF Pulse Shaping with Direct RF Sampling for Future Linear Accelerators

In various of particle accelerator designs, amplitude and phase modulation methods are commonly applied to shape the RF pulses for implementing pulse compressors or compensating for the fluctuations introduced by the high-power RF components and beam loading effects. Phase modulations are typically implemented with additional phase shifters that require drive or control electronics. With our recent next-generation LLRF (NG-LLRF) platform developed based on direct RF sampling technology of RF system-on-chip (RFSoC) devices, RF pulse shaping can be realized without the analogue phase shifters, which can significantly simplify the system architecture. We performed a range of high-power experiments in the C-band to evaluate the RF pulse-shaping capabilities of the NG-LLRF system at different stages of the RF circuits. In this paper, the high-power characterization results with the Cool Copper Collider (C3) structure driven by RF pulses with different modulation schemes will be described. With the pulse modulation and demodulation completely implemented in the digital domain, the RF pulse shaping schemes can be rapidly adapted for X-band structures simply by adding analogue mixers.

physics.acc-ph

Next Generation LLRF Control and Monitoring System for S-Band Linear Accelerators

The low-level RF (LLRF) systems for S-band linear accelerating structures are typically implemented with heterodyne base architectures. We have developed and characterized the next generation LLRF (NG-LLRF) based on the RF system-on-chip (RFSoC) for C-band accelerating structures, and the platform delivered the pulse-to-pulse fluctuation levels considerably better than the requirement of the targeted applications. The NG-LLRF system uses the direct RF sampling technique of the RFSoC, which significantly simplified the architecture compared to the conventional LLRF. We have extended the frequency range of the NG-LLRF to S-band and experimented with different RFSoC devices and system designs to meet the more stringent requirements for S-band LLRF applications. In this paper, the characterization results of the platform with different system architectures will be summarized and the high-power test results of the NG-LLRF with the S-band accelerating structure in the Next Linear Collider Test Accelerator (NLCTA) test facility at the SLAC National Accelerator Laboratory will be presented and analyzed.

physics.acc-ph

ESPPU INPUT: C$^3$ within the "Linear Collider Vision"

The Linear Collider Vision calls for a Linear Collider Facility with a physics reach from a Higgs Factory to the TeV-scale with $e^+e^{-}$ collisions. One of the technologies under consideration for the accelerator is a cold-copper distributed-coupling linac capable of achieving high gradient. This technology is being pursued by the C$^3$ collaboration to understand its applicability to future colliders and broader scientific applications. In this input we share the baseline parameters for a C$^3$ Higgs-factory and the energy reach of up to 3 TeV in the 33 km tunnel foreseen under the Linear Collider Vision. Recent results, near-term plans and future R\&D needs are highlighted.

physics.acc-ph

FPGA-Accelerated SpeckleNN with SNL for Real-time X-ray Single-Particle Imaging

We implement a specialized version of our SpeckleNN model for real-time speckle pattern classification in X-ray Single-Particle Imaging (SPI) using the SLAC Neural Network Library (SNL) on an FPGA. This hardware is optimized for inference near detectors in high-throughput X-ray free-electron laser (XFEL) facilities like the Linac Coherent Light Source (LCLS). To fit FPGA constraints, we optimized SpeckleNN, reducing parameters from 5.6M to 64.6K (98.8% reduction) with 90% accuracy. We also compressed the latent space from 128 to 50 dimensions. Deployed on a KCU1500 FPGA, the model used 71% of DSPs, 75% of LUTs, and 48% of FFs, with an average power consumption of 9.4W. The FPGA achieved 45.015us inference latency at 200 MHz. On an NVIDIA A100 GPU, the same inference consumed ~73W and had a 400us latency. Our FPGA version achieved an 8.9x speedup and 7.8x power reduction over the GPU. Key advancements include model specialization and dynamic weight loading through SNL, eliminating time-consuming FPGA re-synthesis for fast, continuous deployment of (re)trained models. These innovations enable real-time adaptive classification and efficient speckle pattern vetoing, making SpeckleNN ideal for XFEL facilities. This implementation accelerates SPI experiments and enhances adaptability to evolving conditions.

physics.ins-det

Development of a novel bunch oscillation recorder with RFSoC technology

The SuperKEKB accelerator is designed to achieve unprecedented luminosity levels, but this goal is currently hindered by Sudden Beam Loss (SBL) events. These events not only obstruct luminosity improvement but also pose a significant risk to accelerator components, the Belle II detectors, and the superconducting focusing system, potentially leading to severe damage and quenching of the superconducting system. To address this critical challenge, we have developed a novel Bunch Oscillation Recorder (BOR) based on RFSoC technology. The BOR has demonstrated high precision with a position resolution of 0.03 mm, making it a powerful tool for real-time beam monitoring. In its initial deployment, the BOR successfully recorded multiple SBL events, providing valuable data for further analysis. By strategically positioning BORs at the suspected points of SBL origin, we aim to directly identify sources of beam instability. We anticipate that this portable, high-speed BOR monitor will play a crucial role in resolving the SBL issue, ultimately helping achieve SuperKEKB's luminosity targets.

physics.acc-ph

Analysis of Hardware Synthesis Strategies for Machine Learning in Collider Trigger and Data Acquisition

To fully exploit the physics potential of current and future high energy particle colliders, machine learning (ML) can be implemented in detector electronics for intelligent data processing and acquisition. The implementation of ML in real-time at colliders requires very low latencies that are unachievable with a software-based approach, requiring optimization and synthesis of ML algorithms for deployment on hardware. An analysis of neural network inference efficiency is presented, focusing on the application of collider trigger algorithms in field programmable gate arrays (FPGAs). Trade-offs are evaluated between two frameworks, the SLAC Neural Network Library (SNL) and hls4ml, in terms of resources and latency for different model sizes. Results highlight the strengths and limitations of each approach, offering valuable insights for optimizing real-time neural network deployments at colliders. This work aims to guide researchers and engineers in selecting the most suitable hardware and software configurations for real-time, resource-constrained environments.

physics.ins-det

Embedded FPGA Developments in 130nm and 28nm CMOS for Machine Learning in Particle Detector Readout

Embedded field programmable gate array (eFPGA) technology allows the implementation of reconfigurable logic within the design of an application-specific integrated circuit (ASIC). This approach offers the low power and efficiency of an ASIC along with the ease of FPGA configuration, particularly beneficial for the use case of machine learning in the data pipeline of next-generation collider experiments. An open-source framework called "FABulous" was used to design eFPGAs using 130 nm and 28 nm CMOS technology nodes, which were subsequently fabricated and verified through testing. The capability of an eFPGA to act as a front-end readout chip was assessed using simulation of high energy particles passing through a silicon pixel sensor. A machine learning-based classifier, designed for reduction of sensor data at the source, was synthesized and configured onto the eFPGA. A successful proof-of-concept was demonstrated through reproduction of the expected algorithm result on the eFPGA with perfect accuracy. Further development of the eFPGA technology and its application to collider detector readout is discussed.

cs.AR

Next Generation LLRF Control Platform for Compact C-band Linear Accelerator

The Low-Level RF (LLRF) control circuits of linear accelerators (LINACs) are conventionally realized with heterodyne based architectures, which have analog RF mixers for up and down conversion with discrete data converters. We have developed a new LLRF platform for C-band linear accelerator based on the Frequency System-on-Chip (RFSoC) device from AMD Xilinx. The integrated data converters in the RFSoC can directly sample the RF signals in C-band and perform the up and down mixing digitally. The programmable logic and processors required for signal processing for the LLRF control system are also included in a single RFSoC chip. With all the essential components integrated in a device, the RFSoC-based LLRF control platform can be implemented more cost-effectively and compactly, which can be applied to a broad range of accelerator applications. In this paper, the structure and configuration of the newly developed LLRF platform will be described. The LLRF prototype has been tested with high power test setup with a Cool Cooper Collider (C\(^3\)) accelerating structure. The LLRF and the solid state amplifier (SSA) loopback setup demonstrated phase jitter in 1 s as low as 115 fs, which is lower than the requirement of C\(^3\). The rf signals from the klystron forward and accelerating structure captured with peak power up to 16.45 MW will be presented and discussed.

physics.acc-ph

Development of RFSoC-based direct sampling highly multiplexed microwave SQUID readout for future CMB and submillimeter surveys

The SLAC Microresonator Radio Frequency (SMuRF) electronics is being deployed as the readout for the Cosmic Microwave Background (CMB) telescopes of the Simons Observatory (SO). A Radio Frequency System-on-Chip (RFSoC) based readout of microwave frequency resonator based cryogenic sensors is under development at SLAC as an upgrade path for SMuRF with simplified RF hardware, a more compact footprint, and lower total power consumption. The high-speed integrated data converters and digital data path in RFSoC enable direct RF sampling without analog up and down conversion for RF frequencies up to 6 GHz. A comprehensive optimization and characterization study has been performed for direct RF sampling for microwave SQUID multiplexers, which covers noise level, RF dynamic range, and linearity using a prototype implementation. The SMuRF firmware, including the implementation of closed-loop tone tracking, has been ported to the RFSoC platform and interfaced with the quadrature mixers for digital up and down conversion in the data converter data path to realize a full microwave SQUID multiplexer readout. In this paper, a selection of the performance characterization results of direct RF sampling for microwave SQUID multiplexer readout will be summarized and compared with science-driven requirements. Preliminary results demonstrating the read out of cryogenic sensors using the prototype system will also be presented here. We anticipate our new RFSoC-based SMuRF system will be an enabling readout for on-going and future experiments in astronomy and cosmology, which rely on large arrays of cryogenic sensors to achieve their science goals.

astro-ph.IM

End-to-End Modeling of the TDM Readout System for CMB-S4

The CMB-S4 experiment is developing next-generation ground-based microwave telescopes to observe the Cosmic Microwave Background with unprecedented sensitivity. This will require an order of magnitude increase in the 100 mK detector count, which in turn increases the demands on the readout system. The CMB-S4 readout will use time division multiplexing (TDM), taking advantage of faster switches and amplifiers in order to achieve an increased multiplexing factor. To facilitate the design of the new readout system, we have developed a model that predicts the bandwidth and noise performance of this circuity and its interconnections. This is then used to set requirements on individual components in order to meet the performance necessary for the full system. We present an overview of this model and compare the model results to the performance of both legacy and prototype readout hardware.

astro-ph.IM

Evaluating Direct RF Sampling Performance for RFSoC-based Radio-frequency Astronomy Receivers

As the maximum RF input and output frequencies of the integrated data converters in RFSoC increase, it becomes practical to digitize and synthesize RF signals in the majority of C band directly without analogue up and down mixing circuits. The elimination of the mixer circuits can significantly simplify the architecture of the receivers or readouts for radio astronomy telescopes. For the systems with large bandwidth or high channel counts, direct sampling can dramatically reduce the size and cost of overall system. This paper with focus on summarising part of the preliminary characterization results for direct sampling with RFSoC data converters in higher order Nyquist zones.

astro-ph.IM

Higher Order Nyquist Zone Sampling with RFSoC Data Converters for Astronomical and High Energy Physics Readout Systems

From generation to generation, the maximum RF frequency and sampling rate of the integrated data converters in RF system-on-chip (RFSoC) family devices from Xilinx increases significantly. With the integrated digital mixers and up and down conversion blocks in the datapaths of the data converters, those RFSoC devices offer the capability for implementing a full readout system of ground and space-based telescopes and detectors across the electromagnetic spectrum within the devices with minimum or no analog mixing circuit. In this paper, we present the characterization results for the the data converters sampling at higher orders of Nyquist zones to extend the frequency range covered for our targeted readout systems of microwave-frequency resonator-based cryogenic detector and multiplexer systems and other astronomical and high-energy physics instrumentation applications, such as, axion search and dark matter detection. The initial evaluation of the data converters operating higher order Nyquist zones covers two-tones and comb of tones tests to address the concerns in the RF inter-modulation distortion, which is the key performance index for our targeted applications. The characterization of the data converters is performed in the bandwidth of 4-6 GHz and results meet our requirements. The settings and operating strategies of the data converters for our targeted applications will be summarised.

astro-ph.IM

Implementation of a framework for deploying AI inference engines in FPGAs

The LCLS2 Free Electron Laser FEL will generate xray pulses to beamline experiments at up to 1Mhz These experimentals will require new ultrahigh rate UHR detectors that can operate at rates above 100 kHz and generate data throughputs upwards of 1 TBs a data velocity which requires prohibitively large investments in storage infrastructure Machine Learning has demonstrated the potential to digest large datasets to extract relevant insights however current implementations show latencies that are too high for realtime data reduction objectives SLAC has endeavored on the creation of a software framework which translates MLs structures for deployment on Field Programmable Gate Arrays FPGAs deployed at the Edge of the data chain close to the instrumentation This framework leverages Xilinxs HLS framework presenting an API modeled after the open source Keras interface to the TensorFlow library This SLAC Neural Network Library SNL framework is designed with a streaming data approach optimizing the data flow between layers while minimizing the buffer data buffering requirements The goal is to ensure the highest possible framerate while keeping the maximum latency constrained to the needs of the experiment Our framework is designed to ensure the RTL implementation of the network layers supporting full redeployment of weights and biases without requiring resynthesis after training The ability to reduce the precision of the implemented networks through quantization is necessary to optimize the use of both DSP and memory resources in the FPGA We currently have a preliminary version of the toolset and are experimenting with both general purpose example networks and networks being designed for specific LCLS2 experiments.

physics.ins-det