SearcharxivSearch

arXiv subjects

Farah Fahim

Publications and source records attributed to Farah Fahim.

At least 19 recordsLinked to original sources

How to Build a Quantum Supercomputer: Scaling from Hundreds to Millions of Qubits

In the span of four decades, quantum computation has evolved from an intellectual curiosity to a potentially realizable technology. Today, small-scale demonstrations have become possible for quantum algorithmic primitives on hundreds of physical qubits. Nevertheless, there are significant outstanding challenges in quantum hardware, fabrication, software architecture, and algorithms on the path towards a full-stack scalable quantum computing technology. Here, we provide a comprehensive review of these scaling challenges. We show how to facilitate scaling by adopting existing semiconductor technology to build much higher-quality qubits, employing systems engineering approaches, and performing distributed heterogeneous quantum-classical computing. We provide a detailed resource and sensitivity analysis for quantum applications on surface-code error-corrected quantum computers given current, target, and desired hardware specifications based on superconducting qubits, accounting for a realistic distribution of errors. We provide comprehensive resource estimates for several utility-scale applications including quantum chemistry calculations, catalyst design, NMR spectroscopy, and Fermi-Hubbard simulation. We show that orders of magnitude enhancement in performance could be obtained by a combination of hardware improvements and tight quantum-HPC integration. Furthermore, we introduce high-performance architectures for quantum-probabilistic computing with custom-designed accelerators to tackle today's industry-scale classical optimization, machine learning, and quantum simulation tasks in a cost-effective manner.

quant-ph

On-chip probabilistic inference for charged-particle tracking at the sensor edge

Modern scientific instruments operate under increasingly extreme constraints on bandwidth, latency, and power. Inference at the sensor edge determines experimental data collection efficiency by deciding which information to save for further analysis. Particle tracking detectors at the Large Hadron Collider exemplify this challenge: pixelated silicon sensors generate rich spatiotemporal ionization patterns, yet most of this information is discarded due to data-rate limitations. Concurrently, advancements in co-design tools provide rapid turn-around for incorporating machine learning into application-specific integrated circuits, motivating designs for particle detectors with new integrated technologies. We demonstrate that neural networks embedded in the front-end electronics can infer charged-particle kinematic parameters from a single silicon layer. We regress hit positions and incident angles with calibrated uncertainties, while satisfying stringent constraints on numerical precision, latency, and silicon area. Our results establish a path toward probabilistic inference directly at the edge, opening new opportunities for intelligent sensing in high-rate scientific instruments.

physics.ins-det

On-Detector Machine Learning for Beam-Induced Background Rejection at a 10 TeV Muon Collider

A 10 TeV Muon Collider is a compelling candidate for a future energy-frontier facility, offering unprecedented opportunities to explore the fundamental laws of particle physics. Muon decays in the collider ring produce intense beam-induced background (BIB) that can overwhelm detector occupancy and exceed readout bandwidth constraints. We investigate the potential of on-detector Machine Learning for BIB rejection in the vertex detector, exploiting pixel cluster shapes to distinguish background from collision products. We study three classes of lightweight neural-network architectures, and evaluate their implementation feasibility using high-level synthesis. Selected architectures achieve 88 to 90% data reduction at 99% signal efficiency, while requiring hardware resources compatible with potential ASIC implementation. These results demonstrate the potential of performing substantial BIB rejection directly in the pixel readout, providing a strategy for meeting the tracker readout requirements at a future Muon Collider.

hep-ex

PSEC6: an 8-Channel 40 GSa/s Waveform Sampling ASIC in TSMC 65nm with 10.24 GHz PLL

Picosecond level timing resolution is a prerequisite capability for improved coincidence matching, time-of-flight measurements, and secondary vertex reconstruction. Here, we present the specification, design, and simulation results for a new Application Specific Integrated Circuit (ASIC), called PSEC6, in the TSMC 65nm process. It features 8 channels, a maximum sampling rate of 40 GSa/s, a buffer length of 204.8 nanoseconds, and a 10.24 GHz Phase Locked Loop (PLL), which is the first of its kind in the 65nm CMOS process. The event readout rate is 32 kHz, with the digitization done by an off-chip Analog-to-Digital Converter (ADC). Simulations predict a 4.0 GHz analog input bandwidth and 20 mW per channel during sampling; the 10.24 GHz PLL has a predicted jitter of 550 fs RMS at 15.7 mW. The paper describes the sampling architecture, chip signal paths, PLL design, and presents simulation results.

physics.ins-det

A cryogenic readout integrated circuit with analog pile-up and in-Pixel ADC for high frame rate Skipper CCD-in-CMOS Sensors

The Skipper CCD-in-CMOS Parallel Read-Out Circuit V2 (SPROCKET2) is an in-Pixel analog front end and ADC designed for the vertical readout of Skipper CCD-in-CMOS image sensors. SPROCKET2 is fabricated in a 65 nm CMOS process and each pixel occupies a \SI{60}{\micro\meter} $\times$ \SI{60}{\micro\meter} footprint. SPROCKET2 is intended to be heterogeneously integrated with a Skipper-in-CMOS sensor ASIC, such that one readout pixel is connected to a multiplexed array of sixteen \SI{15}{\micro\meter} $\times$ \SI{15}{\micro\meter} Skipper-in-CMOS pixels. In order to fully leverage the Skipper CCD-in-CMOS sensor's ability to achieve exceptionally low noise from repeated sampling while minimizing the power consumption of digitization and data movement, SPROCKET2 is designed to ``pile-up'' ten successive samples in the analog front end before digitizing the result with a compact 10-bit SAR ADC at a rate of 66.7 ksps, leading to an effective throughput of 667 ksps. The SPROCKET2 pixel achieves input-referred noise $ \lt 100 μV_{rms}$ with $\lt 1$ LSB ADC non-linearity while consuming an estimated $44 μW$ per pixel. A SPROCKET2 test pixel was submitted in December 2022, and test results are presented.

physics.ins-det

Characterization of a 28 nm $\textit{smartpixels}$ ASIC With On-Chip ML for Particle Tracking Detectors

We present a 28 nm CMOS pixel readout integrated circuit implementing in-pixel analog signal processing and on-chip machine learning data filtering for particle tracking detectors. Our ASIC comprises two $32 \times 8$ pixel matrices with a pixel pitch of $25 \times 25~μ\mathrm{m}^2$, in which each pixel integrates a charge-sensitive amplifier with synchronous auto-zero offset cancellation and a 2-bit flash ADC with programmable thresholds. Two analog front-end architectures, single-ended and differential, are implemented and characterized. Digitized pixel data are combined into row-wise projections and processed by an on-chip, fully combinational neural network classifier for data reduction. Measurements at room temperature using charge injection demonstrate an equivalent noise charge of $54.6~\mathrm{e}^{-}$ and a threshold dispersion of $\sim$78.2~\unit{\electron} at nominal bias, linear response up to several~\unit{\kilo\electron}, and stable operation at a 10~MHz clock frequency. The neural network output is compared with offline RTL predictions and agrees for $99.06\%$ of $1.5 \times 10^{5}$ test inputs.

physics.ins-det

A Sub-electron-noise Skipper-CCD Readout ASIC with Improved Channel-to-channel Isolation and an Integrated Cryogenic Voltage Reference

The MIDNA application specific integrated circuits (ASICs) are a series of skipper-CCD readout chips fabricated in a 65 nm low-power CMOS process that implement a correlated double sampling signal processing chain based on dual-slope integrators. They are capable of working from room to cryogenic temperatures, down to 84 K. The present iteration of the ASIC has been fabricated including several design updates and the addition of an on-chip voltage reference, resulting in improved performance. This work presents the main vulnerabilities solved, the changes carried out, and the resulting performance benefits. Measurements with a skipper-CCD and the ASIC at 140 K showed that the single-electron resolution can be reached by averaging the measured charge in the analog domain using the analog pile-up technique with a readout noise as low as 0.11 erms of equivalent charge for 1200 samples. The channel-to-channel crosstalk was also characterized showing values better than -62 dB.

physics.ins-det

Machine Learning for Arbitrary Single-Qubit Rotations on an Embedded Device

Here we present a technique for using machine learning (ML) for single-qubit gate synthesis on field programmable logic for a superconducting transmon-based quantum computer based on simulated studies. Our approach is multi-stage. We first bootstrap a model based on simulation with access to the full statevector for measuring gate fidelity. We next present an algorithm, named adapted randomized benchmarking (ARB), for fine-tuning the gate on hardware based on measurements of the devices. We also present techniques for deploying the model on programmable devices with care to reduce the required resources. While the techniques here are applied to a transmon-based computer, many of them are portable to other architectures.

quant-ph

hls4ml: A Flexible, Open-Source Platform for Deep Learning Acceleration on Reconfigurable Hardware

We present hls4ml, a free and open-source platform that translates machine learning (ML) models from modern deep learning frameworks into high-level synthesis (HLS) code that can be integrated into full designs for field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs). With its flexible and modular design, hls4ml supports a large number of deep learning frameworks and can target HLS compilers from several vendors, including Vitis HLS, Intel oneAPI and Catapult HLS. Together with a wider eco-system for software-hardware co-design, hls4ml has enabled the acceleration of ML inference in a wide range of commercial and scientific applications where low latency, resource usage, and power consumption are critical. In this paper, we describe the structure and functionality of the hls4ml platform. The overarching design considerations for the generated HLS code are discussed, together with selected performance results.

cs.AR

Sensor Co-design for $\textit{smartpixels}$

Pixel tracking detectors at upcoming collider experiments will see unprecedented charged-particle densities. Real-time data reduction on the detector will enable higher granularity and faster readout, possibly enabling the use of the pixel detector in the first level of the trigger for a hadron collider. This data reduction can be accomplished with a neural network (NN) in the readout chip bonded with the sensor that recognizes and rejects tracks with low transverse momentum (p$_T$) based on the geometrical shape of the charge deposition (``cluster''). To design a viable detector for deployment at an experiment, the dependence of the NN as a function of the sensor geometry, external magnetic field, and irradiation must be understood. In this paper, we present first studies of the efficiency and data reduction for planar pixel sensors exploring these parameters. A smaller sensor pitch in the bending direction improves the p$_T$ discrimination, but a larger pitch can be partially compensated with detector depth. An external magnetic field parallel to the sensor plane induces Lorentz drift of the electron-hole pairs produced by the charged particle, broadening the cluster and improving the network performance. The absence of the external field diminishes the background rejection compared to the baseline by $\mathcal{O}$(10%). Any accumulated radiation damage also changes the cluster shape, reducing the signal efficiency compared to the baseline by $\sim$ 30 - 60%, but nearly all of the performance can be recovered through retraining of the network and updating the weights. Finally, the impact of noise was investigated, and retraining the network on noise-injected datasets was found to maintain performance within 6% of the baseline network trained and evaluated on noiseless data.

physics.ins-det

End-to-end workflow for machine learning-based qubit readout with QICK and hls4ml

We present an end-to-end workflow for superconducting qubit readout that embeds co-designed Neural Networks (NNs) into the Quantum Instrumentation Control Kit (QICK). Capitalizing on the custom firmware and software of the QICK platform, which is built on Xilinx RFSoC FPGAs, we aim to leverage machine learning (ML) to address critical challenges in qubit readout accuracy and scalability. The workflow utilizes the hls4ml package and employs quantization-aware training to translate ML models into hardware-efficient FPGA implementations via user-friendly Python APIs. We experimentally demonstrate the design, optimization, and integration of an ML algorithm for single transmon qubit readout, achieving 96% single-shot fidelity with a latency of 32ns and less than 16% FPGA look-up table resource utilization. Our results offer the community an accessible workflow to advance ML-driven readout and adaptive control in quantum information processing applications.

quant-ph

Skipper-in-CMOS: Non-Destructive Readout with Sub-Electron Noise Performance for Pixel Detectors

The Skipper-in-CMOS image sensor integrates the non-destructive readout capability of Skipper Charge Coupled Devices (Skipper-CCDs) with the high conversion gain of a pinned photodiode in a CMOS imaging process, while taking advantage of in-pixel signal processing. This allows both single photon counting as well as high frame rate readout through highly parallel processing. The first results obtained from a 15 x 15 um^2 pixel cell of a Skipper-in-CMOS sensor fabricated in Tower Semiconductor's commercial 180 nm CMOS Image Sensor process are presented. Measurements confirm the expected reduction of the readout noise with the number of samples down to deep sub-electron noise of 0.15rms e-, demonstrating the charge transfer operation from the pinned photodiode and the single photon counting operation when the sensor is exposed to light. The article also discusses new testing strategies employed for its operation and characterization.

astro-ph.IM

Intelligent Pixel Detectors: Towards a Radiation Hard ASIC with On-Chip Machine Learning in 28 nm CMOS

Detectors at future high energy colliders will face enormous technical challenges. Disentangling the unprecedented numbers of particles expected in each event will require highly granular silicon pixel detectors with billions of readout channels. With event rates as high as 40 MHz, these detectors will generate petabytes of data per second. To enable discovery within strict bandwidth and latency constraints, future trackers must be capable of fast, power efficient, and radiation hard data-reduction at the source. We are developing a radiation hard readout integrated circuit (ROIC) in 28nm CMOS with on-chip machine learning (ML) for future intelligent pixel detectors. We will show track parameter predictions using a neural network within a single layer of silicon and hardware tests on the first tape-outs produced with TSMC. Preliminary results indicate that reading out featurized clusters from particles above a modest momentum threshold could enable using pixel information at 40 MHz.

physics.ins-det

Design of an 8-Channel 40 GS/s 20 mW/Ch Waveform Sampling ASIC in 65 nm CMOS

1 ps timing resolution is the entry point to signature based searches relying on secondary/tertiary vertices and particle identification. We describe a preliminary design for PSEC5, an 8-channel 40 GS/s waveform-sampling ASIC in the TSMC 65 nm process targetting 1 ps resolution at 20 mW power per channel. Each channel consists of four fast and one slow switched capacitor arrays (SCA), allowing ps time resolution combined with a long effective buffer. Each fast SCA is 1.6 ns long and has a nominal sampling rate of 40 GS/s. The slow SCA is 204.8 ns long and samples at 5 GS/s. Recording of the analog data for each channel is triggered by a fast discriminator capable of multiple triggering during the window of the slow SCA. To achieve a large dynamic range, low leakage, and high bandwidth, the SCA sampling switches are implemented as 2.5 V nMOSFETs controlled by 1.2 V shift registers. Stored analog data are digitized by an external ADC at 10 bits or better. Specifications on operational parameters include a 4 GHz analog bandwidth and a dead time of 20 microseconds, corresponding to a 50 kHz readout rate, determined by the choice of the external ADC.

physics.ins-det

Smart Pixels: In-pixel AI for on-sensor data filtering

We present a smart pixel prototype readout integrated circuit (ROIC) designed in CMOS 28 nm bulk process, with in-pixel implementation of an artificial intelligence (AI) / machine learning (ML) based data filtering algorithm designed as proof-of-principle for a Phase III upgrade at the Large Hadron Collider (LHC) pixel detector. The first version of the ROIC consists of two matrices of 256 smart pixels, each 25$\times$25 $μ$m$^2$ in size. Each pixel consists of a charge-sensitive preamplifier with leakage current compensation and three auto-zero comparators for a 2-bit flash-type ADC. The frontend is capable of synchronously digitizing the sensor charge within 25 ns. Measurement results show an equivalent noise charge (ENC) of $\sim$30e$^-$ and a total dispersion of $\sim$100e$^-$ The second version of the ROIC uses a fully connected two-layer neural network (NN) to process information from a cluster of 256 pixels to determine if the pattern corresponds to highly desirable high-momentum particle tracks for selection and readout. The digital NN is embedded in-between analog signal processing regions of the 256 pixels without increasing the pixel size and is implemented as fully combinatorial digital logic to minimize power consumption and eliminate clock distribution, and is active only in the presence of an input signal. The total power consumption of the neural network is $\sim$ 300 $μ$W. The NN performs momentum classification based on the generated cluster patterns and even with a modest momentum threshold, it is capable of 54.4\% - 75.4\% total data rejection, opening the possibility of using the pixel information at 40MHz for the trigger. The total power consumption of analog and digital functions per pixel is $\sim$ 6 $μ$W per pixel, which corresponds to $\sim$ 1 W/cm$^2$ staying within the experimental constraints.

physics.ins-det

Smartpixels: Towards on-sensor inference of charged particle track parameters and uncertainties

The combinatorics of track seeding has long been a computational bottleneck for triggering and offline computing in High Energy Physics (HEP), and remains so for the HL-LHC. Next-generation pixel sensors will be sufficiently fine-grained to determine angular information of the charged particle passing through from pixel-cluster properties. This detector technology immediately improves the situation for offline tracking, but any major improvements in physics reach are unrealized since they are dominated by lowest-level hardware trigger acceptance. We will demonstrate track angle and hit position prediction, including errors, using a mixture density network within a single layer of silicon as well as the progress towards and status of implementing the neural network in hardware on both FPGAs and ASICs.

hep-ex

Fast and Efficient Type-II Phototransistors Integrated on Silicon

Increasing the efficiency and reducing the footprint of on-chip photodetectors enables dense optical interconnects for emerging computational and sensing applications. Avalanche photodetectors (APD) are currently the dominating on-chip photodetectors. However, the physics of avalanche multiplication leads to low energy efficiencies and prevents device operation at a high gain, due to a high excess noise, resulting in the need for electrical amplifiers. These properties significantly increase power consumption and footprint of current optical receivers. In contrast, heterojunction phototransistors (HPT) exhibit high efficiency and very small excess noise at high gain. However, HPT's gain-bandwidth product (GBP) is currently inferior to that of APDs at low optical powers. Here, we demonstrate that the type-II energy band alignment in an antimony-based HPT results in a significantly smaller junction capacitance and higher GBP at low optical powers. We used a CMOS-compatible heterogeneous integration method to create compact optical receivers on silicon with an energy efficiency that is about one order of magnitude higher than that of the best reported integrated APDs on silicon at a similar GBP of 270 GHz. Bitrate measurements show data rate spatial density above 800 Tbps per mm2, and an energy-per-bit consumption of only 6 fJ/bit at 3 Gbps. These unique features suggest new opportunities for creating highly efficient and compact on-chip optical receivers based on devices with type-II band alignment.

physics.ins-det

Achieving Single-Electron Sensitivity at Enhanced Speed in Fully-Depleted CCDs with Double-Gate MOSFETs

We introduce a new output amplifier for fully-depleted thick p-channel CCDs based on double-gate MOSFETs. The charge amplifier is an n-type MOSFET specifically designed and operated to couple the fully-depleted CCD with high charge-transfer efficiency. The junction coupling between the CCD and MOSFET channels has enabled high sensitivity, demonstrating sub-electron readout noise in one pixel charge measurement. We have also demonstrated the non-destructive readout capability of the device. Achieving single-electron and single-photon per pixel counting in the entire CCD pixel array has been made possible through the averaging of a small number of samples. We have demonstrated fully-depleted CCD readout with better performance than the floating diffusion and floating gate amplifiers available today, in both single and multisampling regimes, boasting at least six times the speed of floating gate amplifiers.

physics.ins-det