SearcharxivSearch

arXiv subjects

Raymond G. Beausoleil

Publications and source records attributed to Raymond G. Beausoleil.

At least 19 recordsLinked to original sources

How to Build a Quantum Supercomputer: Scaling from Hundreds to Millions of Qubits

In the span of four decades, quantum computation has evolved from an intellectual curiosity to a potentially realizable technology. Today, small-scale demonstrations have become possible for quantum algorithmic primitives on hundreds of physical qubits. Nevertheless, there are significant outstanding challenges in quantum hardware, fabrication, software architecture, and algorithms on the path towards a full-stack scalable quantum computing technology. Here, we provide a comprehensive review of these scaling challenges. We show how to facilitate scaling by adopting existing semiconductor technology to build much higher-quality qubits, employing systems engineering approaches, and performing distributed heterogeneous quantum-classical computing. We provide a detailed resource and sensitivity analysis for quantum applications on surface-code error-corrected quantum computers given current, target, and desired hardware specifications based on superconducting qubits, accounting for a realistic distribution of errors. We provide comprehensive resource estimates for several utility-scale applications including quantum chemistry calculations, catalyst design, NMR spectroscopy, and Fermi-Hubbard simulation. We show that orders of magnitude enhancement in performance could be obtained by a combination of hardware improvements and tight quantum-HPC integration. Furthermore, we introduce high-performance architectures for quantum-probabilistic computing with custom-designed accelerators to tackle today's industry-scale classical optimization, machine learning, and quantum simulation tasks in a cost-effective manner.

quant-ph

Demonstration and Design of Uni-Directional and Ultra-Low Threshold Hybrid Quantum Dot III-V/Si Micro-Ring Laser

Micro-ring lasers (MRLs) are attractive light sources for energy-efficient optical interconnects, but their intrinsic directional bistability leads to unpredictable clockwise/counter-clockwise emission. We demonstrate stable unidirectional emission in hybrid quantum-dot (QD) III-V/Si MRLs using passive reflective feedback integrated on the bus waveguide, leaving the ring cavity unperturbed. Three reflector architectures - Y-splitter loop mirrors, adiabatic Y-splitter loop mirrors, and distributed Bragg reflectors (DBRs) - are benchmarked against a reflector-free bidirectional baseline through combined experiment and coupled-mode-theory rate-equation modeling. All designs preserve ultra-low thresholds of 0.79-1.12 mA (112-158 A/cm^2, roughly an order of magnitude below prior quantum-well unidirectional ring lasers) while enhancing single-facet output power and wall-plug efficiency, with directional isolation up to 27.65 dB for the DBR. The reflectors impose no penalty on the 4-5 GHz modulation bandwidth or its thermal robustness, establishing passive external feedback as a practical route to unidirectional QD MRLs for DWDM-scale optical interconnects.

physics.optics

Distributed Quantum Computing via Adaptive Circuit Knitting

Distributing quantum workloads over many Quantum Processing Units (QPUs) is a crucial step in scaling up quantum computers toward practical quantum advantage due to the limitations in size of a single QPU. In the absence of high-fidelity quantum interconnects, circuit knitting could provide a path to computing certain properties of large quantum systems on many QPUs of limited size in a distributed fashion using only classical communication. Circuit knitting partitions large quantum circuits into manageable sub-circuits, however, reconstructing observables in a straightforward manner comes at an exponential cost in sampling and classical post-processing. To mitigate the overhead this technique incurs, we introduce an Adaptive Circuit Knitting (ACK) method that finds efficient partitions of quantum circuits by discovering regions of minimal entanglement between subsystems. We simulate 1D and 2D disordered mixed-field Ising models up to 60 qubits and show that the ACK approach can reduce circuit knitting sampling overheads by up to four orders of magnitude for observables of interest. We highlight our parallel GPU-accelerated implementation and discuss the need for efficient classical simulators to enable distributed quantum algorithm development. Our techniques could enable efficient distribution of quantum simulation for both near-term and fault-tolerant architectures.

quant-ph

Scalable Back-Propagation-Free Training of Optical Physics-Informed Neural Networks

Physics intelligence and digital twins often require rapid and repeated performance evaluation of various engineering systems (e.g. robots, autonomous vehicles, semiconductor chips) to enable (almost) real-time actions or decision making. This has motivated the development of accelerated partial differential equation (PDE) solvers, in resource-constrained scenarios if the PDE solvers are to be deployed on the edge. Physics-informed neural networks (PINNs) have shown promise in solving high-dimensional PDEs, but the training time on state-of-the-art digital hardware (e.g., GPUs) is still orders-of-magnitude longer than the latency required for enabling real-time decision making. Photonic computing offers a potential solution to address this huge latency gap because of its ultra-high operation speed. However, the lack of photonic memory and the large device sizes prevent training real-size PINNs on photonic chips. This paper proposes a completely back-propagation-free (BP-free) and highly salable framework for training real-size PINNs on silicon photonic platforms. Our approach involves three key innovations: (1) a sparse-grid Stein derivative estimator to avoid the BP in the loss evaluation of a PINN, (2) a dimension-reduced zeroth-order optimization via tensor-train decomposition to achieve better scalability and convergence in BP-free training, and (3) a scalable on-chip photonic PINN training accelerator design using photonic tensor cores. We validate our numerical methods on both low- and high-dimensional PDE benchmarks. Through pre-silicon simulation based on real device parameters, we further demonstrate the significant performance benefit (e.g., real-time training, huge chip area reduction) of our photonic accelerator.

cs.LG

A Full Stack Framework for High Performance Quantum-Classical Computing

To address the growing needs for scalable High Performance Computing (HPC) and Quantum Computing (QC) integration, we present our HPC-QC full stack framework and its hybrid workload development capability with modular hardware/device-agnostic software integration approach. The latest development in extensible interfaces for quantum programming, dispatching, and compilation within existing mature HPC programming environment are demonstrated. Our HPC-QC full stack enables high-level, portable invocation of quantum kernels from commercial quantum SDKs within HPC meta-program in compiled languages (C/C++ and Fortran) as well as Python through a quantum programming interface library extension. An adaptive circuit knitting hypervisor is being developed to partition large quantum circuits into sub-circuits that fit on smaller noisy quantum devices and classical simulators. At the lower-level, we leverage Cray LLVM-based compilation framework to transform and consume LLVM IR and Quantum IR (QIR) from commercial quantum software frontends in a retargetable fashion to different hardware architectures. Several hybrid HPC-QC multi-node multi-CPU and GPU workloads (including solving linear system of equations, quantum optimization, and simulating quantum phase transitions) have been demonstrated on HPE EX supercomputers to illustrate functionality and execution viability for all three components developed so far. This work provides the framework for a unified quantum-classical programming environment built upon classical HPC software stack (compilers, libraries, parallel runtime and process scheduling).

cs.DC

SEPhIA: <1 laser/neuron Spiking Electro-Photonic Integrated Multi-Tiled Architecture for Scalable Optical Neuromorphic Computing

Research into optical spiking neural networks (SNNs) has primarily focused on spiking devices, networks of excitable lasers or numerical modelling of large architectures, often overlooking key constraints such as limited optical power, crosstalk and footprint. We introduce SEPhIA, a photonic-electronic, multi-tiled SNN architecture emphasizing implementation feasibility and realistic scaling. SEPhIA leverages microring resonator modulators (MRMs) and multi-wavelength sources to achieve effective sub-one-laser-per-spiking neuron efficiency. We validate SEPhIA at both device and architecture levels by time-domain co-simulating excitable CMOS-MRR coupled circuits and by devising a physics-aware, trainable optoelectronic SNN model, with both approaches utilizing experimentally derived device parameters. The multi-layer optoelectronic SNN achieves classification accuracies over 90% on a four-class spike-encoded dataset, closely comparable to software models. A design space study further quantifies how photonic device parameters impact SNN performance under constrained signal-to-noise conditions. SEPhIA offers a scalable, expressive, physically grounded solution for neuromorphic photonic computing, capable of addressing spike-encoded tasks.

cs.ET

Heterogeneously Integrated Memristive Laser on Silicon with Non-Volatile Wavelength Tuning

The von-Neumann bottleneck has constrained computing systems from efficiently operating on the increasingly large demand in data from networks and devices. Silicon (Si) photonics offers a powerful solution for this issue by providing a platform for high-bandwidth, energy-efficient interconnects. Furthermore, memristors have emerged as a fundamental building block for non-volatile data storage and novel computing architectures with powerful in-memory processing capabilities. In this paper, we integrate an Al2O3 memristor into a heterogeneous Si quantum dot microring laser to demonstrate the first laser with non-volatile optical memory. The memristor alters the effective optical modal index of the microring laser cavity by the plasma dispersion effect in the high resistance state (HRS) or Joule heating in the low resistance state (LRS), subsequently controlling the output wavelength of the laser in a non-volatile manner. This device enables a novel pathway for future optoelectronic neuromorphic computers and optical memory chips.

physics.optics

Silicon Optical Memory: Non-Volatile Optoelectronic Devices via Si-SiO$_2$ Hysteresis Effect

Implementing on-chip non-volatile optical memories has long been an actively pursued goal, promising significant enhancements in the capability and energy efficiency of photonic integrated circuits. Here, a novel optical memory has been demonstrated exclusively using the semiconductor primary material, silicon. By manipulating the optoelectronic effect of this device, we introduce a hysteresis effect at the silicon-silicon oxide interface, which in turn demonstrates multi-level, non-volatile optical data storage with robust retention and endurance. This new silicon optical memory provides a distinctively simple and accessible route to realize optical data storage in standard silicon foundry processes.

physics.app-ph

Real-Time FJ/MAC PDE Solvers via Tensorized, Back-Propagation-Free Optical PINN Training

Solving partial differential equations (PDEs) numerically often requires huge computing time, energy cost, and hardware resources in practical applications. This has limited their applications in many scenarios (e.g., autonomous systems, supersonic flows) that have a limited energy budget and require near real-time response. Leveraging optical computing, this paper develops an on-chip training framework for physics-informed neural networks (PINNs), aiming to solve high-dimensional PDEs with fJ/MAC photonic power consumption and ultra-low latency. Despite the ultra-high speed of optical neural networks, training a PINN on an optical chip is hard due to (1) the large size of photonic devices, and (2) the lack of scalable optical memory devices to store the intermediate results of back-propagation (BP). To enable realistic optical PINN training, this paper presents a scalable method to avoid the BP process. We also employ a tensor-compressed approach to improve the convergence and scalability of our optical PINN training. This training framework is designed with tensorized optical neural networks (TONN) for scalable inference acceleration and MZI phase-domain tuning for \textit{in-situ} optimization. Our simulation results of a 20-dim HJB PDE show that our photonic accelerator can reduce the number of MZIs by a factor of $1.17\times 10^3$, with only $1.36$ J and $1.15$ s to solve this equation. This is the first real-size optical PINN training framework that can be applied to solve high-dimensional PDEs.

cs.LG

Automatic differentiation accelerated shape optimization approaches to photonic inverse design on rectilinear simulation grids

Shape optimization approaches to inverse design offer low-dimensional, physically-guided parameterizations of structures by representing them as combinations of shape primitives. However, on discretized rectilinear simulation grids, computing the gradient of a user objective via the adjoint variables method requires a sum reduction of the forward/adjoint field solutions and the Jacobian of the simulation material distribution with respect to the structural shape parameters. These shape parameters often perturb large or global parts of the simulation grid resulting in many non-zero Jacobian entries, which are typically computed by finite-difference in practice. Consequently, the gradient calculation can be non-trivial. In this work we propose to accelerate the gradient calculation by invoking automatic differentiation (AutoDiff) in instantiations of structural material distributions. In doing so, we develop extensible differentiable mappings from shape parameters to shape primitives and differentiable effective logic operations (denoted AutoDiffGeo). These AutoDiffGeo definitions may introduce some additional discretization error into the field solutions because they relax notions of sub-pixel smoothing along shape boundaries. However, we show that some mappings (e.g. simple cuboids) can achieve zero error with respect to volumetric averaging strategies. We demonstrate AutoDiff enhanced shape optimization using three integrated photonic examples: a multi-etch blazed grating coupler, a non-adiabatic waveguide transition taper, and a polarization-splitting grating coupler. We find accelerations of the gradient calculation by AutoDiff relative to finite-difference often exceed 50x, resulting in total wall time accelerations of 4x or more on the same hardware with little or no compromise to final device performance. Our code is available open source at https://github.com/smhooten/emopt

cs.CE

Corona: System Implications of Emerging Nanophotonic Technology

We expect that many-core microprocessors will push performance per chip from the 10 gigaflop to the 10 teraflop range in the coming decade. To support this increased performance, memory and inter-core bandwidths will also have to scale by orders of magnitude. Pin limitations, the energy cost of electrical signaling, and the non-scalability of chip-length global wires are significant bandwidth impediments. Recent developments in silicon nanophotonic technology have the potential to meet these off- and on- stack bandwidth requirements at acceptable power levels. Corona is a 3D many-core architecture that uses nanophotonic communication for both inter-core communication and off-stack communication to memory or I/O devices. Its peak floating-point performance is 10 teraflops. Dense wavelength division multiplexed optically connected memory modules provide 10 terabyte per second memory bandwidth. A photonic crossbar fully interconnects its 256 low-power multithreaded cores at 20 terabyte per second bandwidth. We have simulated a 1024 thread Corona system running synthetic benchmarks and scaled versions of the SPLASH-2 benchmark suite. We believe that in comparison with an electrically-connected many-core alternative that uses the same on-stack interconnect power, Corona can provide 2 to 6 times more performance on many memory-intensive workloads, while simultaneously reducing power.

cs.AR

Energy-Efficient Photonic Memory Based on Electrically Programmable Embedded III-V/Si Memristors: Switches and Filters

We demonstrate non-volatile optical functionality by embedding multi-layer $HfO_2/Al_2O_3$ memristors with III-V/Si photonics. The wafer-bonded III-V/Si memristor facilitates non-volatile optical functionality for a variety of devices such as Mach-Zehnder Interferometers (MZIs), and (de-)interleaver filters. The MZI optical memristor exhibits non-volatile optical phase shifts $> π(Δn_{g} > 2.70 \times 10^{-3}$) with ~ 30 dB extinction ratio while consuming 0 electrical power consumption in a true "set-and-forget" operation. We demonstrate 6 non-volatile states with each state capable of 4 Gbps modulation. III-V/Si (de-)interleavers were also demonstrated to exhibit memristive non-volatile passband transformation with full set/reset states. Time duration tests were performed on all devices and indicated non-volatility up to 24 hours and most likely beyond. To the best of our knowledge, we have demonstrated for the first time, non-volatile III-V/Si optical memristors with the largest electric-field driven phase shifts and reconfigurable filters with the lowest power consumption.

physics.optics

RETROSPECTIVE: Corona: System Implications of Emerging Nanophotonic Technology

The 2008 Corona effort was inspired by a pressing need for more of everything, as demanded by the salient problems of the day. Dennard scaling was no longer in effect. A lot of computer architecture research was in the doldrums. Papers often showed incremental subsystem performance improvements, but at incommensurate cost and complexity. The many-core era was moving rapidly, and the approach with many simpler cores was at odds with the better and more complex subsystem publications of the day. Core counts were doubling every 18 months, while per-pin bandwidth was expected to double, at best, over the next decade. Memory bandwidth and capacity had to increase to keep pace with ever more powerful multi-core processors. With increasing core counts per die, inter-core communication bandwidth and latency became more important. At the same time, the area and power of electrical networks-on-chip were increasingly problematic: To be reliably received, any signal that traverses a wire spanning a full reticle-sized die would need significant equalization, re-timing, and multiple clock cycles. This additional time, area, and power was the crux of the concern, and things looked to get worse in the future. Silicon nanophotonics was of particular interest and seemed to be improving rapidly. This led us to consider taking advantage of 3D packaging, where one die in the 3D stack would be a photonic network layer. Our focus was on a system that could be built about a decade out. Thus, we tried to predict how the technologies and the system performance requirements would converge in about 2018. Corona was the result this exercise; now, 15 years later, it's interesting to look back at the effort.

cs.AR

Non-volatile heterogeneous III-V/Si photonics via optical charge-trap memory

We demonstrate, for the first time, non-volatile charge-trap flash memory (CTM) co-located with heterogeneous III-V/Si photonics. The wafer-bonded III-V/Si CTM cell facilitates non-volatile optical functionality for a variety of devices such as Mach-Zehnder Interferometers (MZIs), asymmetric MZI lattice filters, and ring resonator filters. The MZI CTM exhibits full write/erase operation (100 cycles with 500 states) with wavelength shifts of $Δλ_{non-volatile} = 1.16 nm$ ($Δn_{eff,non-volatile} ~ 2.5 \times 10^{-4}$) and a dynamic power consumption $<$ 20 pW (limited by measurement). Multi-bit write operation (2 bits) is also demonstrated and verified over a time duration of 24 hours and most likely beyond. The cascaded 2nd order ring resonator CTM filter exhibited an improved ER of ~ 7.11 dB compared to the MZI and wavelength shifts of $Δλ_{non-volatile} = 0.041 nm$ ($Δn_{eff, non-volatile} = 1.5 \times 10^{-4}$) with similar pW-level dynamic power consumption as the MZI CTM. The ability to co-locate photonic computing elements and non-volatile memory provides an attractive path towards eliminating the von-Neumann bottleneck.

physics.optics

High-Speed and Energy-Efficient Non-Volatile Silicon Photonic Memory Based on Heterogeneously Integrated Memresonator

Recently, interest in programmable photonics integrated circuits has grown as a potential hardware framework for deep neural networks, quantum computing, and field programmable arrays (FPGAs). However, these circuits are constrained by the limited tuning speed and large power consumption of the phase shifters used. In this paper, introduced for the first time are memresonators, or memristors heterogeneously integrated with silicon photonic microring resonators, as phase shifters with non-volatile memory. These devices are capable of retention times of 12 hours, switching voltages lower than 5 V, an endurance of 1,000 switching cycles. Also, these memresonators have been switched using voltage pulses as short as 300 ps with a record low switching energy of 0.15 pJ. Furthermore, these memresonators are fabricated on a heterogeneous III-V/Si platform capable of integrating a rich family of active, passive, and non-linear optoelectronic devices, such as lasers and detectors, directly on-chip to enable in-memory photonic computing and further advance the scalability of integrated photonic processor circuits.

physics.optics

Fast and energy-efficient non-volatile III-V-on-silicon photonic phase shifter based on memristors

Silicon photonics has evolved from lab research to commercial products in the past decade as it plays an increasingly crucial role in data communication for next-generation data centers and high performance computing1. Recently, programmable silicon photonics has also found new applications in quantum2 and classical 3 information processing. A key component of programmable silicon photonic integrated circuits (PICs) is the phase shifter, traditionally realized via the thermo-optic or plasma dispersion effect which are weak, volatile, and power hungry. A non-volatile phase shifter can circumvent these limitations by requiring zero power to maintain the switched phases. Previously non-volatile phase modulation was achieved via phase-change4 or ferroelectric materials5, but the switching energy remains high (pico to nano joules) and the speed is slow (micro to milli seconds). Here, we report a non-volatile III-V-on-silicon photonic phase shifter based on HfO2 memristor with sub-pJ switching energy (~400fJ), representing over an order of magnitude improvement in energy efficiency compared to the state of the art. The non-volatile phase shifter can be switched reversibly using a single 100ns pulse and exhibits an excellent endurance over 800 cycles. This technology can enable future energy-efficient programmable PICs for data centers, optical neural networks, and quantum information processing.

physics.optics

Tensorized Optical Multimodal Fusion Network

We propose the first tensorized optical multimodal fusion network architecture with a self-attention mechanism and low-rank tensor fusion. Simulation results show $51.3 \times$ less hardware requirement and $3.7\times 10^{13}$ MAC/J energy efficiency.

eess.SP