SearcharxivSearch

arXiv subjects

Ryan Hamerly

Publications and source records attributed to Ryan Hamerly.

At least 19 recordsLinked to original sources

Homodyne Photonic Tensor Processor exceeds 1,000-TOPS

High-performance computing underpins modern artificial intelligence (AI), enabling foundation models, real-time inference and perception in autonomous systems, and data-intensive scientific simulations. Recent advances in quantization techniques utilizing low-precision computation without degrading model accuracy, create new opportunities for analog photonic computing characterized by ultra-high clock rates and low energy consumption. Here we propose and demonstrate a coherent homodyne integrated circuit capable of general matrix multiplication (GEMM) with aggregate throughput that exceeds 1,000 TOPS (tera-operations per second), enabled by massive on-chip optical fanout and parallelism. By leveraging time multiplexing, the required modulator count is reduced from O($N^2$) to O(N), allowing dense integration of record-scale 256 $\times$ 256 homodyne units (each <0.0064 $mm^2$) within a single reticle. We employ wafer-scale fabricated 64 thin-film lithium niobate (TFLN) transmitters (each over 40-GHz bandwidth with propagation loss of 0.2 dB/cm) to encode data and chip-to-chip coupled to Si/SiN computing circuits (64 channels). Our system achieves up to 7-bit computational accuracy across 8 $\times$ 8 parallel channels at record computing clockrate 120 Gbaud/s, and 6-bit statistical accuracy across 256 $\times$ 100 channels at 20-128 Gbaud/s, representing a total throughput of 1,000-6,000 TOPS. Massive parallelism amortizes the optoelectronic (OE) conversion to allow 330-TOPS/W efficiency using foundry-available packaging technology. The system throughput is benchmarked with Qwen2.5-0.5 billion parameter models that generate accurate tokens. High throughput and energy efficiency establish a near-term pathway toward light-based accelerators for large-scale training and low-latency inference from datacenters to edges, accelerating new models toward artificial general intelligence.

cs.ET

Quantization-aware Photonic Homodyne computing for Accelerated Artificial Intelligence and Scientific Simulation

Modern problems in high-performance computing, ranging from training and inferencing deep learning models in computer vision and language models to simulating complex physical systems with nonlinearly-coupled equations, require exponential growth of computational resources. Photonic analog systems are emerging with solutions of intrinsic parallelism, high bandwidth, and low propagation loss. However, their application has been hindered by the low analog accuracy due to the electro-optic distortion, material nonlinearities, and signal-to-noise ratios. Here we overcome this barrier with a quantization-aware digital-photonic mixed-precision framework across chiplets for accelerated AI processing and physical simulation. Using Lithium Niobate photonics with channel equalization techniques, we demonstrate linear multiplication (9-bit amplitude-phase decoupling) in homodyne optical logics with 6-bit precision at the clock rate of 128 giga-symbol-per-second (128 GS/s), enabling AI processing with 6 ns latency. Codesign hardware-algorithms, including iterative solvers, sparse-dense quantization, and bit-sliced matrix multiplication, explore photonic amplitude and phase coherence for complex-valued, physics-inspired computation. In electromagnetic problems, our approach yields 12-bit solutions for partial differential equations (PDEs) in scattering problems that would conventionally require up to 32-bit and often even 64-bit precision. These results preserve digital-level fidelity while leveraging the high-speed low-energy photonic hardware, establishing a pathway toward general-purpose optical acceleration for generative artificial intelligence, real-time robotics, and accurate simulation for climate challenges and biological discoveries.

cs.ET

Demonstration and Non-volatile Trimming of a Highly-Parallel, High-Capacity Silicon Microdisk Transmitter

Optical interconnects are the most promising solution to address the data-movement bottleneck in data centers. Silicon microdisks, benefiting from their compact footprint, low energy consumption, and wavelength division multiplexing (WDM) capability, have emerged as an attractive and scalable platform for optical modulation. However, microdisk resonators inherently exhibit low fabrication error tolerance, limiting their practical deployment. Here, utilizing a CMOS photonics platform, we demonstrate 1.2 Tb/s of off-die bandwidth through a 64 microdisk modulator system. In addition, we develop an automated, close-looped, non-reversible, low-loss, and picometer-precision permanent wavelength tuning technique using laser trimming. The trimming technique reduces 33 % of the energy consumption needed to thermally tune the microdisk resonant wavelength. Using this technique, we achieve a fully passive, 5-channel dense wavelength division multiplexing (DWDM, 50 GHz spacing) transmitter. The integration of the high speed (1.2 Tb/s), low energy consumption (29 fJ/bit) and the permanent wavelength trimming lays a robust foundation for next-generation optical interconnect systems, poised to facilitate scaling of future AI and computing hardware.

physics.optics

Single-Shot Matrix-Matrix Multiplication Optical Tensor Processor for Deep Learning

The ever-increasing data demand craves advancements in high-speed and energy-efficient computing hardware. Analog optical neural network (ONN) processors have emerged as a promising solution, offering benefits in bandwidth and energy consumption. However, existing ONN processors exhibit limited computational parallelism, and while certain architectures achieve high parallelism, they encounter serious scaling roadblocks for large-scale implementation. This restricts the throughput, latency, and energy efficiency advantages of ONN processors. Here, we introduce a spatial-wavelength-temporal hyper-multiplexed ONN processor that supports high data dimensionality, high computing parallelism and is feasible for large-scale implementation, and in a single time step, a three-dimensional matrix-matrix multiplication (MMM) optical tensor processor is demonstrated. Our hardware accelerates convolutional neural networks (CNNs) and deep neural networks (DNNs) through parallel matrix multiplication. We demonstrate benchmark image recognition using a CNN and a subsequently fully connected DNN in the optical domain. The network works with 292,616 weight parameters under ultra-low optical energy of 20 attojoules (aJ) per multiply and accumulate (MAC) at 96.4% classification accuracy. The system supports broad spectral and spatial bandwidths and is capable for large-scale demonstration, paving the way for highly efficient large-scale optical computing for next-generation deep learning.

physics.optics

Towards the Information-Theoretic Limit of Programmable Photonics

The scalability of many programmable photonic circuits is limited by the $2\pi$ tuning range needed for the constituent phase shifters. To address this problem, we introduce the concept of a phase-efficient circuit architecture, where the average phase shift is $\ll 2\pi$. We derive a universal information-theoretic limit to the phase-shift efficiency of universal multiport interferometers, and propose a "3-MZI" architecture that approaches this limit to within a factor of $2\times$, approximately a $10\times$ reduction in average phase shift over the prior art, where the average phase shift scales inversely with system size as $O(1/\sqrt{N})$. For non-unitary circuits, we show that the 3-MZI saturates the theoretical bound for Gaussian-distributed target matrices. Using this architecture, we show optical neural network training with all phase shifters constrained to $\lesssim 0.2$ radians without loss of accuracy.

physics.optics

Quantum-secure multiparty deep learning

Secure multiparty computation enables the joint evaluation of multivariate functions across distributed users while ensuring the privacy of their local inputs. This field has become increasingly urgent due to the exploding demand for computationally intensive deep learning inference. These computations are typically offloaded to cloud computing servers, leading to vulnerabilities that can compromise the security of the clients' data. To solve this problem, we introduce a linear algebra engine that leverages the quantum nature of light for information-theoretically secure multiparty computation using only conventional telecommunication components. We apply this linear algebra engine to deep learning and derive rigorous upper bounds on the information leakage of both the deep neural network weights and the client's data via the Holevo and the Cram\'er-Rao bounds, respectively. Applied to the MNIST classification task, we obtain test accuracies exceeding $96\%$ while leaking less than $0.1$ bits per weight symbol and $0.01$ bits per data symbol. This weight leakage is an order of magnitude below the minimum bit precision required for accurate deep learning using state-of-the-art quantization techniques. Our work lays the foundation for practical quantum-secure computation and unlocks secure cloud deep learning as a field.

quant-ph

Hybrid AM/FM Mode-Locking of Singly-Resonant OPOs

We investigate a new mode-locking regime in the singly-resonant OPO employing simultaneous amplitude- and frequency-modulation of the intracavity field. This OPO exhibits deterministic, "turn-key" formation of a stable, broadband, chirped frequency comb with high conversion efficiency. Comb-forming dynamics follow a simple phase-space dynamical model, governed by cavity dispersion and modulator chirp, which agrees closely with full numerical simulations. The comb exhibits fast, mode-hop-free tuning over the full gain window of the OPA crystal, controlled by the modulator frequency. Conditions for comb stability, and techniques to enhance comb bandwidth through intentional phase-mismatch and chirping, are investigated.

physics.optics

Hypermultiplexed Integrated-Photonics-based Tensor Optical Processor

The escalating data volume and complexity resulting from the rapid expansion of artificial intelligence (AI), internet of things (IoT) and 5G/6G mobile networks is creating an urgent need for energy-efficient, scalable computing hardware. Here we demonstrate a hypermultiplexed integratedphotonics-based tensor optical processor (HITOP) that can perform trillions of operations per second (TOPS) at the energy efficiency of 40 TOPS/W. Space-time-wavelength three-dimensional (3D) optical parallelism enables O($N^{2}$) operations per clock-cycle using O($N$) modulator devices. The system is built with wafer-fabricated III/V micron-scale lasers and high-speed thin-film Lithium-Niobate electro-optics for encoding at 10s femtojoule/symbol. Lasing threshold incorporates analog inline rectifier (ReLu) nonlinearity for low-latency activation. The system scalability is verified with machine learning models of 405,000 parameters. A combination of high clockrates, energy-efficient processing and programmability unlocks the potential of light for large-scale AI accelerators in applications ranging from training of large AI models to real-time decision making in edge deployment.

cs.ET

Ultrafast second-order nonlinear photonics -- from classical physics to non-Gaussian quantum dynamics

Photonic integrated circuits with second-order ($\chi^{(2)}$) nonlinearities are rapidly scaling to remarkably low powers. At this time, state-of-the-art devices achieve saturated nonlinear interactions with thousands of photons when driven by continuous-wave lasers, and further reductions in these energy requirements enabled by the use of ultrafast pulses may soon push nonlinear optics into the realm of single-photon nonlinearities. This tutorial reviews these recent developments in ultrafast nonlinear photonics, discusses design strategies for realizing few-photon nonlinear interactions, and presents a unified treatment of ultrafast quantum nonlinear optics using a framework that smoothly interpolates from classical behaviors to the few-photon scale. These emerging platforms for quantum optics fundamentally differ from typical realizations in cavity quantum electrodynamics due to the large number of coupled optical modes. Classically, multimode behaviors have been well studied in nonlinear optics, with famous examples including soliton formation and supercontinuum generation. In contrast, multimode quantum systems exhibit a far greater variety of behaviors, and yet closed-form solutions are even sparser than their classical counterparts. In developing a framework for ultrafast quantum optics, we will identify what behaviors carry over from classical to quantum devices, what intuition must be abandoned, and what new opportunities exist at the intersection of ultrafast and quantum nonlinear optics. While this article focuses on establishing connections between the classical and quantum behaviors of devices with $\chi^{(2)}$ nonlinearities, the frameworks developed here are general and are readily extended to the description of dynamical processes based on third-order ($\chi^{(3)}$) nonlinearities.

physics.optics

Mesoscopic ultrafast nonlinear optics -- The emergence of multimode quantum non-Gaussian physics

Over the last few decades, nonlinear optics has become significantly more nonlinear, traversing nearly a billionfold improvement in energy efficiency, with ultrafast nonlinear nanophotonics in particular emerging as a frontier for combining both spatial and temporal engineering. At present, cutting-edge experiments in nonlinear nanophotonics place us just above the mesoscopic regime, where a few hundred photons suffice to trigger nonlinear saturation. In contrast to classical or deep-quantum optics, the mesoscale is characterized by dynamical interactions between mean-field, Gaussian, and non-Gaussian quantum features, all within a close hierarchy of scales. When combined with the inherent multimode complexity of optical fields, such hybrid quantum-classical dynamics present theoretical, experimental, and engineering challenges to the contemporary framework of quantum optics. In this review, we highlight the unique physics that emerges in multimode nonlinear optics at the mesoscale and outline key principles for exploiting both classical and quantum features to engineer novel functionalities. We briefly survey the experimental landscape and draw attention to outstanding technical challenges in materials, dispersion engineering, and device design for accessing mesoscopic operation. Finally, we speculate on how these capabilities might usher in some new paradigms in quantum photonics, from quantum-augmented information processing to nonclassical-light-driven dynamics and phenomena to all-optical non-Gaussian measurement and sensing. The physics unlocked at the mesoscale present significant challenges and opportunities in theory and experiment alike, and this review is intended to serve as a guidepost as we begin to navigate this new frontier in ultrafast quantum nonlinear optics.

quant-ph

Asymptotically Fault-Tolerant Programmable Photonics

Component errors limit the scaling of programmable coherent photonic circuits. These errors arise because the standard tunable photonic coupler -- the Mach-Zehnder interferometer (MZI) -- cannot be perfectly programmed to the cross state. Here, we introduce two modified circuit architectures that overcome this limitation: (1) a 3-splitter MZI mesh for generic errors, and (2) a broadband MZI+Crossing design for correlated errors. Because these designs allow for perfect realization of the cross state, the matrix fidelity no longer decreases with mesh size, allowing scaling to arbitrarily large meshes. The proposed architectures support progressive self-configuration, are more compact than previous MZI-doubling schemes, and do not require additional phase shifters. This eliminates a major obstacle to the development of very-large-scale linear photonic circuits.

physics.optics

Stability of Self-Configuring Large Multiport Interferometers

Realistic multiport interferometers (beamsplitter meshes) are sensitive to component imperfections, and this sensitivity increases with size. Self-configuration techniques can be employed to correct these imperfections, but not all techniques are equal. This paper highlights the importance of algorithmic stability in self-configuration. Naive approaches based on sequentially setting matrix elements are unstable and perform poorly for large meshes, while techniques based on power ratios perform well in all cases, even in the presence of large errors. Based on this insight, we propose a self-configuration scheme for triangular meshes that requires only external detectors and works without prior knowledge of the component imperfections. This scheme extends to the rectangular mesh by adding a single array of detectors along the diagonal.

cs.ET

Accurate Self-Configuration of Rectangular Multiport Interferometers

Multiport interferometers based on integrated beamsplitter meshes are widely used in photonic technologies. While the rectangular mesh is favored for its compactness and uniformity, its geometry resists conventional self-configuration approaches, which are essential to programming large meshes in the presence of fabrication error. Here, we present a new configuration algorithm, related to the $2\times 2$ block decomposition of a unitary matrix, that overcomes this limitation. Our proposed algorithm is robust to errors, requires no prior knowledge of the process variations, and relies only on external sources and detectors. We show that self-configuration using this technique reduces the effect of fabrication errors by the same quadratic factor observed in triangular meshes. This relaxes a significant limit to the size of multiport interferometers, removing a major roadblock to the scaling of optical quantum and machine-learning hardware.

physics.optics

Temporal trapping: a route to strong coupling and deterministic optical quantum computation

The realization of deterministic photon-photon gates is a central goal in optical quantum computation and engineering. A longstanding challenge is that optical nonlinearities in scalable, room-temperature material platforms are too weak to achieve the required strong coupling, due to the critical loss-confinement tradeoff in existing photonic structures. In this work, we introduce a novel confinement method, dispersion-engineered temporal trapping, to circumvent the tradeoff, paving a route to all-optical strong coupling. Temporal confinement is imposed by an auxiliary trap pulse via cross-phase modulation, which, combined with the spatial confinement of a waveguide, creates a "flying cavity" that enhances the nonlinear interaction strength by at least an order of magnitude. Numerical simulations confirm that temporal trapping confines the multimode nonlinear dynamics to a single-mode subspace, enabling high-fidelity deterministic quantum gate operations. With realistic dispersion engineering and loss figures, we show that temporally trapped ultrashort pulses could achieve strong coupling on near-term nonlinear nanophotonic platforms. Our results highlight the potential of ultrafast nonlinear optics to become the first scalable, high-bandwidth, and room-temperature platform that achieves a strong coupling, opening a new path to quantum computing, simulation, and light sources.

quant-ph

Transferable Learning on Analog Hardware

While analog neural network (NN) accelerators promise massive energy and time savings, an important challenge is to make them robust to static fabrication error. Present-day training methods for programmable photonic interferometer circuits, a leading analog NN platform, do not produce networks that perform well in the presence of static hardware errors. Moreover, existing hardware error correction techniques either require individual re-training of every analog NN (which is impractical in an edge setting with millions of devices), place stringent demands on component quality, or introduce hardware overhead. We solve all three problems by introducing one-time error-aware training techniques that produce robust NNs that match the performance of ideal hardware and can be exactly transferred to arbitrary highly faulty photonic NNs with hardware errors up to 5x larger than present-day fabrication tolerances.

cs.ET

A Self-Similar Sine-Cosine Fractal Architecture for Multiport Interferometers

Multiport interferometers based on integrated beamsplitter meshes have recently captured interest as a platform for many emerging technologies. In this paper, we present a novel architecture for multiport interferometers based on the Sine-Cosine fractal decomposition of a unitary matrix. Our architecture is unique in that it is self-similar, enabling the construction of modular multi-chiplet devices. Due to this modularity, our design enjoys improved resilience to hardware imperfections as compared to conventional multiport interferometers. Additionally, the structure of our circuit enables systematic truncation, which is key in reducing the hardware footprint of the chip as well as compute time in training optical neural networks, while maintaining full connectivity. Numerical simulations show that truncation of these meshes gives robust performance even under large fabrication errors. This design is a step forward in the construction of large-scale programmable photonics, removing a major hurdle in scaling up to practical machine learning and quantum computing applications.

physics.optics

Quantum nondemolition measurements with optical parametric amplifiers for ultrafast universal quantum information processing

Realization of a room-temperature ultra-fast photon-number-resolving (PNR) quantum nondemolition (QND) measurement would have significant implications for photonic quantum information processing (QIP), enabling, e.g., deterministic quantum computation in discrete-variable architectures, but the requirement for strong coupling has hampered the development of scalable implementations. In this work, we propose and analyze a nonlinear-optical route to PNR QND using quadratic (i.e., $χ^{(2)}$) nonlinear interactions. We show that the coherent pump field driving a phase-mismatched optical parametric amplifier (OPA) experiences displacements conditioned on the number of signal Bogoliubov excitations. A measurement of the pump displacement thus provides a QND measurement of the signal Bogoliubov excitations, projecting the signal mode to a squeezed photon-number state. We then show how our nonlinear OPA dynamics can be utilized for deterministically generating Gottesman-Kitaev-Preskill states only with additional Gaussian resources, offering an all-optical route for fault-tolerant QIP in continuous-variable systems. Finally, we place these QND schemes into a more traditional context by highlighting analogies between the phase-mismatched optical parametric oscillator and multilevel atom-cavity QED systems, by showing how continuous monitoring of the outcoupled pump quadrature induces conditional localization of the intracavity signal mode onto squeezed photon-number states. Our analysis suggests that our proposal may be viable in near-term $χ^{(2)}$ nonlinear nanophotonics, highlighting the rich potential of OPA as a universal tool for ultrafast non-Gaussian quantum state engineering and quantum computation.

quant-ph

Single chip photonic deep neural network with accelerated training

As deep neural networks (DNNs) revolutionize machine learning, energy consumption and throughput are emerging as fundamental limitations of CMOS electronics. This has motivated a search for new hardware architectures optimized for artificial intelligence, such as electronic systolic arrays, memristor crossbar arrays, and optical accelerators. Optical systems can perform linear matrix operations at exceptionally high rate and efficiency, motivating recent demonstrations of low latency linear algebra and optical energy consumption below a photon per multiply-accumulate operation. However, demonstrating systems that co-integrate both linear and nonlinear processing units in a single chip remains a central challenge. Here we introduce such a system in a scalable photonic integrated circuit (PIC), enabled by several key advances: (i) high-bandwidth and low-power programmable nonlinear optical function units (NOFUs); (ii) coherent matrix multiplication units (CMXUs); and (iii) in situ training with optical acceleration. We experimentally demonstrate this fully-integrated coherent optical neural network (FICONN) architecture for a 3-layer DNN comprising 12 NOFUs and three CMXUs operating in the telecom C-band. Using in situ training on a vowel classification task, the FICONN achieves 92.7% accuracy on a test set, which is identical to the accuracy obtained on a digital computer with the same number of weights. This work lends experimental evidence to theoretical proposals for in situ training, unlocking orders of magnitude improvements in the throughput of training data. Moreover, the FICONN opens the path to inference at nanosecond latency and femtojoule per operation energy efficiency.

cs.ET