SearcharxivSearch

arXiv subjects

Yue Hao

Publications and source records attributed to Yue Hao.

At least 19 recordsLinked to original sources

Gaussian process learning with flow map refinement for parameter estimation in dynamical systems

Parameter estimation is a central task in data-driven learning of dynamical systems. It aims to recover the underlying physical parameters from observed time-series data, thereby providing interpretable insights into the physical mechanisms governing the system. Gradient/derivative matching methods based on Gaussian process provide an efficient way to perform parameter estimation. Those methods avoid repeated numerical integration and enforce local derivative consistency. However, such local matching may result in global inconsistency with the governing flow map, particularly under scarce and noisy observations. To address this limitation, we propose a framework based on Gaussian process learning with flow map refinement (GPL-FMR), a two-stage parameter estimation framework. The first stage is based on Gaussian process learning algorithm and the posterior obtained from which is transferred as an informative prior to the second stage based on flow-map refinement. The second stage further improves the parameter estimation via optimisation based on global dynamical constraints. We demonstrate and analyse its performance on multiple numerical examples, including the Van der Pol oscillator, the Lotka-Volterra model, and the Lorenz-63 system. The results show that the proposed framework consistently improves parameter estimation accuracy, particularly under scarce and noisy observations.

cs.LG

zenDot: An LLM-integrated quantum TCAD platform for semiconductor quantum-device design and optimization automation

Semiconductor quantum-device design still lacks an integrated Technology Computer-Aided Design (TCAD)-like environment that connects material geometry, quantum many-body simulation, and automated design. Here we introduce zenDot, a large-language model (LLM)-integrated quantum TCAD platform that links a material-labelled device state to a unified condensed-matter physics toolbox. The device and calculation components are integrated into a desktop workbench, Python API, and an embedded LLM agent, allowing electrostatics, charge and transport characterization, correlated-state calculations, and qubit modelling to be executed within one reproducible environment. We demonstrate zenDot on a Si/SiO2 double quantum dot, where a single device state reproduces the characterization workflow and supports hybrid, tunnel-charge, and singlet-triplet qubit analyses. A platform-level universal-control scan revises the singlet-triplet operating point and reduces the predicted worst-gate infidelity by nearly 30-fold. Beyond analysis, the LLM agent directly operates the same physics environment as human users, proposing design changes, executing registered simulations, and iterating on solver-returned metrics under physics-aware validation. Across three demonstration tasks it completes 18 validated design iterations, including geometry modification followed by a full re-solve from the material stack. zenDot establishes a machine-operable quantum TCAD workflow that connects device physics with LLM-driven design exploration.

quant-ph

A Unified Electrostatic-to-Spin Framework for Asymmetric Multi-Gate CMOS Quantum Devices

In advanced complementary metal-oxide-semiconductor (CMOS) quantum chips, compact gate stacks make it difficult to connect lithographic geometry, electrostatic confinement and many-electron spin filling in one transparent model. This connection is central to design-technology co-optimization (DTCO). Here we develop a reduced-order analytical framework for asymmetric multigate silicon quantum-dot devices. Its electrostatic core, the Poisson-kernel coupled-interface Green-function (PK-GF) model, agrees with an independent finite-volume solution at the millivolt scale for the matched two-dimensional problem, without fitting to that solution. We then pass the gate-derived confinement, rather than a harmonic or fitted potential, to a spin-valley many-body calculation for a jellybean quantum dot with N = 2-17 electrons at B = 5 T. The unrestricted Hartree-Fock (UHF) solution supports occupation-dependent, Wigner-molecule-like charge localization but likely overestimates spin polarization. Complete active-space configuration interaction (CASCI) supports a low-spin branch within the tested active spaces, which aligns with the experiments. The workflow therefore connects CMOS layout, device electrostatics, and potential-determined quantum observables, providing an auditable modelling layer for CMOS-based qubit design and DTCO.

cond-mat.mes-hall

Drift-free characterization of electro-optic tuning efficiency in lithium niobate photonic nanocavities

Lithium niobate photonic crystal nanobeam cavity (PCNBC) represents a premier platform for integrated electro-optics, offering deep sub-wavelength mode confinement, enhanced light-matter interactions, and ultralow power consumption. However, accurate characterization of the electro-optic (EO) tuning efficiency in such high-Q devices is fundamentally impeded by DC drift, a time-dependent spectral instability arising from charge redistribution, surface screening, or buffer layer relaxation under sustained electric fields. Here, we report the systematic analysis of DC drift dynamics in lithium niobate nanocavities and demonstrate that conventional quasi-static DC voltage scanning yields highly unreliable characterization data. To circumvent this limitation, we introduce a drift-free, dynamic measurement methodology that employs high-frequency triangular-wave voltage sweeps to effectively decouple the instantaneous electronic Pockels response from slow charge-relaxation processes. Validated across 35 devices with varying electrode geometries, our method delivers reproducible tuning efficiency of 4.3-4.5 pm/V with a low coefficient of variation of 1.1%, showing excellent quantitative agreement with three-dimensional finite-element simulations. This robust, drift-free measurement technique establishes a rigorous standard for the characterization and optimization of resonant cavity electro-optics, accelerating the development of high-performance thin-film lithium niobate photonic integrated circuits.

physics.optics

Approximate Invariant Analysis: An Efficient Framework for Nonlinear Beam Dynamics, Part I: Geometric Approaches of the Poincar\'e Rotation Number

We present the first part of an efficient framework for nonlinear beam dynamics, termed Approximate Invariant Analysis (AIA). The framework is based on the construction of approximate invariants~[Y.~Li, D.~Xu, and Y.~Hao, Phys.\ Rev.\ Accel.\ Beams \textbf{28}, 074001 (2025)] and on the extraction of the betatron frequency with the geometric foundations of Poincar\'e rotation number~[S.~Nagaitsev and T.~Zolkin, Phys.\ Rev.\ Accel.\ Beams \textbf{23}, 054001 (2020)]. The method is demonstrated using the National Synchrotron Light Source~II (NSLS-II) storage ring as an illustrative example.

physics.acc-ph

Programmable Packet Scheduling with Dynamic Reordering at Line Rate

High-speed switch packet scheduling demands both line-rate performance and programmability. Existing programmable hardware scheduling models, such as PIFO and PIEO, can express a broad range of scheduling algorithms; however, their semantics are restricted to packet-level ordering and cannot dynamically reorder buffered packets, which limits the support for dynamic-ordering algorithms such as pFabric. To overcome this limitation, we propose UIFO (Update-In-First-Out), a new programmable scheduling model that introduces a two-level abstraction over classes and packets. UIFO enables dynamic updates to the scheduling order at the class level while preserving in-order packet scheduling within each class, thereby supporting dynamic reordering of already-buffered packets. Furthermore, UIFO remains fully compatible with and generalizes existing PIFO and PIEO models. We implement a hardware prototype of UIFO based on priority-queue designs and evaluate it on an FPGA platform and in a 28 nm ASIC process. Overall, UIFO significantly enhances scheduling expressiveness and maintains favorable scalability while sustaining 100 Gbps line-rate throughput.

cs.NI

2D Ferroelectric Ruddlesden-Popper Perovskites: an Emerging Fully Electronically Controllable Shift Current and Persistent Spin Helix

Two-dimensional (2D) hybrid organic--inorganic perovskites (HOIPs) are promising candidates for next-generation optoelectronic and spintronic applications. This work systematically investigates the relationship between structural distortions and functional responses in three $C_{2v}$-symmetric Ruddlesden--Popper (RP) ferroelectric perovskites, $(4,4\text{-DFPD})_{2}\mathrm{PbI}_{4}$, $(\mathrm{DFCHA})_{2}\mathrm{PbI}_{4}$, and PEPI, using first-principles calculations combined with irreducible representation decomposition and wave-vector point-group symmetry (WPGS) analysis. The results reveal that the lead--iodide framework yields shift-current (SC) magnitudes comparable to, and in specific cases even an order of magnitude larger than, those of traditional ferroelectric oxides, with PEPI reaching a maximum of $69.16\ \mu\mathrm{A}/\mathrm{V}^{2}$. The SC magnitude correlates positively with the octahedral distortion index ($D_i$), while a competition mechanism is identified between covalent bond strength and structural asymmetry, where increased average bond lengths can offset the enhancement induced by $D_i$. Regarding spintronics, $C_{2v}$ symmetry-protected persistent spin textures (PST) are identified. A transition to $C_2$-protected quasi-PST occurs in monoclinic $(4,4\text{-DFHHA})_{2}\mathrm{PbI}_{4}$, leading to a persistent spin helix (PSH) with long-distance spin transport. The synergy among ferroelectricity, SC, and PST enables nonvolatile electrical control of both photocurrent direction and spin configurations. This work provides evaluation criteria and practical guidance for designing high-performance integrated spintronic--photovoltaic devices.

cond-mat.mtrl-sci

Multi-Agent Collaboration for Automated Design Exploration on High Performance Computing Systems

Today's scientific challenges, from climate modeling to Inertial Confinement Fusion design to novel material design, require exploring huge design spaces. In order to enable high-impact scientific discovery, we need to scale up our ability to test hypotheses, generate results, and learn from them rapidly. We present MADA (Multi-Agent Design Assistant), a Large Language Model (LLM) powered multi-agent framework that coordinates specialized agents for complex design workflows. A Job Management Agent (JMA) launches and manages ensemble simulations on HPC systems, a Geometry Agent (GA) generates meshes, and an Inverse Design Agent (IDA) proposes new designs informed by simulation outcomes. While general purpose, we focus development and validation on Richtmyer--Meshkov Instability (RMI) suppression, a critical challenge in Inertial Confinement Fusion. We evaluate on two complementary settings: running a hydrodynamics simulations on HPC systems, and using a pre-trained machine learning surrogate for rapid design exploration. Our results demonstrate that the MADA system successfully executes iterative design refinement, automatically improving designs toward optimal RMI suppression with minimal manual intervention. Our framework reduces cumbersome manual workflow setup, and enables automated design exploration at scale. More broadly, it demonstrates a reusable pattern for coupling reasoning, simulation, specialized tools, and coordinated workflows to accelerate scientific discovery.

cs.AI

Hardware Implementation of Photonic Spiking Hash Retrieval

Hashing retrieval is a pivotal technology for large-scale similarity search, widely applied in retrieval-augmented generation (RAG) for large language models (LLMs), massive image repositories, and bioinformatics sequence matching. However, traditional electronic hashing implementations face severe bottlenecks in power consumption and latency when processing high-dimensional data, while existing photonic neural networks often lack robust mechanisms for direct binary code generation under analog noise. To address these challenges, we propose a hardware-software co-designed photonic spiking hashing framework. We utilize the nonlinear thresholding dynamics of a distributed feedback laser with saturable absorber (DFB-SA) to realize the final binarization of a single-step spiking neural network (SNN). Crucially, a hardware-aware quantization margin loss is introduced to maximize the decision margin, effectively mitigating bit flips caused by optical intensity fluctuations. Validated on MNIST (image) and 20 Newsgroups (text) datasets, our system demonstrates robust binary code generation and high retrieval accuracy comparable to digital baselines. Most significantly, the proposed photonic architecture exhibits superior efficiency with an encoding latency of 2.294 ns/query and an energy consumption of 73.70 pJ/query. This work offers a robust and viable path for ultra-fast, energy-efficient optoelectronic neuromorphic computing in high-throughput information retrieval tasks.

physics.optics

Ion Implantation Enhanced Nucleation Facilitates Heat Transport across Atomically-Sharp Semiconductor Interfaces

Overheating is a critical bottleneck limiting the performance and reliability of next-generation high-power and high-frequency electronics. Interfacial thermal resistance constitutes a significant portion of the total thermal resistance. In this study, we report an ultrahigh thermal boundary conductance (TBC) of approximately 800 MW/m2-K at the atomically-sharp AlN-SiC interface, achieved through an ion implantation-enhanced nucleation epitaxy technique. This value is among the highest TBC values reported for semiconductor interfaces, confirmed by structural characterizations which show an ultrahigh-quality interface. Atomistic Green Function calculations reveal that elastic phonon transmission dominates the interface, with nearly half of the acoustic modes (0-15 THz) exhibiting near-unity transmission due to the atomically sharp structure. Furthermore, using high-energy-resolution electron energy loss spectroscopy, we probe vibrational properties with nanometer spatial resolution and identify unique interfacial phonon modes connecting the mismatched phonon spectra, confirmed by molecular dynamics simulations. The ultrahigh TBC is attributed to both the high elastic phonon transmission due to the high quality interfaces and the inelastic phonon scattering channel due to interfacial phonon modes. These findings not only advance the fundamental understanding of interfacial thermal transport but also provide a pathway for effective thermal management in emerging electronic devices.

cond-mat.mes-hall

Hardware implementation of photonic neuromorphic autonomous navigation

Reinforcement learning (RL) is a core technology enabling the transition of artificial intelligence (AI) from perception to decision-making, but its deployment on conventional electronic hardware suffers from high latency and energy consumption imposed by the von Neumann architecture. Here, we propose a photonic spiking twin delayed deep deterministic policy gradient (TD3) reinforcement learning architecture for neuromorphic autonomous navigation and experimentally validate it on a distributed feedback laser with a saturable absorber (DFB-SA) array. The hybrid architecture integrates a photonic spiking Actor network with dual continuous-valued Critic networks, where the final nonlinear spiking activation layer of the Actor is deployed on the DFB-SA laser array. In autonomous navigation tasks, the system achieves an average reward of 58.22 plus-minus 17.29 and a success rate of 80% plus-minus 8.3%. Hardware-software co-inference demonstrates an estimated energy consumption of 0.78 nJ/inf and an ultra-low latency of 191.20 ps/inf, with co-inference error rates of 0.051% and 0.059% in task scenarios with and without obstacle interference, respectively. Simulations for error-activated channels show full agreement with the expected responses, validating the dynamic characteristics of the DFB-SA laser. The architecture shows strong potential for integration with large-scale photonic linear computing chips, enabling fully-functional photonic computation and low-power, low-latency neuromorphic autonomous navigation.

physics.optics

Photonic spiking reinforcement learning for intelligent routing

Intelligent routing plays a key role in modern communication infrastructure, including data centers, computing networks, and future 6G networks. Although reinforcement learning (RL) has shown great potential for intelligent routing, its practical deployment remains constrained by high energy consumption and decision latency. Here, we propose a photonic spiking RL architecture that implements a proximal policy optimization (PPO)-based intelligent routing algorithm. The performance of the proposed approach is systematically evaluated on a software-defined network (SDN) with a fat-tree topology. The results demonstrate that, under various baseline traffic rate conditions, the PPO-based routing strategy significantly outperforms the conventional Dijkstra algorithm in several key performance metrics. Furthermore, a hardware-software collaborative framework of the spiking Actor network is realized for three typical baseline traffic rates, utilizing a photonic synapse chip based on a Mach-Zehnder interferometer (MZI) array and a photonic spiking neuron chip based on distributed feedback lasers with a saturable absorber (DFB-SAs). Experimental validation on 640 state-action pairs shows that the inference accuracy of the hardware-software collaborative framework is consistent with that of the pure algorithmic implementation. The impacts of different hidden-layer scales in the spiking Actor network and varying network size of fat-tree topology are further analyzed. The integration of photonic spiking RL with SDN-based routing establishes a novel paradigm for intelligent routing optimization, featuring ultra-low latency and high energy efficiency. This approach exhibits broad application prospects in real-time network optimization scenarios, including large-scale data centers, computing networks, satellite Internet systems, and future 6G networks.

physics.optics

A Grouped Sorting Queue Supporting Dynamic Updates for Timer Management in High-Speed Network Interface Cards

With the hardware offloading of network functions, network interface cards (NICs) undertake massive stateful, high-precision, and high-throughput tasks, where timers serve as a critical enabling component. However, existing timer management schemes suffer from heavy software load, low precision, lack of hardware update support, and overflow. This paper proposes two novel operations for priority queues--update and group sorting--to enable hardware timer management. To the best of our knowledge, this work presents the first hardware priority queue to support an update operation through the composition and propagation of basic operations to modify the priorities of elements within the queue. The group sorting mechanism ensures correct timing behavior post-overflow by establishing a group boundary priority to alter the sorting process and element insertion positions. Implemented with a hybrid architecture of a one-dimension (1D) systolic array and shift registers, our design is validated through packet-level simulations for flow table timeout management. Results demonstrate that a 4K-depth, 16-bit timer queue achieves over 500 MHz (175 Mpps, 12 ns precision) in a 28nm process and over 300 MHz (116 Mpps) on an FPGA. Critically, it reduces LUTs and FFs usage by 31% and 25%, respectively, compared to existing designs.

cs.DS

Square matrix-based six-dimensional convergence map for nonlinear beam dynamics analysis

The square matrix-based convergence map (CM) method has proven effective in characterizing nonlinear dynamics in several 4-D dynamical systems. However, when time-dependent perturbations, such as crabbing kicks in colliders, are present, a comprehensive 6-D analysis becomes essential to accurately capture the coupling between transverse and longitudinal motions. In this work, we extend the CM method to the full 6-D phase space by employing an eigen-decomposition-based formulation of the square matrix combined with iterative procedures. The proposed 6-D CM approach is first validated using a simplified crabbing map. We demonstrate that the 6-D CM preserves computational efficiency by using only one-turn map, while successfully resolving high-order resonance structures that remain unresolved by conventional frequency map analysis (FMA). This method is subsequently applied to the dynamic aperture (DA) study of the future Electron-Ion Collider (EIC). The results obtained from the CM analysis exhibit close agreement with those derived from FMA, demonstrating its potential as a powerful tool for nonlinear beam dynamics analysis and DA evaluation, as well as for broader applications in other nonlinear dynamical systems.

physics.acc-ph

Photonic Spiking Graph Neural Network for Energy-Efficient Structured Data Processing

Photonic computing shows great potential for signal processing and artificial intelligence (AI) acceleration due to its ultra-high speed, low energy consumption, and inherent parallelism. Existing photonic computing research has mainly focused on convolutional neural networks (CNNs) and fully connected neural networks (FCNNs), which are well suited for tasks such as image classification and object detection but face limitations in handling graph-structured data. Graph neural networks (GNNs) are specifically designed to model complex relational structures. In this work, we propose a photonic spiking graph neural network (PSGNN) architecture that integrates the structural modeling capability of GNNs, the temporal dynamics of spiking neurons, and the parallel computing advantages of photonic hardware. Through hardware-software co-optimization, a bias-term simulation method tailored for photonic chips is implemented using feature-dimension expansion, enabling effective network training. Experiments on the KarateClub and PubMed datasets achieve training accuracies of 100 percent (92 +/- 2 percent) and test accuracies of 97 percent (90 +/- 1 percent). A silicon photonics 4 x 4 Mach-Zehnder interferometer (MZI) array is further constructed for hardware validation, achieving a test accuracy of 93 percent. The system demonstrates an inference latency of 97 ps, with an energy efficiency of 280 GOPS/W and a computational density of 52 GOPS/mm^2. These results highlight the potential of PSGNN for structured-data processing applications.

physics.optics

Frequency Extraction from Invariant Flows

In non-degenerate integrable Hamiltonian systems, invariant tori can be parameterized equivalently by action variables or by their fundamental frequencies. We introduce an invariant-flow formulation for extracting fundamental frequencies of integrable Hamiltonian systems. By treating invariants as generators of commuting Hamiltonian flows, the frequencies are obtained from time-of-flight parameters along these flows, providing a direct alternative to action-angle constructions and spectral methods based on long time series. The approach yields an explicit numerical procedure that extends naturally to systems with multiple degrees of freedom. Its effectiveness is demonstrated using the McMillan map, where machine-precision accuracy is achieved.

nlin.SI

Hardware-aware Lightweight Photonic Spiking Neural Network for Pattern Classification

There exists a significant scale gap between photonic neural network integrated chips and neural networks, which hinders the deployment and application of photonic neural network. Here, we propose hardware-aware lightweight spiking neural networks (SNNs) architecture tailored to our photonic neuromorphic chips, and conducts hardware-software collaborative computing for solving patter classification tasks. Here, we employed a simplified Mach-Zehnder interferometer (MZI) mesh for performing linear computation, and 16-channel distributed feedback lasers with saturable absorber (DFB-SA) array for performing nonlinear spike activation. Both photonic neuromorphic chips based on the MZI mesh and DFB-SA array were designed, optimized and fabricated. Furthermore, we propose a lightweight spiking neural network (SNN) with discrete cosine transform to reduce input dimension and match the input/output ports number of the photonic neuromorphic chips. We demonstrated an end-to-end inference of an entire layer of the lightweight photonic SNN. The hardware-software collaborative inference accuracy is 90% and 80.5% for MNIST and Fashion-MNIST datasets, respectively. The energy efficiency is 1.39 TOPS/W for the MZI mesh, and is 987.65 GOPS/W for the DFB-SA array. The lightweight architecture and experimental demonstration address the challenge of scale mismatch between the photonic chip and SNN, paving the way for the hardware deployment of photonic SNNs.

physics.optics

Hardware-Software Collaborative Computing of Photonic Spiking Reinforcement Learning for Robotic Continuous Control

Robotic continuous control tasks impose stringent demands on the energy efficiency and latency of computing architectures due to their high-dimensional state spaces and real-time interaction requirements. Conventional electronic computing platforms face computational bottlenecks, whereas the fusion of photonic computing and spiking reinforcement learning (RL) offers a promising alternative. Here, we propose a novel computing architecture based on photonic spiking RL, which integrates the Twin Delayed Deep Deterministic policy gradient (TD3) algorithm with spiking neural network (SNN). The proposed architecture employs an optical-electronic hybrid computing paradigm wherein a silicon photonic Mach-Zehnder interferometer (MZI) chip executes linear matrix computations, while nonlinear spiking activations are performed in the electronic domain. Experimental validation on the Pendulum-v1 and HalfCheetah-v2 benchmarks demonstrates the system capability for software-hardware co-inference, achieving a control policy reward of 5831 on HalfCheetah-v2, a 23.33% reduction in convergence steps, and an action deviation below 2.2%. Notably, this work represents the first application of a programmable MZI photonic computing chip to robotic continuous control tasks, attaining an energy efficiency of 1.39 TOPS/W and an ultralow computational latency of 120 ps. Such performance underscores the promise of photonic spiking RL for real-time decision-making in autonomous and industrial robotic systems.

cs.RO