Searcharxiv⌕ Search

arXiv subjects

Xiaodong Zheng

Publications and source records attributed to Xiaodong Zheng.

13 recordsLinked to original sources

From $O(mn)$ to $O(r^2)$: Two-Sided Low-Rank Communication for Adam in Distributed Training with Memory Efficiency

As foundation models continue to scale, pretraining increasingly relies on data-parallel distributed optimization, making bandwidth-limited gradient synchronization a key bottleneck. Orthogonally, projection-based low-rank optimizers were mainly designed for memory efficiency, but remain suboptimal for communication-limited training: one-sided synchronization still transmits an $O(rn)$ object for an $m\times n$ matrix gradient and refresh steps can dominate peak communicated bytes. We propose TSR, which brings two-sided low-rank communication to Adam-family updates (TSR-Adam) by synchronizing a compact core $U^\top G V\in\mathbb{R}^{r\times r}$, reducing the dominant per-step payload from $O(mn)$ to $O(r^2)$ while keeping moment states in low-dimensional cores. To further reduce the peak communication from subspace refresh, TSR-Adam adopts a randomized SVD-based refresh that avoids full-gradient synchronization. We additionally extend low-rank communication to embedding gradients with embedding-specific ranks and refresh schedules, yielding additional communication and memory savings over keeping embeddings dense. Across pretraining from 60M to 1B model scales, TSR-Adam reduces average communicated bytes per step by $13\times$, and on GLUE fine-tuning it reduces communication by $25\times$, while achieving comparable performance; we further provide a theoretical stationarity analysis for the proposed update. Code is available at https://github.com/DKmiyan/TSR-Adam.

cs.LG↗

Efficient and broadband quantum frequency comb generation in a monolithic AlGaAs-on-insulator microresonator

The exploration of photonic systems for quantum information processing has generated widespread interest in multiple cutting-edge research fields. Photonic frequency encoding stands out as an especially viable approach, given its natural alignment with established optical communication technologies, including fiber networks and wavelength-division multiplexing systems. Substantial reductions in hardware resources and improvements in quantum performance can be expected by utilizing multiple frequency modes. The integration of nonlinear photonics with microresonators provides a compelling way for generating frequency-correlated photon pairs across discrete spectral modes. Here, by leveraging the high material nonlinearity and low nonlinear loss, we demonstrate an efficient chip-scale multi-wavelength quantum light source based on AlGaAs-on-insulator, featuring a free spectral range of approximately 200 GHz at telecom wavelengths. The optimized submicron waveguide geometry provides both high effective nonlinearity (~550 m$^{-1}$W$^{-1}$) and broad generation bandwidth, producing eleven distinct wavelength pairs across a 35.2 nm bandwidth with an average spectral brightness of 2.64 GHz mW$^{-2}$nm$^{-1}$. The generation of energy-time entanglement for each pair of frequency modes is verified through Franson interferometry, yielding an average net visibility of 93.1%. With its exceptional optical gain and lasing capabilities, the AlGaAs-on-insulator platform developed here shows outstanding potential for realizing fully integrated, ready-to-deploy quantum photonic systems on chip.

quant-ph↗

FZOO: Fast Zeroth-Order Optimizer for Fine-Tuning Large Language Models towards Adam-Scale Speed

Fine-tuning large language models (LLMs) often faces GPU memory bottlenecks: the backward pass of first-order optimizers like Adam increases memory usage to more than 10 times the inference level (e.g., 633 GB for OPT-30B). Zeroth-order (ZO) optimizers avoid this cost by estimating gradients only from forward passes, yet existing methods like MeZO usually require many more steps to converge. Can this trade-off between speed and memory in ZO be fundamentally improved? Normalized-SGD demonstrates strong empirical performance with greater memory efficiency than Adam. In light of this, we introduce FZOO, a Fast Zeroth-Order Optimizer toward Adam-Scale Speed. FZOO reduces the total forward passes needed for convergence by employing batched one-sided estimates that adapt step sizes based on the standard deviation of batch losses. It also accelerates per-batch computation through the use of Rademacher random vector perturbations coupled with CUDA's parallel processing. Extensive experiments on diverse models, including RoBERTa-large, OPT (350M-66B), Phi-2, and Llama3, across 11 tasks validate FZOO's effectiveness. On average, FZOO outperforms MeZO by 3 percent in accuracy while requiring 3 times fewer forward passes. For RoBERTa-large, FZOO achieves average improvements of 5.6 percent in accuracy and an 18 times reduction in forward passes compared to MeZO, achieving convergence speeds comparable to Adam. We also provide theoretical analysis proving FZOO's formal equivalence to a normalized-SGD update rule and its convergence guarantees. FZOO integrates smoothly into PEFT techniques, enabling even larger memory savings. Overall, our results make single-GPU, high-speed, full-parameter fine-tuning practical and point toward future work on memory-efficient pre-training.

cs.LG↗

A measurement-device-independent quantum key distribution network using optical frequency comb

Quantum key distribution (QKD), which promises secure key exchange between two remote parties, is now moving toward the realization of scalable and secure QKD networks (QNs). Fully connected, trusted node-free QNs have been realized based on entanglement distribution, in which the low key rate as well as the large overhead makes their practical deployment and application challenging. Here, we propose and experimentally demonstrate a fully connected multi-user QKD network based on a wavelength-multiplexed measurement-device-independent (MDI) QKD protocol. By combining this novel protocol with integrated optical frequency combs, we achieve an average secure key rate of 267 bits per second for about 30 dB of link attenuation per user pair -- more than three orders of magnitude higher than previous entanglement-based works. More importantly, we realize secure key sharing between two different pairs of users simultaneously, which requires four-photon detection and is not possible with the previous two-photon entanglement distribution. Our work paves the way for the realization of large-scale QKD networks with full connectivity and simultaneous communication capability among multiple users.

quant-ph↗

Experimental Measurement-Device-Independent Quantum Cryptographic Conferencing

Quantum cryptographic conferencing (QCC) allows sharing secret keys among multiple distant users and plays a crucial role in quantum networks. Because of the fragility and low generation rate of genuine multipartite entangled states required in QCC, realizing and extending QCC with the entanglement-based protocol is challenging. Measurement-device-independent (MDI) QCC, which removes all detector side channels, is a feasible long-distance quantum communication scheme to practically generate multipartite correlation with multiphoton projection measurement. Here we experimentally realize the three-user MDIQCC protocol with four-intensity decoy-state method, in which we employ the polarization encoding and the Greenberger-Horne-Zeilinger state projection measurement. Our work demonstrates the experimental feasibility of the MDI QCC, which lays the foundation for the future realization of quantum networks with multipartite communication tasks.

quant-ph↗

Successful Transmission Probability and SIR Meta Distribution Analysis for Multi-Antenna Cache-Enabled Networks with Interference Nulling

This paper investigates a multi-antenna cache-enabled network with interference nulling (IN) employed at base stations. Two IN schemes, namely, the fixed IN scheme and the flexible IN scheme are considered to improve the received signal-to-interference ratio (SIR) at users. To thoroughly explore the effects of the caching parameter and the IN parameters on the network performance, we focus on the analysis of not only the successful transmission probability (STP) but the SIR meta distribution. For each IN scheme, the expression for the STP is derived and an approximated expression for the SIR meta distribution is also obtained by deriving the first and second moments of an upper bound of the link reliability and utilizing the beta distribution. With this analytical framework, we compare the performance of these two IN schemes and gain some useful system design guidelines from the perspectives of the STP and the SIR meta distribution by numerical simulations.

cs.IT↗

Cooperating Cracks in Two-Dimensional Crystals

The pattern development of multiple cracks in extremely anisotropic solids such as bilayer or multilayer two-dimensional (2D) crystals contains rich physics, which, however, remains largely unexplored. We studied crack interaction across neighboring 2D layers by transmission electron microscopy and molecular dynamics simulations. Parallel and anti-parallel ('En-Passant') cracks attract and repel each other in bilayer 2D crystals, respectively, in stark contrast to the behaviors of co-planar cracks. We show that the misfit between in-plane displacement fields around the crack tips results in non-uniform interlayer shear, which modifies the crack driving forces by creating an antisymmetric component of the stress intensity factor. The cross-layer interaction between cracks directly leads to material toughening, the strength of which increases with the shear stiffness and decreases with the crack spacings. Backed by the experimental findings and simulation results, a theory that marries the theory of linear elastic fracture mechanics and the shear-lag model is presented, which guides the unconventional approach to engineer fracture patterns and enhance material resistance to cracking.

cond-mat.mtrl-sci↗

Quantum storage of entangled photons at telecom wavelengths in a crystal

The quantum internet -- in synergy with the internet that we use today -- promises an enabling platform for next-generation information processing, including exponentially speed-up distributed computation, secure communication, and high-precision metrology. The key ingredients for realizing such a global network are the distribution and storage of quantum entanglement. As ground-based quantum networks are likely to be based on existing fiber networks, telecom-wavelength entangled photons and corresponding quantum memories are of central interest. Recently, $\rm^{167}Er^{3+}$ ions have been identified as a promising candidate for an efficient, broadband quantum memory at telecom wavelength. However, to date, no storage of entangled photons, the crucial step of quantum memory using these promising ions, $\rm^{167}Er^{3+}$, has been reported. Here, we demonstrate the storage and recall of the entangled state of two telecom photons generated from an integrated photonic chip based on a silicon nitride micro-ring resonator. Combining the natural narrow linewidth of the entangled photons and long storage time of $\rm^{167}Er^{3+}$ ions, we achieve storage time of 1.936 $μ$s, more than 387 times longer than in previous works. Successful storage of entanglement in the crystal is certified by a violation of an entanglement witness with more than 23 standard deviations (-0.234 $\pm$ 0.010) at 1.936 $μ$s storage time. These results pave the way for realizing quantum networks based on solid-state devices.

quant-ph↗

Heterogeneously integrated, superconducting silicon-photonic platform for measurement-device-independent quantum key distribution

Integrated photonics provides a route both to miniaturize quantum key distribution (QKD) devices and to enhance their performance. A key element for achieving discrete-variable QKD is a single-photon detector. It is highly desirable to integrate detectors onto a photonic chip to enable the realization of practical and scalable quantum networks. We realize an integrated heterogeneous superconducting-silicon-photonic chip. Harnessing the unique high-speed feature of our optical waveguide-integrated superconducting detector, we perform the first optimal Bell-state measurement (BSM) of time-bin encoded qubits generated from two independent lasers. The optimal BSM enables an increased key rate of measurement-device-independent QKD, which is immune to all attacks against the detection system, and hence provides the basis for a QKD network with untrusted relays. Together with the time-multiplexed technique, we have enhanced the sifted key rate by almost one order of magnitude. With a 125 MHz clock rate, we obtain a secure key rate of 6.166 kbps over 24.0 dB loss, which is comparable to the state-of-the-art MDI-QKD experimental results with GHz clock rate. Combined with integrated QKD transmitters, a scalable, chip-based and cost-effective QKD network should become realizable in the near future.

quant-ph↗

High-quality quantum process tomography of time-bin qubit's transmission over a metropolitan fiber network and its application

We employ quantum state and process tomography with time-bin qubits to benchmark a city-wide metropolitan quantum communication system. Over this network, we implement real-time feedback control systems for stabilizing the phase of the time-bin qubits, and obtain a 99.3% quantum process fidelity to the ideal channel, indicating the high quality of the whole quantum communication system. This allows us to implement field trial of high performance quantum key distribution using coherent one way protocol with average quantum bit error rate and visibility of 0.25% and 99.2% during 12 hours over 61 km. Our results pave the way for the high-performance quantum network with metropolitan fibers.

quant-ph↗

Co-optimisation and Settlement of Power-Gas Coupled System in Day-ahead Market under Multiple Uncertainties

The interdependency of power systems and natural gas systems is being reinforced by the emerging power-to-gas facilities (PtGs), and the existing gas-fired generators. To jointly improve the efficiency and security under diverse uncertainties from renewable energy resources and load demands, it is essential to co-optimise these two energy systems for day-ahead market clearance. In this paper, a data-driven integrated electricity-gas system stochastic co-optimisation model is proposed. The model is accurately approximated by sequential mixed integer second-order cone programming, which can then be solved in parallel and decentralised manners by leveraging generalised Benders decomposition. Since the price formation and settlement issues have rarely been investigated for integrated electricity-gas systems in an uncertainty setting, a novel concept of expected locational marginal value is proposed to credit the flexibility of PtGs that helps hedging uncertainties. By comparing with a deterministic model and a distributionally robust model, the advantage of the proposed stochastic model and the efficiency of the proposed solution method are validated. Detailed results of pricing and settlement for PtGs are presented, showing that the expected locational marginal value can fairly credit the contribution of PtGs and reflect the system deficiency of capturing uncertainties.

eess.SY↗

A Mixed-Integer SDP Solution Approach to Distributionally Robust Unit Commitment with Second Order Moment Constraints

A power system unit commitment (UC) problem considering uncertainties of renewable energy sources is investigated in this paper, through a distributionally robust optimization approach. We assume that the first and second order moments of stochastic parameters can be inferred from historical data, and then employed to model the set of probability distributions. The resulting problem is a two-stage distributionally robust unit commitment with second order moment constraints, and we show that it can be recast as a mixed-integer semidefinite programming (MI-SDP) with finite constraints. The solution algorithm of the problem comprises solving a series of relaxed MI-SDPs and a subroutine of feasibility checking and vertex generation. Based on the verification of strong duality of the semidefinite programming (SDP) problems, we propose a cutting plane algorithm for solving the MI-SDPs; we also introduce a SDP relaxation for the feasibility checking problem, which is an intractable biconvex optimization. Experimental results on a IEEE 6-bus system are presented, showing that without any tunings of parameters, the real-time operation cost of distributionally robust UC method outperforms those of deterministic UC and two-stage robust UC methods in general, and our method also enjoys higher reliability of dispatch operation.

math.OC↗

A Global Solution Method for Decentralized Multi-Area SCUC and Savings Allocation Based on MILP Value Functions

To address the issue that Lagrangian dual function based algorithms cannot guarantee convergence and global optimality for decentralized multi-area security constrained unit commitment (M-SCUC) problems, a novel decomposition and coordination method using MILP (mixed integer linear programming) value functions is proposed in this paper. Each regional system operator sets the tie-line power injections as variational parameters in its regional SCUC model, and utilizes a finite algorithm to generate a MILP value function, which returns the optimal generation cost for any given interchange scheduling. With the value functions available from all system operators, theoretically, a coordinator is able to derive a globally optimal interchange scheduling. Since power exchanges may alter the financial position of each area considerably from what it would have been via scheduling independently, we then propose a fair savings allocation method using the values functions derived above and the Shapley value in cooperative game theory. Numerical experiments on a two-area 12-bus system and a three-area 457-bus system are carried out. The validness of the value functions based method is verified for the decentralized M-SCUC problems. The outcome of savings allocation is compared with that of the locational marginal cost based method.

eess.SY↗