SearcharxivSearch

arXiv subjects

Xiao Gong

Publications and source records attributed to Xiao Gong.

At least 19 recordsLinked to original sources

Integrated photonic platform with high-speed entanglement generation and witnessing

High-speed generation and efficient entanglement detection on a photonic chip are essential for quantum information applications but hard to achieve due to common photonic chips' material properties and limited component performance. In this work, we experimentally demonstrate entanglement witness on a silicon photonic chip, with multi-rail single-photon entanglement generation based on decoy-state techniques. The detection is based on balanced homodyne detectors on the same photonic chip with a bandwidth of up to 12.5 GHz, which allows room-temperature operation. A loss-equivalent analysis method compensates for optical losses and system noises. Experimental results quantify an entangled state fidelity of 92% in quantum state tomography and a Clauser-Horne-Shimony-Holt (CHSH) violation lower bound of 2.59. These results establish a viable path toward fully integrated, high-bandwidth, room-temperature quantum photonic systems, with potential applications in on-chip quantum optics and quantum random number generation.

quant-ph

Optoelectronically Active GaAs/GeSn-MQW/Ge Heterojunctions Created via Semiconductor Grafting

Traditionally, advancements in semiconductor devices have been driven by lattice-matched heterojunctions with tailored band alignments through heteroepitaxy techniques. However, there is significant interest in expanding the capabilities of heterojunction devices, in particular utilizing extreme lattice mismatches. We demonstrate the manipulation of device behaviors and performance enhancement achievable through a lattice-mismatched, single-crystalline GaAs/GeSn-multi-quantum well (MQW)/Ge n-i-p heterojunction by employing advanced semiconductor grafting technology. With engineered band alignment and optical field distribution, the grafted GaAs/GeSn-MQW/Ge n-i-p photodiode achieved outstanding performance: a record-low dark current density of 1.22E10^-7 A/cm^2, an extended spectral response from ~0.5 to 2 um, and improved photoresponsivity of RVIS of 0.85 A/W and RNIR of 0.40 A/W at 520 and 1570 nm, respectively. The dark current density is at least 5 orders of magnitude lower than state-of-the-art GeSn photodiodes. The photoresponsivity demonstrates an approximately sevenfold enhancement in the VIS range and a threefold improvement in the NIR range compared to the reference epitaxial photodiode. This work presents a unique strategy for constructing lattice-mismatched semiconductor heterojunction devices. More importantly, the implications transcend the current GaAs/GeSn-MQW/Ge example, offering potential applications in other material systems and freeing device design from the stringent lattice-matching constraints of conventional heteroepitaxy.

cond-mat.mtrl-sci

Solving Boolean Satisfiability Problems Using A Hypergraph-based Probabilistic Computer

Boolean Satisfiability (SAT) problems are critical in fields such as artificial intelligence and cryptography, where efficient solutions are essential. Conventional probabilistic solvers often encounter scalability issues due to complex logic synthesis steps. In this work, we present a novel approach for solving the 3-SAT Boolean satisfiability problem using hypergraph-based probabilistic computers obtained through direct mapping. This method directly translates 3-SAT logical expressions into hypergraph structures, thereby circumventing conventional logic decomposition and synthesis procedures, and offering a more streamlined solver architecture. For representative uf100-430 instances, the proposed approach reduces the node count from 631 to 100 and the edge count from ~2,423 to ~1,013. Under identical simulated annealing conditions, the conventional simple undirected graph (SUG)-based solver achieves a 0% success rate across the tested instances, whereas the hypergraph-based solver attains an average success rate of ~77.6%. In addition, the hypergraph-based method reaches an average minimum energy of ~0.24, close to the theoretical ground state, while the SUG-based architecture remains trapped at substantially higher energy levels (~9.12 on average). The direct hypergraph mapping can further be extended to k-SAT formulations, providing a scalable framework for more complex satisfiability problems in probabilistic computing.

physics.comp-ph

A Full Spectrum of 3D Ferroelectric Memory Architectures Shaped by Polarization Sensing

Ferroelectric memories have attracted significant interest due to their non-volatile storage, energy efficiency, and fast operation, making them prime candidates for future memory technologies. As commercial Dynamic Random Access Memory (DRAM) and NAND flash memory are transiting or have moved toward three-dimensional (3D) integration, 3D ferroelectric memory architectures are also emerging, provided they can achieve a competitive position within the modern memory hierarchy. Given the excellent scalability of ferroelectric HfO2, various dense 3D integrated ferroelectric memory architectures are feasible, each offering unique strengths and facing distinct challenges. In this work, we present a comprehensive classification of 3D ferroelectric memory architectures based on polarization sensing methods, highlighting their critical role in shaping memory cell design and operational efficiency. Through a systematic evaluation of these architectures, we develop a unified framework to assess their advantages and trade-offs. This classification not only enhances the understanding of current 3D ferroelectric memory technologies but also lays the foundation for designing next-generation architectures optimized for advanced computing and high-performance applications.

cs.ET

An In-Situ Spatial-Temporal Sequence Detector for Neuromorphic Vision Sensor Empowered by High Density Vertical NAND Storage

Neuromorphic vision sensors require efficient real-time pattern recognition, yet conventional architectures struggle with energy and latency constraints. Here, we present a novel in-situ spatiotemporal sequence detector that leverages vertical NAND storage to achieve massively parallel pattern detection. By encoding each cell with two single-transistor-based multi-level cell (MLC) memory elements, such as ferroelectric field-effect transistors (FeFETs), and mapping a pixel's temporal sequence onto consecutive word lines (WLs), we enable direct temporal pattern detection within NAND strings. Each NAND string serves as a dedicated reference for a single pixel, while different blocks store patterns for distinct pixels, allowing large-scale spatial-temporal pattern recognition via simple direct bit-line (BL) sensing, a well-established operation in vertical NAND storage. We experimentally validate our approach at both the cell and array levels, demonstrating that vertical NAND-based detector achieves more than six orders of magnitude improvement in energy efficiency and more than three orders of magnitude reduction in latency compared to conventional CPU-based methods. These findings establish vertical NAND storage as a scalable and energy-efficient solution for next-generation neuromorphic vision processing.

cs.ET

The First Hardware Demonstration of a Universal Programmable RRAM-based Probabilistic Computer for Molecular Docking

Molecular docking is a critical computational strategy in drug design and discovery, but the complex diversity of biomolecular structures and flexible binding conformations create an enormous search space that challenges conventional computing methods. Although quantum computing holds promise for these challenges, it remains constrained by scalability, hardware limitations, and precision issues. Here, we report a prototype of a probabilistic computer (p-computer) that efficiently and accurately solves complex molecular docking for the first time, overcoming previously encountered challenges. At the core of the system is a p-computing chip based upon our artificial tunable probabilistic bits (p-bits), which are compatible with computing-in-memory schemes, based upon 180 nm CMOS technology and BEOL HfO2 RRAM. We successfully demonstrated the superior performance of the p-computer in practical ligand-protein docking scenarios. A 42-node molecular docking problem of lipoprotein with LolA-LolCDE complex-a key point in developing antibiotics against Gram-negative bacteria, was successfully solved. Our results align well with the Protein-Ligand Interaction Profiler tool. This work marks the first application of p-computing in molecular docking-based computational biology, which has great potential to overcome the limitations in success rate and efficiency of current technologies in addressing complex bioinformatics problems.

physics.comp-ph

A Novel P-bit-based Probabilistic Computing Approach for Solving the 3-D Protein Folding Problem

In the post-Moore era, the need for efficient solutions to non-deterministic polynomial-time (NP) problems is becoming more pressing. In this context, the Ising model implemented by the probabilistic computing systems with probabilistic bits (p-bits) has attracted attention due to the widespread availability of p-bits and support for large-scale simulations. This study marks the first work to apply probabilistic computing to tackle protein folding, a significant NP-complete problem challenge in biology. We represent proteins as sequences of hydrophobic (H) and polar (P) beads within a three-dimensional (3-D) grid and introduce a novel many-body interaction-based encoding method to map the problem onto an Ising model. Our simulations show that this approach significantly simplifies the energy landscape for short peptide sequences of six amino acids, halving the number of energy levels. Furthermore, the proposed mapping method achieves approximately 100 times acceleration for sequences consisting of ten amino acids in identifying the correct folding configuration. We predicted the optimal folding configuration for a peptide sequence of 36 amino acids by identifying the ground state. These findings highlight the unique potential of the proposed encoding method for solving protein folding and, importantly, provide new tools for solving similar NP-complete problems in biology by probabilistic computing approach.

physics.app-ph

Exploring the distribution of connectivity weights in resting-state EEG networks

The resting-state brain networks (RSNs) reflects the functional connectivity patterns between brain modules, providing essential foundations for decoding intrinsic neural information within the brain. It serves as one of the primary tools for describing the spatial dynamics of the brain using various neuroimaging techniques, such as electroencephalography (EEG) and magnetoencephalography (MEG). However, the distribution rules or potential modes of functional connectivity weights in the resting state remain unclear. In this context, we first start from simulation, using forward solving model to generate scalp EEG with four channel densities (19, 32, 64, 128). Subsequently, we construct scalp brain networks using five coupling measures, aiming to explore whether different channel density or coupling measures affect the distribution pattern of functional connectivity weights. Next, we quantify the distribution pattern by calculating the skewness, kurtosis, and Shannon entropy of the functional connectivity network weights. Finally, the results of the simulation were validated in a normative database. We observed that: 1) The functional connection weights exhibit a right-skewed distribution, and are not influenced by channel density or coupling measures; 2) The functional connection weights exhibit a relatively uniform distribution, with the potential for volume conduction to affect the degree of uniformity in the distribution; 3) Networks constructed using coupling measures influenced by volume conduction exhibit significant correlations between the average connection weight and measures of skewness, kurtosis, and Shannon entropy. This study contributes to a deeper understanding of RSNs, providing valuable insights for research in the field of neuroscience, and holds promise for being associated with brain cognition and disease diagnosis.

cs.HC

Deep-Learning-Aided Alternating Least Squares for Tensor CP Decomposition and Its Application to Massive MIMO Channel Estimation

CANDECOMP/PARAFAC (CP) decomposition is the mostly used model to formulate the received tensor signal in a massive MIMO system, as the receiver generally sums the components from different paths or users. To achieve accurate and low-latency channel estimation, good and fast CP decomposition (CPD) algorithms are desired. The CP alternating least squares (CPALS) is the workhorse algorithm for calculating the CPD. However, its performance depends on the initializations, and good starting values can lead to more efficient solutions. Existing initialization strategies are decoupled from the CPALS and are not necessarily favorable for solving the CPD. This paper proposes a deep-learning-aided CPALS (DL-CPALS) method that uses a deep neural network (DNN) to generate favorable initializations. The proposed DL-CPALS integrates the DNN and CPALS to a model-based deep learning paradigm, where it trains the DNN to generate an initialization that facilitates fast and accurate CPD. Moreover, benefiting from the CP low-rankness, the proposed method is trained using noisy data and does not require paired clean data. The proposed DL-CPALS is applied to millimeter wave MIMO-OFDM channel estimation. Experimental results demonstrate the significant improvements of the proposed method in terms of both speed and accuracy for CPD and channel estimation.

eess.SP

Self-testing quantum randomness expansion on an integrated photonic chip

The power of quantum random number generation is more than just the ability to create truly random numbers$\unicode{x2013}$it can also enable self-testing, which allows the user to verify the implementation integrity of certain critical quantum components with minimal assumptions. In this work, we develop and implement a self-testing quantum random number generator (QRNG) chipset capable of generating 15.33 Mbits of certifiable randomness in each run (an expansion rate of $5.11\times 10^{-4}$ at a repetition rate of 10 Mhz). The chip design is based on a highly loss-and-noise tolerant measurement-device-independent protocol, where random coherent states encoded using quadrature phase shift keying are used to self-test the quantum homodyne detection unit: well-known to be challenging to characterise in practice. Importantly, this proposal opens up the possibility to implement miniaturised self-testing QRNG devices at production scale using standard silicon photonics foundry platforms.

quant-ph

A Ferroelectric Compute-in-Memory Annealer for Combinatorial Optimization Problems

Computationally hard combinatorial optimization problems (COPs) are ubiquitous in many applications, including logistical planning, resource allocation, chip design, drug explorations, and more. Due to their critical significance and the inability of conventional hardware in efficiently handling scaled COPs, there is a growing interest in developing computing hardware tailored specifically for COPs, including digital annealers, dynamical Ising machines, and quantum/photonic systems. However, significant hurdles still remain, such as the memory access issue, the system scalability and restricted applicability to certain types of COPs, and VLSI-incompatibility, respectively. Here, a ferroelectric field effect transistor (FeFET) based compute-in-memory (CiM) annealer is proposed. After converting COPs into quadratic unconstrained binary optimization (QUBO) formulations, a hardware-algorithm co-design is conducted, yielding an energy-efficient, versatile, and scalable hardware for COPs. To accelerate the core vector-matrix-vector (VMV) multiplication of QUBO formulations, a FeFET based CiM array is exploited, which can accelerate the intended operation in-situ due to its unique three-terminal structure. In particular, a lossless compression technique is proposed to prune typically sparse QUBO matrix to reduce hardware cost. Furthermore, a multi-epoch simulated annealing (MESA) algorithm is proposed to replace conventional simulated annealing for its faster convergence and better solution quality. The effectiveness of the proposed techniques is validated through the utilization of developed chip prototypes for successfully solving graph coloring problem, indicating great promise of FeFET CiM annealer in solving general COPs.

cs.ET

Embedding Security into Ferroelectric FET Array via In-Situ Memory Operation

Non-volatile memories (NVMs) have the potential to reshape next-generation memory systems because of their promising properties of near-zero leakage power consumption, high density and non-volatility. However, NVMs also face critical security threats that exploit the non-volatile property. Compared to volatile memory, the capability of retaining data even after power down makes NVM more vulnerable. Existing solutions to address the security issues of NVMs are mainly based on Advanced Encryption Standard (AES), which incurs significant performance and power overhead. In this paper, we propose a lightweight memory encryption/decryption scheme by exploiting in-situ memory operations with negligible overhead. To validate the feasibility of the encryption/decryption scheme, device-level and array-level experiments are performed using ferroelectric field effect transistor (FeFET) as an example NVM without loss of generality. Besides, a comprehensive evaluation is performed on a 128x128 FeFET AND-type memory array in terms of area, latency, power and throughput. Compared with the AES-based scheme, our scheme shows around 22.6x/14.1x increase in encryption/decryption throughput with negligible power penalty. Furthermore, we evaluate the performance of our scheme over the AES-based scheme when deploying different neural network workloads. Our scheme yields significant latency reduction by 90% on average for encryption and decryption processes.

cs.ET

Probabilistic-Bits based on Ferroelectric Field-Effect Transistors for Stochastic Computing

A probabilistic-bit (p-bit) is the fundamental building block in the circuit network of a stochastic computing, and it could produce a continuous random bit-stream with tunable probability. Utilizing the stochasticity in few-domain ferroelectric material(FE), we propose for the first time, the p-bits based on ferroelectric FET. The stochasticity of the FE p-bits stems from the thermal noise-induced lattice vibration, which renders dipole fluctuations and is tunable by an external electric field. The impact of several key FE parameters on p-bits' stochasticity is evaluated, where the domain properties are revealed to play crucial roles. Furthermore, the integer factorization based on FE p-bits circuit network is performed to verify its functionality, and the accuracy is found to depend on FE p-bits' stochasticity. The proposed FE p-bits possess the advantages of both extremely low hardware coast and the compatibility with CMOS-technology, rendering it a promising candidate for stochastic computing applications.

cs.ET

Eliminating Leakage in Volatile Memory with Anti-Ferroelectric Transistors

Cache serves as a temporary data memory module in many general-purpose processors and domain-specific accelerators. Its density, power, speed, and reliability play a critical role in enhancing the overall system performance and quality of service. Conventional volatile memories, including static random-access memory (SRAM) and embedded dynamic random-access memory (eDRAM) in the complementary metal-oxide-semiconductor technology, have high performance and good reliability. However, the inherent leakage in both SRAM and eDRAM hinders further improvement towards smaller feature sizes and higher energy efficiency. Although the emerging nonvolatile memories can eliminate the leakage efficiently, the penalties of lower speed and degraded reliability are significant. This article reveals a new opportunity towards leakage-free volatile static memory beyond the known paradigms of existing volatile and nonvolatile memories. By engineering a double-well energy landscape with the assistance of a clamping voltage bias, leakage-free and refresh-free state retention of volatile memory is achieved for the first time. This new memory is highlighted by both the ultra-low leakage of nonvolatile memories and the speed, energy, and reliability advantages of volatile memories. A proof-of-concept memory is demonstrated using in-house anti-ferroelectric field-effect transistors (AFeFETs), delivering an extrapolated endurance of about 1012 cycles, a retention time of over 10 years, and no subthreshold channel leakage current. Such a new concept of AFeFET-based memory enables an improved balance between density, power, and reliability beyond all existing memory solutions.

cs.ET

Computational Associative Memory with Amorphous InGaZnO Channel 3D NAND-Compatible FG Transistors

3D NAND enables continuous NAND density and cost scaling beyond conventional 2D NAND. However, its poly-Si channel suffers from low mobility, large device variations, and instability caused by grain boundaries. Here, we overcome these drawbacks by introducing an amorphous indium-gallium-zinc-oxide (a-IGZO) channel, which has the advantages of ultra-low OFF current, back-end-of-line compatibility, higher mobility and better uniformity than poly-Si, and free of grain boundaries due to the amorphous nature. Ultra-scaled floating-gate (FG) transistors with a channel length of 60 nm are reported, achieving the highest ON current of 127 uA/um among all reported a-IGZO-based flash devices for high-density, low-power, and high-performance 3D NAND applications. Furthermore, a non-volatile and area-efficient ternary content-addressable memory (TCAM) with only two a-IGZO FG transistors is experimentally demonstrated. Array-level simulations using experimentally calibrated models show that this design achieves at least 240x array-size scalability and 2.7-fold reduction in search energy than 16T-CMOS, 2T2R, and 2FeFET TCAMs.

cond-mat.mes-hall

Ferroelectric FET based Context-Switching FPGA Enabling Dynamic Reconfiguration for Adaptive Deep Learning Machines

Field Programmable Gate Array (FPGA) is widely used in acceleration of deep learning applications because of its reconfigurability, flexibility, and fast time-to-market. However, conventional FPGA suffers from the tradeoff between chip area and reconfiguration latency, making efficient FPGA accelerations that require switching between multiple configurations still elusive. In this paper, we perform technology-circuit-architecture co-design to break this tradeoff with no additional area cost and lower power consumption compared with conventional designs while providing dynamic reconfiguration, which can hide the reconfiguration time behind the execution time. Leveraging the intrinsic transistor structure and non-volatility of ferroelectric FET (FeFET), compact FPGA primitives are proposed and experimentally verified, including 1FeFET look-up table (LUT) cell, 1FeFET routing cell for connection blocks (CBs) and switch boxes (SBs). To support dynamic reconfiguration, two local copies of primitives are placed in parallel, which enables loading of arbitrary configuration without interrupting the active configuration execution. A comprehensive evaluation shows that compared with the SRAM-based FPGA, our dynamic reconfiguration design shows 63.0%/71.1% reduction in LUT/CB area and 82.7%/53.6% reduction in CB/SB power consumption with minimal penalty in the critical path delay (9.6%). We further implement a Super-Sub network model to show the benefit from the context-switching capability of our design. We also evaluate the timing performance of our design over conventional FPGA in various application scenarios. In one scenario that users switch between two preloaded configurations, our design yields significant time saving by 78.7% on average. In the other scenario of implementing multiple configurations with dynamic reconfiguration, our design offers time saving of 20.3% on average.

cs.AR

A Novel Non-Volatile Inverter-based CiM: Continuous Sign Weight Transition and Low Power on-Chip Training

In this work, we report a novel design, one-transistor-one-inverter (1T1I), to satisfy high speed and low power on-chip training requirements. By leveraging doped HfO2 with ferroelectricity, a non-volatile inverter is successfully demonstrated, enabling desired continuous weight transition between negative and positive via the programmable threshold voltage (VTH) of ferroelectric field-effect transistors (FeFETs). Compared with commonly used designs with the similar function, 1T1I uniquely achieves pure on-chip-based weight transition at an optimized working current without relying on assistance from off-chip calculation units for signed-weight comparison, facilitating high-speed training at low power consumption. Further improvements in linearity and training speed can be obtained via a two-transistor-one-inverter (2T1I) design. Overall, focusing on energy and time efficiencies, this work provides a valuable design strategy for future FeFET-based computing-in-memory (CiM).

cond-mat.mes-hall

Securing practical quantum communication systems with optical power limiters

Controlling the energy of unauthorized light signals in a quantum cryptosystem is an essential criterion for implementation security. Here, we propose a passive optical power limiter device based on thermo-optical defocusing effects providing a reliable power limiting threshold which can be readily adjusted to suit various quantum applications. In addition, the device is robust against a wide variety of signal variations (e.g. wavelength, pulse width), which is important for implementation security. Moreover, we experimentally show that the proposed device does not compromise quantum communication signals, in that it has only a very minimal impact (if not, negligible impact) on the intensity, phase, or polarization degrees of freedom of the photon, thus making it suitable for general communication purposes. To show its practical utility for quantum cryptography, we demonstrate and discuss three potential applications: (1) measurement-device-independent quantum key distribution with enhanced security against a general class of Trojan-horse attacks, (2) using the power limiter as a countermeasure against bright illumination attacks, and (3) the application of power limiters to potentially enhance the implementation security of plug-and-play quantum key distribution.

quant-ph