SearcharxivSearch

arXiv subjects

Kai Guo

Publications and source records attributed to Kai Guo.

At least 19 recordsLinked to original sources

NeuCME: Toward Dynamic Multimodal Continual Learning via Neural Combinatorics of Multiple Experts

Multimodal continual learning has recently shown great potential for developing agents with human-like intelligence by continuously learning new tasks across multiple modalities. However, existing methods typically assume that the set of modalities per task is predefined and fixed. In this paper, we investigate a more realistic learning setting, referred to as dynamic multimodal continual learning, in which the set of modalities may vary across tasks rather than remaining fixed. This setting involves two primary challenges: (i) spatio-temporal catastrophic forgetting and (ii) adaptive multimodal fusion. To address these challenges, we propose NeuCME (as shorthand for \textbf{Neu}ral \textbf{C}ombinatorics of \textbf{M}ultiple \textbf{E}xperts), a novel framework designed to effectively learn and integrate knowledge across tasks with varying modalities. The proposed NeuCME model comprises three key components, namely modality-combinational rehearsal, multi-gated mixture-of-experts, and task relevance-guided distillation. Furthermore, we formulate an evaluation metric to quantify the dynamism of task sequences and then set up a comprehensive benchmark with different degrees of dynamism. Extensive experiments using four real-world datasets demonstrate that the proposed NeuCME outperforms state-of-the-art methods markedly.

cs.LG

Observation of Hong-Ou-Mandel interference between photon and polariton

Light-matter interactions underlie many quantum technologies, yet whether quasiparticles formed from such interactions preserve the full quantum state of light remains unresolved. Surface plasmon polaritons (SPPs), a class of polaritons formed by interacting photons with free-electron oscillations at metal-dielectric interfaces, are prime candidates to explore this question. Here we demonstrate quantum interference between single photons and SPPs using an Au-SiN$_{\mathrm{x}}$ integrated photonic-plasmonic device. Our results reveal that SPPs retain the indistinguishability of their excitation photons, establishing SPP as a viable quantum information carrier and opening a potential route toward photonic-plasmonic quantum circuitry.

quant-ph

Quantum-interference metrology of dissipative Kerr solitons

Dissipative Kerr solitons in optical microresonators underpin chip-scale frequency combs with applications ranging from coherent telecommunications to precision spectroscopy. Yet the characterization of their intrinsic femtosecond temporal structure remains challenging, as the low pulse energy and broad spectral bandwidth necessitate optical amplification and careful dispersion compensation in conventional ultrafast diagnostics, both of which can significantly distort the waveform. Here we demonstrate a quantum-interference metrology of microcomb solitons based on Hong-Ou-Mandel interference. By attenuating the soliton stream to the single-photon level and measuring fourth-order interference, we directly retrieve near transform-limited pulse durations without amplification or dispersion management, remaining accurate even after propagation through 25 km of standard fiber. The same interferogram also provides direct access to the temporal separations in multi-soliton states by converting inter-soliton separations into additional interference dips at corresponding delays, enabling sub-picosecond characterization of their intracavity temporal structure. This quantum-inspired paradigm introduces a fundamentally new metrological approach that is immune to amplification and dispersion distortions, offering a powerful tool for the characterization of complex soliton physics.

quant-ph

Synergy between laser linewidth and frequency chirp in mesospheric magnetometry based on the sodium laser guide star

Mesospheric sodium magnetometry with a laser guide star measures the geomagnetic field near 90~km. Its sensitivity hinges on laser linewidth and chirp, yet prior work optimized these two parameters only separately. We use velocity-resolved density-matrix simulations of Larmor-synchronous pulsed Na D$_2$ pumping to scan both parameters jointly. Linewidth and chirp exhibit a synergy: when chirping carries the recoil-mitigation role, the return flux stays within 1\% of its peak across linewidths of 2--10~MHz. The flux-optimal chirp is $0.19~\mathrm{MHz/\upmu s}$, one fifth of the continuous-wave rate; this synergy guides high-sensitivity mesospheric magnetometer design.

physics.app-ph

Which transition can be used in sodium mirrorless lasing for mesospheric magnetometry?

Directed mirrorless lasing from the mesospheric sodium layer has been proposed as a way to overcome the isotropy of laser-guide-star fluorescence, with demonstrated cell-scale analogues and a demonstrated stand-off magnetometry application. Several transition paths on the Na ladder compete for the same pump photons, and which of them can sustain a population inversion---and under what pumping format---is usually settled by numerical scans of a multi-level rate model. Here the same question is answered in closed form, and three design-relevant results follow from the atomic data alone. (1) A cascade inversion on $u\to l$ requires $\tau_u/\tau_l>(g_u/g_l)b_{ul}$. The inequality reproduces the full continuous-wave classification---admitting $4P_{3/2}\to4S_{1/2}$ (2.21~\textmu m), $4P_{3/2}\to3D_{5/2}$ (9.1~\textmu m) and the fine-structure companion $4S_{1/2}\to3P_{1/2}$ (1138~nm), excluding $4D_{5/2}\to4P_{3/2}$ (2.34~\textmu m) and $3D_{5/2}\to3P_{3/2}$ (819~nm)---and shows the 2.34~\textmu m line to miss by only 9\%, which is why it is transiently accessible. (2) Velocity selectivity reverses the naive scheme ranking because a Doppler-averaged treatment, exact for one-step pumping, underestimates two-step pumping by a participation factor of order the Doppler-to-natural width ratio, $\sim$$10^{2}$; equal division of power between two pump beams is exactly optimal. (3) The transient window on 2.34~\textmu m closes after $\sim$2$\tau(4P)\approx210$~ns, set by the reservoir lifetime rather than by the pump. Each conclusion is traceable to a lifetime, a branching ratio or a linewidth, and transfers to another species without rerunning a model.

physics.app-ph

Optimization of the Repumping Parameters for a Sodium Laser Guide Star Magnetometer

A sodium laser guide star operated as a mesospheric magnetometer modulates a 589 nm laser at the local Larmor frequency and usually diverts a fraction of its power to a repumping light that recovers atoms lost to the dark ground state.The polarization, read out for the most strongly driven velocity group, calls for 2.8 times the flux optimal fraction, and a shot noise figure of merit combining the two observables for 2 times, beyond the range commercial guide star lasers provide.

physics.app-ph

Population-inversion map of the mesospheric sodium ladder: continuous-wave and pulsed pumping schemes for directed emission

Directed mirrorless lasing from the mesospheric sodium layer has been proposed as a way to overcome the isotropy of laser guide star fluorescence, with demonstrated cell-scale analogues and a demonstrated stand-off magnetometry application. Several transition paths on the Na ladder compete for the same pump photons. We build a ten-level rate-equation model of the ladder from NIST transition probabilities and evaluate every electric-dipole line under four continuous-wave pumping schemes, both in a Doppler-averaged treatment and in a velocity-selective treatment appropriate for the collision-poor mesosphere. Three design-relevant results emerge. At practically accessible continuous-wave irradiances the column-gain exponents remain far below unity, consistent with published feasibility estimates; the value of the classification is to identify which lines, schemes, and pulse formats merit further study.

physics.optics

Entanglement-based quantum key distribution with data in hollow-core fiber

The coexistence of quantum information and classical signals in a single fiber is essential for future quantum networks that leverage the well-established optical fiber infrastructure. Although multiplexing technologies can separate quantum and classical signals, pure silica core fibers (PSCFs) remain fundamentally limited by the high nonlinearity, which generates substantial Raman scattering and four-wave mixing noise. Hollow-core fibers (HCFs), guiding light predominantly in air, offer an attractive solution with intrinsically ultra-low nonlinearity and strongly suppressed nonlinear noise. In this work, we demonstrate the entanglement-based key coexisting with data over an 18-km HCF link. We achieve time-encoded high-dimensional quantum key distribution (HD-QKD) carrying 0 dBm of bidirectional received power, corresponding to a theoretical data capacity of up to 2.3 Tbps. During 24 hours of continuous operation, an average secret key rate (SKR) of 10.56 kbps is obtained. Theoretical analysis further predicts SKRs above 135 kbps over transmission distances exceeding 200 km using state-of-the-art low-loss HCFs. These results show significantly improved performance compared with PSCF-based systems and highlight the potential of HCFs for scalable quantum-classical coexistence compatible with the architectures of established fiber-optic networks.

quant-ph

Quantum teleportation over a field-deployed hollow-core fibre network

When a photon and one member of an entangled photon pair are jointly projected onto a Bell-state measurement (BSM), the quantum state of the photon can be transferred to the distant partner of the pair without physically transmitting this information carrier. In real-world deployment, however, teleportation performance is fundamentally bottlenecked by quantum channel impairments, such as loss, noise, and fluctuations, which induce severe decoherence and degrade fidelity. This vulnerability is further exacerbated in scenarios with intense classical data traffic or background light. Realizing scalable quantum networks, therefore, hinges on developing advanced channel architectures capable of supporting both high-fidelity quantum operations and high-capacity classical communications within a shared infrastructure. Towards this end, hollow core fibre (HCF) offers a promising quantum channel resource by combining free-space-like weak light-matter interaction with the stability of fibre-based systems. Here, utilizing a field-deployed metropolitan HCF network spanning three spatially separated nodes in Chengdu, we achieve quantum teleportation with an intermediate BSM under co-propagating classical traffic. Crucially, the HCF links preserve the long-term indistinguishability of photonic qubits without active stabilization, and exhibit a Raman noise approximately three orders of magnitude lower than that of standard solid-core counterparts. This noise suppression enables robust quantum teleportation even alongside classical launch powers up to 160 mW. Our findings establish a classical-data-compatible framework for quantum networking over deployed fibre infrastructure and offer a wavelength-agnostic, plug-and-play, and free-running pathway toward the quantum internet.

quant-ph

Quantum Teleportation toward the Quantum Internet: A Concise Review

Quantum networks play a pivotal role in quantum information science, which not only provide a secure communication platform for remote access to quantum computers but also serve as the strategic core for achieving large-scale quantum information processing, forming the foundational infrastructure for the future global-scale quantum internet. Quantum teleportation, which enables the transmission of unknown quantum states over long distances by employing quantum entanglement together with classical communication, is essential for the distribution of quantum resources in the construction of the global-scale quantum internet. To realize a global-scale quantum internet, quantum repeater protocols represent one of the most promising approaches for enabling quantum communication between any nodes. This concise review presents representative experimental demonstrations of quantum teleportation for constructing quantum networks across different physical platforms. Along this trajectory, the review discusses current challenges, open issues, and future perspectives toward scalable and practical quantum internet.

quant-ph

Extreme-RGMT: Continual Learning of Highly Dynamic Skills for Robust Generalist Humanoid Control

Humans can progressively acquire highly dynamic motor skills while preserving reliable everyday motor abilities. In contrast, existing humanoid controllers face a trade-off between generalist and specialist capabilities: generalist motion tracking policies struggle to reliably execute rare highly dynamic motions, whereas specialist training can degrade previously acquired behaviors. We introduce Extreme-RGMT, a two-stage continual learning framework for robust generalist humanoid control. The method first learns a generalist motion-tracking base policy from diverse multi-source motion data, then employs an asymmetric skill acquisition and capability consolidation mechanism to constrain policy drift on mastered motions while emphasizing difficult dynamic segments. To address the scarcity of highly dynamic motions, their high failure rates, and the resulting shortage of informative samples, Extreme-RGMT combines difficulty-aware sampling with advantage-prioritized trajectory resampling to emphasize critical segments. Experiments show that Extreme-RGMT achieves state-of-the-art generalist whole-body motion-tracking performance, including substantially improved completion of challenging highly dynamic motions. The resulting controller directly executes diverse unseen highly dynamic motions under fixed references and online inertial motion-capture inputs, advancing generalist whole-body motion-tracking controllers toward highly dynamic motor capabilities at the human-expert level.

cs.RO

Quantum LiDAR with non-local modulation

Quantum light detection and ranging (LiDAR) utilizes quantum entanglement and correlation to improve precision, noise resilience and covertness of target detection. Despite recent advances, the development of a quantum LiDAR system that simultaneously achieves high precision and a large measurement range remains challenging. Here, we demonstrate a quantum amplitude-modulated continuous wave LiDAR with micrometer precision achievable via increased acquisition time and meter-scale measurement range. In our demonstration, the signal photons directly illuminate the target, while the idler photons are non-locally modulated with a high-frequency cosine wave and never interact with the target. By leveraging the non-local modulation and the quantum correlation, the target detection is achieved with a precision of 0.64 $\pm$ 0.06 mm within one second over a measurement range of 2-8 m. As the acquisition time is up to 500 s, the system achieves a precision of 29 $\pm\ 4{\ \mathrm{\mu m}}$. Furthermore, our system realizes a 50 times precision improvement over the classical single-photon scheme in a background noise 37 dB stronger than the returned probe photons. With these advantages, our method will open venues for the development of high-precision, long-range, and noise-resilient target detection.

quant-ph

MemTrace: Probing What Final Accuracy Misses in Long-Term Memory

LLM agents increasingly maintain long-term memory of user facts across sessions. Yet such memory is usually evaluated by aggregating accuracy over question rows or episodes. Because this approach scores question rows independently, even when several questions probe the same fact, it cannot show how that fact behaves as conditions change. We introduce MemTrace, a benchmark whose unit of measurement is the knowledge point: a single typed fact about the user, rather than an individual question. MemTrace probes each fact along three controlled dimensions: memory age, defined by how many sessions ago the fact appeared in the history; question type, covering current state, earlier state, and trajectory of change; and evidence condition, covering present, missing, and contradicted-by-false-premise settings. Evaluating 13 memory-system configurations across four paradigms, we find that similar pooled accuracy hides different failures: recovering a fact's current and earlier states does not imply tracking how it changed, and safe abstention does not imply correcting a false premise. The dominant bottleneck is evidence use, not retrieval: when systems fail, the evidence was retrievable 10 times more often than it was missing. These results suggest that improving long-term memory requires better use of reachable evidence, not simply more storage or retrieval.

cs.AI

Magnifying What Matters: Attention-Guided Adaptive Rendering for Visual Text Comprehension

Visual Text Comprehension (VTC) renders text into images for a vision-language model (VLM) to read, sidestepping LLM context-window limits and powering applications from long-page OCR to multi-page memory QA. Yet existing VTC pipelines treat rendering and layout as a fixed, content-agnostic preprocessing step and offer little mechanistic understanding of how VLMs internally process visualized text. Through a focused empirical study on VTC QA tasks, we reveal that VLMs exhibit a localization-without-utilization regime: evidence-localizing attention emerges sharply in the middle-to-late layers and is largely decoupled from answer correctness, yet simply enlarging the localized spans on the rendered page recovers a large fraction of the failures. Building on these observations, we propose AGAR (Attention-Guided Adaptive Rendering), a training-free, model-agnostic method that leverages a VLM's own middle-to-late layer attention to identify the top-K important visual patches, maps them back to word spans, and re-renders the page with those spans enlarged before re-inferring the answer. Extensive experiments across nine VTC benchmarks (short-form, long-context, and multi-page memory QA) and four VLM backbones show that AGAR (i)consistently improves off-the-shelf VLMs as a plug-and-play enhancement, (ii)composes with VLM post-training to yield further gains, and (iii)remains robust under both visual- and text-side input degradation.

cs.CV

Exploring Cross-Scenario Generality of Agentic Memory Systems: Diagnostics and a Strong Baseline

LLM agents accumulate histories that outgrow their context windows, motivating a growing literature on memory systems. Yet most existing designs are tuned to a single scenario (multi-session chat or a single trajectory format), and there is little evidence that they generalize across the heterogeneous trajectories agents encounter in deployment. We revisit eight memory systems plus an agentic harness for search problems, on five scenarios: single-turn QA, multi-session chat, agentic-trajectory QA, memory stress tests, and long-horizon agentic tasks. The harness, which self-manages flat text-file storage via tool calls, achieves the best cross-task ranking, suggesting that memory performance hinges on giving the agent active control over storage and retrieval rather than on a passive store behind a fixed pipeline. We instantiate this insight in AutoMEM, an agentic memory harness with a self-managed tool interface that achieves the best cross-scenario generality among the systems we evaluate.

cs.AI

OpenRFM: Dissecting Relational In-Context Learning

Relational Foundation Models (RFMs) promise a single pre-trained predictor that, given any relational database, returns predictions in one forward pass via relational in-context learning (ICL). Yet a substantial gap separates open RFMs from their commercial counterparts, and the origin of this gap has not been systematically understood. We dissect a representative framework, the Relational Transformer (RT), from two perspectives. Model side: we show that RT performs relation-level ICL, and a kernel regression view shows it fails when sparse label-cell coverage yields an underdetermined regression. Data side: we ablate RT's pre-training source and find that existing synthetic-only pre-training and in-distribution pre-training drive the same architecture into different regimes, lazy vs. feature-learning. Probing this gap reveals that the missing ingredient is a support-identifiable relational latent in the label-generation process. These two diagnoses translate into (1) a dual-stage ICL architecture that combines the relational backbone with a batch-level ICL layer lifted from a pre-trained tabular foundation model to overcome relation-level label scarcity, and (2) a homophily-aware synthetic plus continual real-data pre-training mixture, augmented with a prototype-based regularization. These choices define OpenRFM, a simple yet effective RFM that improves average task performance by approximately 30% over the RT backbone and surpasses the commercial model KumoRFMv1 on a large set of evaluation tasks.

cs.LG

Quantum light source with lithium tantalate for scalable photonic quantum circuits

Thin-film lithium tantalate (TFLT) has emerged as a promising integrated photonic platform owing to its low photorefractive noise, high optical damage threshold, and reduced birefringence, attracting increasing interest for scalable photonic technologies. Here, to the best of our knowledge, we demonstrate the first quantum light source with TFLT via spontaneous four-wave mixing, bridging the gap between the rapidly advancing classical TFLT ecosystem and integrated quantum photonics. The fabricated microring exhibits a free spectral range of 350~GHz and an optical quality factor of $10^6$, enabling efficient cavity-enhanced nonlinear interactions. Correlated photon pairs are generated across the telecom band from 1510 to 1570~nm, with a photon pair generation rate of 24 $\mathrm{MHz/mW^{2}}$ at a wavelength of 1535.04 nm. The source delivers strongly antibunched heralded single photons with $g^{(2)}_{H}(0)=0.071\pm0.004$ at a heralding rate of 170 kHz, while the unheralded statistics yield $g^{(2)}(0)=1.93 \pm 0.05$, indicating near-single-temporal-mode emission. Energy-time entanglement is further confirmed by a raw two-photon interference visibility of $92.55\pm0.94\%$, well above the Bell-inequality violation threshold. These results establish TFLT as a manufacturing-compatible platform for scalable photonic quantum circuits, paving the way for the monolithic co-integration of classical and quantum photonic functionalities.

quant-ph

Integrated time-bin entangled quantum light source on a 4H-SiC microring chip

Integrated time-bin-entangled photon-pair source with cavity-enhanced nonlinear optical processes is essential for quantum information technologies. However, microcavities with a high quality factor inherently introduce a trade-off between generation efficiency and photon bandwidth, which hinders the development of high-speed quantum networks with an integrated source. Here, we address this challenge by optimizing the nonlinearity property of the material and the geometry of the integrated microring resonator with a 4H-silicon carbide platform. Operating at a loaded quality factor of 1.9 $\times$ 10^5 - spectral bandwidth of 1.0 GHz and pumped with 300-ps double pulses separated by 1.25 ns at a repetition rate of 160 MHz, the device achieves a time-bin-entangled photon-pair generation rate of 1.35 $\times$ 10^7 s^-1 mW^-2. A raw visibility of 95.55 $\pm$ 0.18% is measured, showing a violation of Bell's inequality by more than 138 standard deviations, and a fidelity of 94.37 $\pm$ 0.22% is obtained by quantum state tomography. These results provide a scalable pathway to an efficient and broadband time-bin entangled quantum light source, overcoming intrinsic limitations of cavity-based designs and advancing integrated platforms for future quantum communication networks.

quant-ph