Searcharxiv⌕ Search

arXiv subjects

Cheng Li

Publications and source records attributed to Cheng Li.

At least 109 records · Page 6Linked to original sources

Fast control of the transverse structure of a light beam using acousto-optic modulators

Fast, reprogrammable control over the transverse structure of light beams plays an essential role in applications such as structured illumination microscopy, optical trapping, and quantum information processing. Existing technologies, such as liquid crystal on silicon spatial light modulators (LCoS-SLMs) and digital micromirror devices (DMDs), suffer from limited refresh rates, low damage thresholds, and high insertion loss. Acousto-optic modulators (AOMs) can resolve the above issues, as they typically handle higher laser power and offer lower insertion loss. By effectively mapping the temporal radio-frequency (RF) waveforms onto the spatial diffraction patterns of the optical field, individual AOMs have been shown to generate one-dimensional (1D) spatial modes at a pixel refresh rate of nearly 20 MHz. We extend this concept to enable fast modulation in a two-dimensional (2D) space using a double-AOM scheme. We demonstrate the generation of 2D Hermite-Gaussian (HG_nm) modes with an average fidelity of 81%, while the highest-order mode generated, HG_53, retains a fidelity of 56%.

physics.optics↗

Enhanced Open-Source NWDAF for Event-Driven Analytics in 5G Networks

The network data analytics function (NWDAF) has been introduced in the fifth-generation (5G) core standards to enable event-driven analytics and support intelligent network automation. However, existing implementations remain largely proprietary, and open-source alternatives lack comprehensive support for end-to-end event subscription and notification. In this paper, we present an open source NWDAF framework integrated into an existing Free5GC implementation, which serves as an open-source 5G core implementation. Our implementation extends the session management function to support standardized event exposure interfaces and introduces custom-built notification mechanisms into the SMF and the access and mobility management function for seamless data delivery. The NWDAF subscribes to events and generates analytics on user equipment (UE) behavior, session lifecycle, and handover dynamics. We validate our system through a two-week deployment involving four virtual next-generation NodeBs (gNBs) and multiple virtual UEs with dynamic mobility patterns. To demonstrate predictive capabilities, we incorporate a mobility-aware module that achieves 80.65\% accuracy in forecasting the next gNB handover cell. The framework supports reliable UE registration, state tracking, and cross-cell handovers.

cs.NI↗

Crystal-KV: Efficient KV Cache Management for Chain-of-Thought LLMs via Answer-First Principle

Chain-of-Thought (CoT) reasoning in large language models (LLMs) significantly improves accuracy on complex tasks, yet incurs excessive memory overhead due to the long think-stage sequences stored in the Key-Value (KV) cache. Unlike traditional generation tasks where all tokens are uniformly important, CoT emphasizes the final answer, rendering conventional KV compression strategies ineffective. In this paper, we present Crystal-KV, an efficient KV cache management framework tailored for CoT reasoning. Our key insight is the answer-first principle. By mapping answer preferences into think-stage attention map, we distinguish between SlipKV, which mainly maintains the reasoning flow but may occasionally introduce misleading context, and CrystalKV, which truly contributes to the correctness of the final answer. Next, we propose an attention-based Least Recently Frequently Used algorithm. It precisely identifies when a SlipKV entry's utility expires and evicts it, retaining CrystalKV without disrupting reasoning flow. Finally, we introduce an adaptive cache budget allocation algorithm. Based on the dynamic proportion of CrystalKV, it estimates the importance of each layer/head and adjusts the KV cache budget during inference, amplifying critical components to improve budget utilization. Results show that Crystal-KV achieves state-of-the-art KV cache compression, significantly improves throughput, and enables faster response time, while maintaining, or even improving, answer accuracy for CoT reasoning.

cs.CL↗

Tracing the Flow of Knowledge From Science to Technology Using Deep Learning

We develop a language similarity model suitable for working with patents and scientific publications at the same time. In a horse race-style evaluation, we subject eight language (similarity) models to predict credible Patent-Paper Citations. We find that our Pat-SPECTER model performs best, which is the SPECTER2 model fine-tuned on patents. In two real-world scenarios (separating patent-paper-pairs and predicting patent-paper-pairs) we demonstrate the capabilities of the Pat-SPECTER. We finally test the hypothesis that US patents cite papers that are semantically less similar than in other large jurisdictions, which we posit is because of the duty of candor. The model is open for the academic community and practitioners alike.

cs.CL↗

Early-stopping for Transformer model training

This work, based on Random Matrix Theory (RMT), introduces a novel early-stopping strategy for Transformer training dynamics. Utilizing the Power Law (PL) fit to tansformer attention matrices as a probe, we demarcate training into three stages: structural exploration, heavy-tailed structure stabilization, and convergence saturation. Empirically, we observe that the spectral density of the shallow self-attention matrix $V$ consistently evolves into a heavy-tailed distribution. Crucially, we propose two consistent and validation-set-free criteria: a quantitative metric for heavy-tailed dynamics and a novel spectral signature indicative of convergence. The strong alignment between these criteria highlights the utility of RMT for monitoring and diagnosing the progression of Transformer model training.

cs.LG↗

Scalable Distributed Vector Search via Accuracy Preserving Index Construction

Scaling Approximate Nearest Neighbor Search (ANNS) to billions of vectors requires distributed indexes that balance accuracy, latency, and throughput. Yet existing index designs struggle with this tradeoff. This paper presents SPIRE, a scalable vector index based on two design decisions. First, it identifies a balanced partition granularity that avoids read-cost explosion. Second, it introduces an accuracy-preserving recursive construction that builds a multi-level index with predictable search cost and stable accuracy. In experiments with up to 8 billion vectors across 46 nodes, SPIRE achieves high scalability and up to 9.64X higher throughput than state-of-the-art systems.

cs.DC↗

Angular dependence of third-order law in anisotropic MHD turbulence

In solar wind turbulence, the energy transfer/dissipation rate is typically estimated using MHD third-order structure functions calculated using spacecraft observations. However, the inherent anisotropy of solar wind turbulence leads to significant variations in structure functions along different observational directions, thereby affecting the accuracy of energy-dissipation rate estimation. An unresolved issue is how to optimise the selection of observation angles under limited directional sampling to improve estimation precision. We conduct a series of MHD turbulence simulations with different mean magnetic field strengths, $ B_0 $. Our analysis of the third-order structure functions reveals that the global energy dissipation rate estimated around a polar angle of $ θ= 60^\circ$ agrees reasonably with the exact one for $ 0 \le B_0/b_{rms} \le 5 $, where $b_{rms}$ denotes the root-mean-square magnetic field fluctuation. The speciality of $60^\circ$ polar angle can be understood by the Mean Value Theorem of Integrals, since the spherical integral of the polar-angle component ($\widetilde{T_θ}$) of the divergence of Yaglom flux is zero, and $\widetilde{T_θ}$ changes sign around 60$^\circ$. Existing theory on the energy flux vector as a function of the polar angle is assessed, and supports the speciality of $60^\circ$ polar angle. The angular dependence of the third-order structure functions is further assessed with virtual spacecraft data analysis. The present results can be applied to measure the turbulent dissipation rates of energy in the solar wind, which are of potential importance to other areas in which turbulence takes place, such as laboratory plasmas and astrophysics.

physics.space-ph↗

Turning Noise into Value: Uncovering Service Preferences from Ambiguous Interaction in E-commerce

In e-commerce service recommendation, utilizing auxiliary behaviors to alleviate data sparsity often relies on the flawed assumption that auxiliary behaviors that fail to trigger target actions are negative samples. This approach is fundamentally flawed as it ignores false negatives where users actually harbor latent intent or interest but have not yet converted due to external factors. Consequently, existing methods suffer from sample selection bias and a severe distribution shift between the auxiliary and target behaviors, leading to the erroneous suppression of potential user needs. To address these challenges, we propose a Noise-to-Value Adapter (NoVa), an e-commerce service recommendation framework that re-examines the problem through the lens of positive-unlabeled learning. Instead of treating ambiguous auxiliary behaviors as definite negatives, NoVa aims to uncover high-quality preferences from noise via two key mechanisms. First, to bridge the distribution gap, we employ adversarial feature alignment. This module aligns the auxiliary behavior distribution with the target space to identify high-confidence false negatives, which are instances that statistically resemble confirmed target behaviors and thus represent latent conversion intents. Second, to mitigate label noise caused by accidental clicks or random browsing, we introduce a semantic consistency constraint. This mechanism implements semantic-aware filtering based on the content similarity of services, acting as a bias correction step to filter out low-confidence interactions that lack semantic relevance to historical user preferences. Extensive experiments on three real-world datasets demonstrate that NoVa outperforms state-of-the-art baselines.

cs.IR↗

Scalable and Provable Kemeny Constant Computation on Static and Dynamic Graphs: A 2-Forest Sampling Approach

Kemeny constant, defined as the expected hitting time of random walks from a source node to a randomly chosen target node, is a fundamental metric in graph data management with many real-world applications. However, computing it exactly on large graphs is highly challenging, as it requires inverting large graph matrices. Existing solutions mainly rely on approximate random-walk-based methods, which still need large sample sizes and lack strong theoretical guarantees. In this paper, we propose a new approach for approximating the Kemeny constant via 2-forest sampling. We first derive an unbiased estimator expressed through spanning trees by introducing a path mapping technique that establishes a direct correspondence between spanning trees and certain classes of 2-forests. Compared to random walk-based estimators, 2-forest-based estimators yield leads to a better theoretical bound. We further design efficient algorithms to sample and traverse spanning trees, leveraging data structures such as the Binary Indexed Tree (BIT) for optimization. Our theoretical analysis shows that the Kemeny constant can be approximated with relative error $ε$ in $O\left(\frac{Δ^2\bar{d}^2}{ε^2}(τ+ n\min(\log n, Δ))\right)$ time, where $τ$ is the tree-sampling time, $\bar{d}$ is the average degree, and $Δ$ is the graph diameter. This complexity is near-linear in practice. Moreover, existing methods largely target static graphs and lack efficient mechanisms for dynamic updates. To address this, we propose two sample maintenance strategies that partially update samples while preserving accuracy on dynamic graphs. Extensive experiments on 10 large real-world datasets demonstrate that our method consistently outperforms state-of-the-art approaches in both efficiency and accuracy on static and dynamic graphs.

cs.DS↗

Virtual Width Networks

We introduce Virtual Width Networks (VWN), a framework that delivers the benefits of wider representations without incurring the quadratic cost of increasing the hidden size. VWN decouples representational width from backbone width, expanding the embedding space while keeping backbone compute nearly constant. In our large-scale experiment, an 8-times expansion accelerates optimization by over 2 times for next-token and 3 times for next-2-token prediction. The advantage amplifies over training as both the loss gap grows and the convergence-speedup ratio increases, showing that VWN is not only token-efficient but also increasingly effective with scale. Moreover, we identify an approximately log-linear scaling relation between virtual width and loss reduction, offering an initial empirical basis and motivation for exploring virtual-width scaling as a new dimension of large-model efficiency.

cs.LG↗

Violation of local realism with spatially multimode parametric down-conversion pumped by spatially incoherent light

We experimentally demonstrate a violation of local realism with highly spatially multimode polarization-entangled two-photon states produced by spontaneous parametric down-conversion (SPDC) pumped by a spatially incoherent light source-a light-emitting diode (LED). While existing studies have observed such a violation only by post-selecting the LED-pumped SPDC photons into a single spatial detection mode, we achieve a Clauser-Horne-Shimony-Holt inequality violation of $S = 2.532 \pm 0.069 > 2$ using a spatially multimode detection setup that collects nearly 4,080 SPDC spatial modes. These results indicate that coherent pump sources, such as lasers, are not required for SPDC-based entanglement generation. Our work could enable novel and practical sources of entangled photons for quantum technologies such as device-independent quantum key distribution and quantum-enhanced sensing.

quant-ph↗

Global Distribution of the Key Species on the Surface of Europa

The icy surface of Europa is continuously bombarded by ions and electrons from Jupiter's magnetosphere. The bombardment of the particles dissociates water molecules on the surface of Europa and introduces impurities to the icy surface. Such processes lead to the generation of the nonwater species on the surface of Europa. These chemical species are closely related to the chemistry of the icy crust and the subsurface ocean, as well as Europa's habitability. However, our knowledge of the global distribution of these species is limited due to the sparse satellite and telescope observations on Europa. In this study, we combine a Europa plasma model and a chemical-transport model to simulate the global distribution of the key nonwater species on the surface of Europa. The initial results from our model agree well with the existing observations on the distributions of H2SO4 and SO2 but they show a significant discrepancy with the observed distribution of H2O2. Sensitivity tests on the reaction rate coefficients indicate that the simulated global distribution of all three species fit the observations well if the reaction rate coefficients in the ice are reduced by one order of magnitude. This finding provides a useful constraint on the rate coefficient of the chemical reactions in the ice. Furthermore, our model predicts that the O2 on the surface ice of Europa is concentrated on the leading hemisphere. The simulated global distribution of the key species on Europa may provide useful guidance for future missions to Europa, such as Europa Clipper and JUICE.

astro-ph.EP↗

A multi-modal vision-language model for generalizable annotation-free pathology localization

Existing deep learning models for defining pathology from clinical imaging data rely on expert annotations and lack generalization capabilities in open clinical environments. Here, we present a generalizable vision-language model for Annotation-Free pathology Localization (AFLoc). The core strength of AFLoc is extensive multi-level semantic structure-based contrastive learning, which comprehensively aligns multi-granularity medical concepts with abundant image features to adapt to the diverse expressions of pathologies without the reliance on expert image annotations. We conduct primary experiments on a dataset of 220K pairs of image-report chest X-ray images and perform validation across eight external datasets encompassing 34 types of chest pathologies. The results demonstrate that AFLoc outperforms state-of-the-art methods in both annotation-free localization and classification tasks. Additionally, we assess the generalizability of AFLoc on other modalities, including histopathology and retinal fundus images. We show that AFLoc exhibits robust generalization capabilities, even surpassing human benchmarks in localizing five different types of pathological images. These results highlight the potential of AFLoc in reducing annotation requirements and its applicability in complex clinical environments.

cs.CV↗

All-optical turbulence mitigation for free-space quantum key distribution using stimulated parametric down-conversion

In this work, we propose and demonstrate a turbulence-resilient scheme for free-space quantum communication. By leveraging the phase conjugation property of stimulated parametric down-conversion, our scheme enables all-optical dynamic correction of spatial-mode distortion induced by atmospheric turbulence, thereby enhancing the secure key rate in high-dimensional quantum key distribution. We develop a theoretical model that provides detailed guidelines for selecting the optimal basis and spatial properties needed to maximize the efficiency of the proposed scheme. Both numerical simulations and experimental results show that, even under strong turbulence, our scheme can reduce the quantum error rates well below the security threshold. These results highlight the potential of nonlinear optical approaches as powerful tools for robust quantum communication in realistic free-space environments. Our work could have important implications for the practical implementation of secure quantum channels over long free-space distances.

quant-ph↗

ReVeal: Self-Evolving Code Agents via Reliable Self-Verification

Reinforcement learning with verifiable rewards (RLVR) has advanced the reasoning capabilities of large language models. However, existing methods rely solely on outcome rewards, without explicitly optimizing verification or leveraging reliable signals from realistic environments, leading to unreliable self-verification and limited test-time scaling. To address this, we widen the verification-generation asymmetry by explicitly optimizing self-verification, making it a reliable driver of deeper test-time scaling. We introduce ReVeal, a multi-turn reinforcement learning framework that evolves code generation through self-verification and tool-based evaluation. ReVeal structures long-horizon reasoning as iterative generation-verification turns and incorporates TAPO for turn-level credit assignment, fostering the co-evolution of code and test generation. At inference, this strengthened self-verification enables the model to use self-constructed tests and tool feedback to continuously evolve code for 20+ turns on LiveCodeBench despite training on only three. It also significantly improves Pass@k, indicating stronger exploration that expands the reasoning boundaries of the base model. These findings highlight the promise of ReVeal as a scalable paradigm for RL training and test-time scaling, paving the way for more robust and autonomous AI agents.

cs.SE↗

Nonuniform Water Distribution in Jupiter's Mid Latitudes: Influence of Precipitation and Planetary Rotation

Knowing the composition of Jupiter's atmosphere is crucial for constraining Jupiter's bulk metallicity and formation history. Yet, constraining Jupiter's atmospheric water abundance is challenging due to its potential non-uniform distribution. Here, we explicitly resolve the water hydrological cycle in Jupiter's mid-latitudes using high-resolution simulations. Falling precipitation leads to a significant large-scale depletion of water vapor beneath the lifting condensation level. A non-uniform water vapor distribution emerges in the mid-latitude simulation with a changing Coriolis parameter across latitudes and spatially uniform cooling and heating. Water abundance at the 7-bar level varies by up to a factor of ten across latitudes, from sub-solar to super-solar values. We propose that nonlinear large-scale eddies and waves tend to drift air parcels across latitudes along constant potential vorticity (PV) surfaces, thereby sustaining latitudinal dependencies in water vapor and the interplay between water distribution and large-scale dynamics. Therefore, water distribution is influenced by the vertical structure of density stratification and changing Coriolis parameter across Jupiter's mid-latitudes, as quantified by PV. Additionally, the water hydrological cycle amplifies the specific energy of air parcels through the latent heat effect, thereby slowing down vertical mixing with a latent heat flux. The horizontal gradient of water is expected to be more pronounced with a super-solar water abundance. We suggest that similar interplays between precipitating condensates, planetary rotation, and distribution of condensable species generally exist in the weather layer of fast-rotating giant planets. The ongoing Juno mission and future Uranus mission may further reveal the non-uniform distribution of condensed species and their interplay with large-scale dynamics.

astro-ph.EP↗

Introduction to the Chinese Space Station Survey Telescope (CSST)

The Chinese Space Station Survey Telescope (CSST) is an upcoming Stage-IV sky survey telescope, distinguished by its large field of view (FoV), high image quality, and multi-band observation capabilities. It can simultaneously conduct precise measurements of the Universe by performing multi-color photometric imaging and slitless spectroscopic surveys. The CSST is equipped with five scientific instruments, i.e. Multi-band Imaging and Slitless Spectroscopy Survey Camera (SC), Multi-Channel Imager (MCI), Integral Field Spectrograph (IFS), Cool Planet Imaging Coronagraph (CPI-C), and THz Spectrometer (TS). Using these instruments, CSST is expected to make significant contributions and discoveries across various astronomical fields, including cosmology, galaxies and active galactic nuclei (AGN), the Milky Way and nearby galaxies, stars, exoplanets, Solar System objects, astrometry, and transients and variable sources. This review aims to provide a comprehensive overview of the CSST instruments, observational capabilities, data products, and scientific potential.

astro-ph.IM↗

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts

Transformer models based on the Mixture of Experts (MoE) architecture have made significant progress in long-sequence modeling, but existing models still have shortcomings in computational efficiency and the ability to capture long-range dependencies, especially in terms of the dynamic adaptability of expert resource allocation. In this paper, we propose a Dynamic Adaptive Shared Expert and Grouped Multi-Head Attention Hybrid Model (DASG-MoE) to enhance long-sequence modeling capabilities by integrating three modules. First, we employ the Grouped Multi-Head Attention (GMHA) mechanism to effectively reduce the computational complexity of long sequences. By parallel processing through sequence grouping, local sliding window attention, and feature aggregation, we address long-range dependency issues and the model's lack of generalization for local information. Second, we design a Dual-Scale Shared Expert Structure (DSSE), where shallow experts use lightweight computations to quickly respond to low-dimensional features, while deep experts process high-dimensional complex semantics through pre-training transfer and post-training optimization, achieving a dynamic balance between efficiency and accuracy. Third, we propose a hierarchical Adaptive Dynamic Routing (ADR) mechanism that dynamically selects expert levels based on feature complexity and task requirements, and optimizes resource allocation through a local expert activation strategy. Experiments on multiple long-sequence benchmark datasets demonstrate that our DASG-MoE model outperforms state-of-the-art models.

cs.LG↗