Searcharxiv⌕ Search

arXiv subjects

Cheng Li

Publications and source records attributed to Cheng Li.

At least 91 records · Page 5Linked to original sources

ScoutAttention: Efficient KV Cache Offloading via Layer-Ahead CPU Pre-computation for LLM Inference

Large language models encounter critical GPU memory capacity constraints during long-context inference, where KV cache memory consumption severely limits decode batch sizes. While existing research has explored offloading KV cache to DRAM, these approaches either demand frequent GPU-CPU data transfers or impose extensive CPU computation requirements, resulting in poor GPU utilization as the system waits for I/O operations or CPU processing to complete. We propose ScoutAttention, a novel KV cache offloading framework that accelerates LLM inference through collaborative GPU-CPU attention computation. To prevent CPU computation from bottlenecking the system, ScoutAttention introduces GPU-CPU collaborative block-wise sparse attention that significantly reduces CPU load. Unlike conventional parallel computing approaches, our framework features a novel layer-ahead CPU pre-computation algorithm, enabling the CPU to initiate attention computation one layer in advance, complemented by asynchronous periodic recall mechanisms to maintain minimal CPU compute load. Experimental results demonstrate that ScoutAttention maintains accuracy within 2.4% of baseline while achieving 2.1x speedup compared to existing offloading methods.

cs.LG↗

LiteCache: A Query Similarity-Driven, GPU-Centric KVCache Subsystem for Efficient LLM Inference

During LLM inference, KVCache memory usage grows linearly with sequence length and batch size and often exceeds GPU capacity. Recent proposals offload KV states to host memory and reduce transfers using top-k attention. But their CPU-centric management of the on-GPU cache and CPU-GPU data movement incurs high overhead and fragments the bulk GPU execution that CUDA Graph relies on. To close this gap, we observe that adjacent queries within the same attention head exhibit strong directional similarity and retrieve highly overlapping top-k KV states. This insight enables a simple head granularity cache algorithm, QSAC, in which each head reuses its previously cached KV states whenever the current query is sufficiently similar to the prior one. QSAC further simplifies cache management primitives and cuts CPU involvement almost entirely. We develop LiteCache, a KVCache subsystem that incorporates QSAC. LiteCache introduces a GPU-centric synchronization controller and speculative sparse prefetching, enabling fully overlapped data movement and computation. These mechanisms produce a stable and predictable execution pattern that remains compatible with the bulk execution mode required by CUDA Graphs. Evaluation on two widely-used LLMs indicates that LiteCache achieves comparable accuracy to baselines, while sharply minimizing CPU overhead, fully utilizing PCIe bandwidth, thus improving decoding throughput by 10.7-224.2% on both H100 and A40 GPUs and easily supporting sequence lengths beyond 1M. We opensource LiteCache at https://anonymous.4open.science/r/LiteCache-888D.

cs.LG↗

When Brain Foundation Model Meets Cauchy-Schwarz Divergence: A New Framework for Cross-Subject Motor Imagery Decoding

Decoding motor imagery (MI) electroencephalogram (EEG) signals, a key non-invasive brain-computer interface (BCI) paradigm for controlling external systems, has been significantly advanced by deep learning. However, cross-subject MI-EEG decoding remains challenging due to substantial inter-subject variability and limited labeled target data, which necessitate costly calibration for new users. Many existing multi-source domain adaptation (MSDA) methods indiscriminately incorporate all available source domains, disregarding the large inter-subject differences in EEG signals, which leads to negative transfer and excessive computational costs. Moreover, while many approaches focus on feature distribution alignment, they often neglect the explicit dependence between features and decision-level outputs, limiting their ability to preserve discriminative structures. To address these gaps, we propose a novel MSDA framework that leverages a pretrained large Brain Foundation Model (BFM) for dynamic and informed source subject selection, ensuring only relevant sources contribute to adaptation. Furthermore, we employ Cauchy-Schwarz (CS) and Conditional CS (CCS) divergences to jointly perform feature-level and decision-level alignment, enhancing domain invariance while maintaining class discriminability. Extensive evaluations on two benchmark MI-EEG datasets demonstrate that our framework achieves average accuracies of 86.17% and 78.41%, outperforming a broad range of state-of-the-art baselines. Additional experiments with a large source pool validate the scalability and efficiency of BFM-guided selection.

cs.LG↗

Interaction-induced HI gas concentration with centrally-enhanced star formation in ALFALFA-SDSS galaxies

We present a statistical analysis for the interaction-induced central concentration of HI gas distributions and its connection with interaction-induced central star formation enhancement, using a large sample of $\sim 10^4$ galaxies from the ALFALFA and SDSS surveys. By adopting the HI profile parameter $K$, an indicator of gas concentration inferred from the integrated 21 cm emission line, we find that galaxies with more centrally concentrated HI (higher $K$ values) or enhanced specific star foramtion rate (sSFR) exhibit significantly stronger clustering and higher probability of hosting a nearby neighbor on scales below $100h^{-1}\mathrm{kpc}$, which is more pronounced in low-mass galaxies. Furthermore, by utilizing the enhancement functions for a sample of galaxy pairs, we directly trace the evolution of HI concentration and sSFR enhancement as a function of projected separation. Our findings indicate that tidal interactions drive a statistical synchrony between the central concentration of atomic gas and the enhancement of central star formation. Gas concentration appears to be a necessary condition for central star formation enhancement in interacting systems at all but the smallest separations. Compared to satellite galaxies, central galaxies exhibit stronger enhancement of gas fraction, gas concentration and sSFR, suggesting the role of environmental regulation.

astro-ph.GA↗

A Tug-of-War Between Baroclinic Eddies and Convection: Implications for Icy Moon Oceans

In many geophysical and planetary environments, such as Earth's ocean and atmosphere as well as subsurface oceans of icy satellites, convection driven by bottom geothermal heating usually coexists with baroclinic eddies driven by lateral buoyancy/temperature gradients. These processes compete against each other, with convection destabilizing the stratification and baroclinic eddies re-stabilizing it, thereby controlling whether the bottom heat flux is significantly redistributed as it is transmitted to the upper surface. Using scaling analysis and numerical simulations, we show that a stratified layer persists near the upper surface up to ${\rm Ra}_{v}\sim {\rm Ra}_h^{5/2}$, where ${\rm Ra}_h\equiv Δb_0/(L_zf^2)$ measures the imposed upper-surface buoyancy contrast $Δb_0$ and ${\rm Ra}_v\equiv B_0/(L_z^2f^3)$ measures the strength of the bottom buoyancy flux $B_0$, $L_z$ is the domain depth and $f$ is the Coriolis parameter. For ${\rm Ra}_v<{\rm Ra}_h^{5/2}$, baroclinic eddies dominate over convection, maintain the upper stratified layer, and completely deflect the bottom buoyancy/heat input into meridional transport. In contrast, when ${\rm Ra}_v>{\rm Ra}_h^{5/2}$, convective plumes penetrate the stratification and transport buoyancy/heat vertically with negligible deflection. Building on these results, we further propose a scaling law for the meridional buoyancy/heat transport in this system. Applications to icy satellites are discussed.

physics.ao-ph↗

Let It Flow: Agentic Crafting on Rock and Roll, Building the ROME Model within an Open Agentic Learning Ecosystem

Agentic crafting requires LLMs to operate in real-world environments over multiple turns by taking actions, observing outcomes, and iteratively refining artifacts. Despite its importance, the open-source community lacks a principled, end-to-end ecosystem to streamline agent development. We introduce the Agentic Learning Ecosystem (ALE), a foundational infrastructure that optimizes the production pipeline for agentic model. ALE consists of three components: ROLL, a post-training framework for weight optimization; ROCK, a sandbox environment manager for trajectory generation; and iFlow CLI, an agent framework for efficient context engineering. We release ROME, an open-source agent grounded by ALE and trained on over one million trajectories. Our approach includes data composition protocols for synthesizing complex behaviors and a novel policy optimization algorithm, Interaction-Perceptive Agentic Policy Optimization (IPA), which assigns credit over semantic interaction chunks rather than individual tokens to improve long-horizon training stability. Empirically, we evaluate ROME within a structured setting and introduce Terminal Bench Pro, a benchmark with improved scale and contamination control. ROME demonstrates strong performance across benchmarks like SWE-bench Verified and Terminal Bench, proving the effectiveness of ALE.

cs.AI↗

Wukong-Omni: Design, Modeling and Control of a Multi-mode Robot for Air, Land, and Underwater Exploration with All-in-One Propulsion Unit

In flood disaster rescue scenarios, partially submerged buildings prevent aerial robots from accessing lower levels, limiting mission effectiveness. To address this challenge, this paper presents Wukong-Omni, a novel multimode robot capable of operating across land, air, and underwater using a unified propulsion system. The system is enabled by an innovative mechanical design that allows motor reuse and improves thrust generation. Efficiency and peak thrust are enhanced through simulation and tank-based optimization. Experimental results show a 100 percent improvement in propulsion efficiency and a 150 percent increase in maximum thrust compared with direct installation methods. Dynamic models for the three operating domains are developed, and a unified cross-domain control framework is proposed. Comprehensive experiments validate stable locomotion and smooth transition across domains. Outdoor experiments further demonstrate robustness and adaptability in real-world environments.

cs.RO↗

Lagom: Unleashing the Power of Communication and Computation Overlapping for Distributed LLM Training

Overlapping communication with computation is crucial for distributed large-model training, yet optimizing it - especially when computation becomes the bottleneck-remains challenging. We present Lagom, a system that co-tunes communication parameters to balance resource usage between computation and communication. By introducing a unified cost model and a priority-based search algorithm, Lagom reduces optimization complexity from exponential to linear. Evaluations on high- and low-bandwidth GPU clusters show that Lagom achieves 1.07-1.33x and 1.03-1.27x speedup over NCCL and AutoCCL across diverse models and parallelizations.

cs.DC↗

BayeSED-GALAXIES II. Bayesian full spectrum analysis of galaxies and application in the CSST wide-field slitless spectroscopy survey

The China Space Station Telescope (CSST) will conduct wide-field multiband photometric imaging and slitless spectroscopic surveys, advancing cosmology and galaxy evolution studies. Achieving CSST's cosmological goals requires precise redshifts ($σ_{\rm NMAD}\lesssim 0.002-0.005$) from low-resolution ($R\sim200$) and potentially blended slitless spectra. We present BayeSED3, extended for Bayesian full-spectrum analysis, including nebular emission modeling (via \textsc{Cloudy}) and a Bayesian treatment of the model scaling factor, improving reliability over optimization methods for low SNR spectra. Validated on realistic mock data generated with the CESS emulator (median SNR=1.65, including instrumental and self-blending effects), our method achieves excellent redshift precision with three-band (GU+GV+GI) spectroscopy: $σ_{\rm NMAD}=0.0008$ ($\sim$80% success) for star-forming and $σ_{\rm NMAD}=0.0015$ ($\sim$50% success) for quiescent galaxies. Stellar mass ($σ_{\rm NMAD}\approx0.015$ dex for SF, $\approx0.016$ dex for quiescent) and SFR ($σ_{\rm NMAD}\approx0.05$ dex for SF, especially at SNR>1) are reliably recovered. Self-blending increases scatter by $\gtrsim30%$, but combining spectroscopy with CSST's seven-band photometry significantly improves accuracy, especially for quiescent galaxies and data-limited cases. Single-band spectroscopy plus photometry yields reasonable redshifts: GU+photometry is limited, GI+photometry gives >60% (SF) and >40% (quiescent) success at $σ_{\rm NMAD}\lesssim0.002$, GV+photometry gives >35% (SF) and $\sim$40% (quiescent) at similar precision. The Bayesian framework offers a powerful method for accurate galaxy characterization, enhancing CSST's scientific outcomes despite the challenges of slitless spectroscopy.

astro-ph.GA↗

Kimi K2: Open Agentic Intelligence

We introduce Kimi K2, a Mixture-of-Experts (MoE) large language model with 32 billion activated parameters and 1 trillion total parameters. We propose the MuonClip optimizer, which improves upon Muon with a novel QK-clip technique to address training instability while enjoying the advanced token efficiency of Muon. Based on MuonClip, K2 was pre-trained on 15.5 trillion tokens with zero loss spike. During post-training, K2 undergoes a multi-stage post-training process, highlighted by a large-scale agentic data synthesis pipeline and a joint reinforcement learning (RL) stage, where the model improves its capabilities through interactions with real and synthetic environments. Kimi K2 achieves state-of-the-art performance among open-source non-thinking models, with strengths in agentic capabilities. Notably, K2 obtains 66.1 on Tau2-Bench, 76.5 on ACEBench (En), 65.8 on SWE-Bench Verified, and 47.3 on SWE-Bench Multilingual -- surpassing most open and closed-sourced baselines in non-thinking settings. It also exhibits strong capabilities in coding, mathematics, and reasoning tasks, with a score of 53.7 on LiveCodeBench v6, 49.5 on AIME 2025, 75.1 on GPQA-Diamond, and 27.1 on OJBench, all without extended thinking. These results position Kimi K2 as one of the most capable open-source large language models to date, particularly in software engineering and agentic tasks. We release our base and post-trained model checkpoints to facilitate future research and applications of agentic intelligence.

cs.LG↗

Neuro-Symbolic Verification on Instruction Following of LLMs

A fundamental problem of applying Large Language Models (LLMs) to important applications is that LLMs do not always follow instructions, and violations are often hard to observe or check. In LLM-based agentic workflows, such violations can propagate and amplify along reasoning chains, causing task failures and system incidents. This paper presents NSVIF, a neuro-symbolic framework for verifying whether an LLM's output follows the instructions used to prompt the LLM. NSVIF is a universal, general-purpose verifier; it makes no assumption about the instruction or the LLM. NSVIF formulates instruction-following verification as a constraint-satisfaction problem by modeling user instructions as constraints. NSVIF models both logical and semantic constraints; constraint solving is done by a unified solver that orchestrates logical reasoning and semantic analysis. To evaluate NSVIF, we develop VIFBENCH, a new benchmark for instruction-following verifiers with fine-grained data labels. Experiments show that NSVIF significantly outperforms LLM-based approaches and provides interpretable feedback. We also show that feedback from NSVIF helps improve LLMs' instruction-following capability without post-training.

cs.AI↗

Pathways of Thoughts: Multi-Directional Thinking for Long-form Personalized Question Answering

Personalization is well studied in search and recommendation, but personalized question answering remains underexplored due to challenges in inferring preferences from long, noisy, implicit contexts and generating responses that are both accurate and aligned with user expectations. To address this, we propose Pathways of Thoughts (PoT), an inference-stage method that applies to any large language model (LLM) without task-specific fine-tuning. PoT models the thinking as an iterative decision process, where the model dynamically selects among cognitive operations such as reasoning, revision, personalization, and clarification. This enables exploration of multiple reasoning trajectories, producing diverse candidate responses that capture different perspectives. PoT then aggregates and reweights these candidates according to inferred user preferences, yielding a final personalized response that benefits from the complementary strengths of diverse reasoning paths. Experiments on the LaMP-QA benchmark show that PoT consistently outperforms competitive baselines, achieving up to a 10.8\% relative improvement. Human evaluation further validates these improvements, with annotators preferring PoT in 66\% of cases compared to the best-performing baseline and reporting ties in 15\% of cases.

cs.CL↗

Exploring Lorentz Invariance Violation from Ultra-high-energy Gamma Rays Observed by LHAASO

Recently the LHAASO Collaboration published the detection of 12 ultra-high-energy gamma-ray sources above 100 TeV, with the highest energy photon reaching 1.4 PeV. The first detection of PeV gamma rays from astrophysical sources may provide a very sensitive probe of the effect of the Lorentz invariance violation (LIV), which results in decay of high-energy gamma rays in the superluminal scenario and hence a sharp cutoff of the energy spectrum. Two highest energy sources are studied in this work. No signature of the existence of LIV is found in their energy spectra, and the lower limits on the LIV energy scale are derived. Our results show that the first-order LIV energy scale should be higher than about 10^5 times the Planck scale M_{pl} and that the second-order LIV scale is >10^{-3}M_{pl}. Both limits improve by at least one order of magnitude the previous results.

astro-ph.HE↗

MVT: Mask-Grounded Vision-Language Models for Taxonomy-Aligned Land-Cover Tagging

Land-cover understanding in remote sensing increasingly demands class-agnostic systems that generalize across datasets while remaining spatially precise and interpretable. We study a geometry-first discovery-and-interpretation setting under domain shift, where candidate regions are delineated class-agnostically and supervision avoids lexical class names via anonymized identifiers. Complementary to open-set recognition and open-world learning, we focus on coupling class-agnostic mask evidence with taxonomy-grounded scene interpretation, rather than unknown rejection or continual class expansion. We propose MVT, a three-stage framework that (i) extracts boundary-faithful region masks using SAM2 with domain adaptation, (ii) performs mask-grounded semantic tagging and scene description generation via dual-step LoRA fine-tuning of multimodal LLMs, and (iii) evaluates outputs with LLM-as-judge scoring calibrated by stratified expert ratings. On cross-dataset segmentation transfer (train on OpenEarthMap, evaluate on LoveDA), domain-adapted SAM2 improves mask quality; meanwhile, dual-step MLLM fine-tuning yields more accurate taxonomy-aligned tags and more informative mask-grounded scene descriptions.

cs.CV↗

Constraints on heavy decaying dark matter from 570 days of LHAASO observations

The Kilometer Square Array~(KM2A) of the Large High Altitude Air Shower Observatory (LHAASO) aims at surveying the northern gamma-ray sky at energies above 10 TeV with unprecedented sensitivity. Gamma-ray observations have long been one of the most powerful tools for dark matter searches, as e.g., high-energy gamma-rays could be produced by the decays of heavy dark matter particles. In this letter, we present the first dark matter analysis with LHAASO-KM2A, using the first 340~days of data from 1/2-KM2A and 230~days of data from 3/4-KM2A. Several regions of interest are used to search for a signal and account for the residual cosmic-ray background after gamma/hadron separation. We find no excess of dark matter signals, and thus place some of the strongest gamma-ray constraints on the lifetime of heavy dark matter particles with mass between 10^5 and 10^9~GeV. Our results with LHAASO are robust, and have important implications for dark matter interpretations of the diffuse astrophysical high-energy neutrino emission.

astro-ph.HE↗

Discovery of a new $γ$-ray source LHAASO J0341+5258 with emission up to 200TeV

We report the discovery of a new unidentified extended $γ$-ray source in the Galactic plane named LHAASO J0341+5258 with a pre-trial significance of 8.2 standard deviations above 25 TeV. The best fit position is R.A.$=55.34^{\circ}\pm0.11^{\circ}$ and Dec$=52.97^{\circ}\pm0.07^{\circ}$. The angular size of LHAASO J0341+5258 is $0.29^\circ \pm 0.06^\circ_{stat} \pm0.02^\circ_{sys}$. The flux above 25 TeV is about $20\%$ of the flux of Crab Nebula. Although a power-law fit of the spectrum from 10 TeV to 200 TeV with the photon index $α=2.98 \pm 0.19_{stat} \pm 0.02_{sys}$ is not excluded, the LHAASO data together with the flux upper limit at 10 GeV set by the Fermi LAT observation, indicate a noticeable steepening of an initially hard power-law spectrum %($α\leq 1.75$) spectrum with a cutoff at $\approx 50$ TeV. We briefly discuss the origin of UHE gamma-rays. The lack of an energetic pulsar and a young SNR inside or in the vicinity of LHAASO J0341+5258 challenge, but do not exclude both the leptonic and hadronic scenarios of gamma-ray production.

astro-ph.HE↗

Discovery of the Ultra-high energy gamma-ray source LHAASO J2108+5157

We report the discovery of a UHE gamma-ray source, LHAASO J2108+5157, by analyzing the LHAASO-KM2A data of 308.33 live days. Significant excess of gamma-ray induced showers is observed in both energy bands of 25-100 TeV and $\gt$100 TeV with 9.5 sigma and 8.5 sigma, respectively. This source is not significantly favored as an extensive source with the angular extension smaller than the point-spread function of KM2A. The measured energy spectrum from 20 to 200 TeV can be approximately described by a power-law function with an index of -2.83$\pm$ 0.18stat. A harder spectrum is demanded at lower energies considering the flux upper limit set by Fermi-LAT observations. The position of the gamma-ray emission is correlated with a giant molecular cloud, which favors a hadronic origin. No obvious counterparts have been found, deeper multiwavelength observations will help to shed new light on this intriguing UHE source.

astro-ph.HE↗

Fast control of the transverse structure of a light beam using acousto-optic modulators

Fast, reprogrammable control over the transverse structure of light beams plays an essential role in applications such as structured illumination microscopy, optical trapping, and quantum information processing. Existing technologies, such as liquid crystal on silicon spatial light modulators (LCoS-SLMs) and digital micromirror devices (DMDs), suffer from limited refresh rates, low damage thresholds, and high insertion loss. Acousto-optic modulators (AOMs) can resolve the above issues, as they typically handle higher laser power and offer lower insertion loss. By effectively mapping the temporal radio-frequency (RF) waveforms onto the spatial diffraction patterns of the optical field, individual AOMs have been shown to generate one-dimensional (1D) spatial modes at a pixel refresh rate of nearly 20 MHz. We extend this concept to enable fast modulation in a two-dimensional (2D) space using a double-AOM scheme. We demonstrate the generation of 2D Hermite-Gaussian (HG_nm) modes with an average fidelity of 81%, while the highest-order mode generated, HG_53, retains a fidelity of 56%.

physics.optics↗