Searcharxiv⌕ Search

arXiv subjects

Rui Li

Publications and source records attributed to Rui Li.

At least 109 records · Page 6Linked to original sources

PestVL-Net: Enabling Multimodal Pest Learning via Fine-grained Vision-Language Interaction

Effective pest recognition and management are crucial for sustainable agricultural development. However, collecting pest data in real scenarios is often challenging. Compared to other domains, pests exhibit a wide variety of species with complex and diverse morphological characteristics. Existing techniques struggle to effectively model the key visual and high-level semantic features of pests in a fine-grained manner. These limitations hinder the practical application of such methods in real agricultural scenarios. To address these critical challenges, we present a synergistic approach that integrates PestVL-Net, a novel vision-language framework, with two multi-species pest datasets to facilitate fine-grained pest learning. The visual pathway of PestVL-Net utilizes the Recurrent Weighted Key Value (RWKV) architecture, incorporating a saliency-guided adaptive window partitioning scheme to effectively model the fine-grained visual characteristics of pests. Concurrently, the linguistic component generates precise pest semantic descriptions by leveraging Multimodal Large Language Models (MLLMs) priors, critically informed by agricultural expert knowledge and structured via multimodal Chain-of-Thought (CoT) reasoning. The deep fusion of these complementary visual and textual representations enables fine-grained multimodal pest learning. Extensive experimental evaluations on multiple pest datasets validate the superior performance of PestVL-Net, highlighting its potential for effective real-world pest management.

cs.CV↗

Antiferromagnetic Dimers in the Parent Phase of a Correlated Kagome Superconductor

Kagome metals are prone to charge-density wave (CDW), magnetic, and superconducting phases, with their flat electronic band conducive for correlated physics. In contrast to the weakly correlated $A$V$_3$Sb$_5$ ($A$ = K, Rb, Cs) kagome metals with a $2\times2$ CDW, CsCr$_3$Sb$_5$ is a correlated metal with a flat band close to the Fermi level, and exhibits a $4\times1$ CDW intertwined with magnetic order. Under pressure, the intertwined orders are suppressed and give way to a dome of superconductivity that emerges from a non-Fermi liquid normal state. Here, we solve the crystal structure of the $4\times 1$ CDW state in CsCr$_3$Sb$_5$, and show it consists of Cr dimers separated by Cr chains. First-principles calculations show the dominant exchange interaction is antiferromagnetic within the dimers, while the intra-chain and dimer-chain couplings are much weaker. The CDW transition of CsCr$_3$Sb$_5$ is found to be more strongly first-order than those in $A$V$_3$Sb$_5$, without significant soft phonons or diffuse scattering above the CDW transition temperature. These findings suggest that fluctuating antiferromagnetic dimers may play a major role in the electron pairing of superconducting CsCr$_3$Sb$_5$.

cond-mat.str-el↗

Stochasticity in Tokenisation Improves Robustness

The widespread adoption of large language models (LLMs) has increased concerns about their robustness. Vulnerabilities in perturbations of tokenisation of the input indicate that models trained with a deterministic canonical tokenisation can be brittle to adversarial attacks. Recent studies suggest that stochastic tokenisation can deliver internal representations that are less sensitive to perturbations. In this paper, we analyse how stochastic tokenisations affect robustness to adversarial attacks and random perturbations. We systematically study this over a range of learning regimes (pre-training, supervised fine-tuning, and in-context learning), data sets, and model architectures. We show that pre-training and fine-tuning with uniformly sampled stochastic tokenisations improve robustness to random and adversarial perturbations. Evaluating on uniformly sampled non-canonical tokenisations reduces the accuracy of a canonically trained Llama-1b model by 29.8%. We find that training with stochastic tokenisation preserves accuracy without increasing inference cost.

cs.CL↗

A Lightweight, Transferable, and Self-Adaptive Framework for Intelligent DC Arc-Fault Detection in Photovoltaic Systems

Arc-fault circuit interrupters (AFCIs) are essential for mitigating fire hazards in residential photovoltaic (PV) systems, yet achieving reliable DC arc-fault detection under real-world conditions remains challenging. Spectral interference from inverter switching, hardware heterogeneity, operating-condition drift, and environmental noise collectively compromise conventional AFCI solutions. This paper proposes a lightweight, transferable, and self-adaptive learning-driven framework (LD-framework) for intelligent DC arc-fault detection. At the device level, LD-Spec learns compact spectral representations enabling efficient on-device inference and near-perfect arc discrimination. Across heterogeneous inverter platforms, LD-Align performs cross-hardware representation alignment to ensure robust detection despite hardware-induced distribution shifts. To address long-term evolution, LD-Adapt introduces a cloud-edge collaborative self-adaptive updating mechanism that detects unseen operating regimes and performs controlled model evolution. Extensive experiments involving over 53,000 labeled samples demonstrate near-perfect detection, achieving 0.9999 accuracy and 0.9996 F1-score. Across diverse nuisance-trip-prone conditions, including inverter start-up, grid transitions, load switching, and harmonic disturbances, the method achieves a 0% false-trip rate. Cross-hardware transfer shows reliable adaptation using only 0.5%-1% labeled target data while preserving source performance. Field adaptation experiments demonstrate recovery of detection precision from 21% to 95% under previously unseen conditions. These results indicate that the LD-framework enables a scalable, deployment-oriented AFCI solution maintaining highly reliable detection across heterogeneous devices and long-term operation.

eess.SP↗

PoTable: Towards Systematic Thinking via Plan-then-Execute Stage Reasoning on Tables

In recent years, table reasoning has garnered substantial research interest, particularly regarding its integration with Large Language Models (LLMs), which have revolutionized natural language applications. Existing LLM-based studies typically achieve step-by-step thinking for table reasoning guided by task semantics. While these approaches emphasize autonomous exploration and enhance fine-grained table understanding, they often overlook systematic thinking in the reasoning process. This oversight can lead to omitted steps, disorganized logic and misleading results, especially in complex scenarios. In this paper, we propose PoTable, a novel stage-oriented plan-then-execute approach that incorporates systematic thinking into table reasoning. Specifically, PoTable involves several distinct analytical stages with clear objectives to provide adequate guidance. To accomplish stage-specific goals, PoTable employs a plan-then-execute mechanism: it first plans the operation chain based on the stage objective, and then executes operations sequentially through code generation, real-time running and feedback processing. Consequently, PoTable produces reliable table reasoning results with highly accurate, step-wise commented and completely executable programs. It mirrors the workflow of a professional data analyst, offering advantages in both accuracy and explainability. Finally, we conduct extensive experiments on four datasets from the WikiTQ and TabFact benchmarks, where the results demonstrate the effectiveness, efficiency and explainability of PoTable. Our code is available at: https://github.com/Double680/PoTable.

cs.IR↗

SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators

The proliferation of 100B+ parameter Large Language Models (LLMs) with 100k+ context length support have resulted in increasing demands for on-chip memory to support large KV caches. Techniques such as StreamingLLM and SnapKV demonstrate how to control KV cache size while maintaining model accuracy. Yet, these techniques are not commonly used within industrial deployments using frameworks like vLLM or SGLang. The reason is twofold: on one hand, the static graphs and continuous batching methodology employed by these frameworks make it difficult to admit modifications to the standard multi-head attention algorithm, while on the other hand, the accuracy implications of such techniques on modern instruction-following and reasoning models are not well understood, obfuscating the need for implementing these techniques. In this paper, we explore these accuracy implications on Llama-3.1-8B-Instruct and DeepSeek-R1, and develop SnapStream, a KV cache compression method that can be deployed at scale. We demonstrate the efficacy of SnapStream in a 16-way tensor-parallel deployment of DeepSeek-671B on SambaNova SN40L accelerators running at 128k context length and up to 1832 tokens per second in a real production setting. SnapStream enables $4\times$ improved on-chip memory usage and introduces minimal accuracy degradation on LongBench-v2, AIME24 and LiveCodeBench. To the best of our knowledge, this is the first implementation of sparse KV attention techniques deployed in a production inference system with static graphs and continuous batching.

cs.AI↗

Select-then-Solve: Paradigm Routing as Inference-Time Optimization for LLM Agents

When an LLM-based agent improves on a task, is the gain from the model itself or from the reasoning paradigm wrapped around it? We study this question by comparing six inference-time paradigms, namely Direct, CoT, ReAct, Plan-Execute, Reflection, and ReCode, across four frontier LLMs and ten benchmarks, yielding roughly 18,000 runs. We find that reasoning structure helps dramatically on some tasks but hurts on others: ReAct improves over Direct by 44pp on GAIA, while CoT degrades performance by 15pp on HumanEval. No single paradigm dominates, and oracle per-task selection beats the best fixed paradigm by 17.1pp on average. Motivated by this complementarity, we propose a select-then-solve approach: before answering each task, a lightweight embedding-based router selects the most suitable paradigm. Across four models, the router improves average accuracy from 47.6% to 53.1%, outperforming the best fixed paradigm at 50.3% by 2.8pp and recovering up to 37% of the oracle gap. In contrast, zero-shot self-routing only works for GPT-5 at 67.1% and fails for weaker models, all trailing the learned router. Our results argue that reasoning paradigm selection should be a per-task decision made by a learned router, not a fixed architectural choice.

cs.CL↗

The Python Simulations of Chemistry Framework: 10 years of an open-source quantum chemistry project

Over the past decade, the Python-based Simulations of Chemistry Framework (PySCF) has developed into a widely used open-source platform for electronic structure theory and quantum chemical method development. This article reviews the major advances since the previous overview in 2020, covering new modules and methodology, infrastructure changes, and performance benchmarks.

physics.chem-ph↗

CoEnv: Driving Embodied Multi-Agent Collaboration via Compositional Environment

Multi-agent embodied systems hold promise for complex collaborative manipulation, yet face critical challenges in spatial coordination, temporal reasoning, and shared workspace awareness. Inspired by human collaboration where cognitive planning occurs separately from physical execution, we introduce the concept of compositional environment -- a synergistic integration of real-world and simulation components that enables multiple robotic agents to perceive intentions and operate within a unified decision-making space. Building on this concept, we present CoEnv, a framework that leverages simulation for safe strategy exploration while ensuring reliable real-world deployment. CoEnv operates through three stages: real-to-sim scene reconstruction that digitizes physical workspaces, VLM-driven action synthesis supporting both real-time planning with high-level interfaces and iterative planning with code-based trajectory generation, and validated sim-to-real transfer with collision detection for safe deployment. Extensive experiments on challenging multi-arm manipulation benchmarks demonstrate CoEnv's effectiveness in achieving high task success rates and execution efficiency, establishing a new paradigm for multi-agent embodied AI.

cs.RO↗

A Self-Evolving Defect Detection Framework for Industrial Photovoltaic Systems

Reliable photovoltaic (PV) power generation requires timely detection of module defects that may reduce energy yield, accelerate degradation, and increase lifecycle operation and maintenance costs during field operation. Electroluminescence (EL) imaging has therefore been widely adopted for PV module inspection. However, automated defect detection in real operational environments remains challenging due to heterogeneous module geometries, low-resolution imaging conditions, subtle defect morphology, long-tailed defect distributions, and continual data shifts introduced by evolving inspection and labeling processes. These factors significantly limit the robustness and long-term maintainability of conventional deep-learning inspection pipelines. To address these challenges, this paper proposes SEPDD, a Self-Evolving Photovoltaic Defect Detection framework designed for evolving industrial PV inspection scenarios. SEPDD integrates automated model optimization with a continual self-evolving learning mechanism, enabling the inspection system to progressively adapt to distribution shifts and newly emerging defect patterns during long-term deployment. Experiments conducted on both a public PV defect benchmark and a private industrial EL dataset demonstrate the effectiveness of the proposed framework. Both datasets exhibit severe class imbalance and significant domain shift. SEPDD achieves a leading mAP50 of 91.4% on the public dataset and 49.5% on the private dataset. It surpasses the autonomous baseline by 14.8% and human experts by 4.7% on the public dataset, and by 4.9% and 2.5%, respectively, on the private dataset.

cs.AI↗

Joint$λ$: Orchestrating Serverless Workflows on Jointcloud FaaS Systems

Existing serverless workflow orchestration systems are predominantly designed for a single-cloud FaaS system, leading to vendor lock-in. This restricts performance optimization, cost reduction, and availability of applications. However, orchestrating serverless workflows on Jointcloud FaaS systems faces two main challenges: (1) additional overhead caused by centralized cross-cloud orchestration; and (2) a lack of reliable failover and fault-tolerant mechanisms for cross-cloud serverless workflows. To address these challenges, we propose Joint$λ$, a distributed runtime system designed to orchestrate serverless workflows on multiple FaaS systems without relying on a centralized orchestrator. Joint$λ$ introduces a compatibility layer, Backend-Shim, leveraging inter-cloud heterogeneity to optimize makespan and reduce costs with on-demand billing. By using function-side orchestration instead of centralized nodes, it enables independent function invocations and data transfers, reducing cross-cloud communication overhead. For high availability, it ensures exactly-once execution via datastores and failover mechanisms for serverless workflows on Jointcloud FaaS systems. We validate Joint$λ$ on two heterogeneous FaaS systems, AWS and Aliyun, with four workflows. Compared to the most advanced commercial orchestration services for single-cloud serverless workflows, Joint$λ$ reduces makespan by up to 3.3$\times$ while saving up to 65% in cost. Joint$λ$ is also up to 4.0$\times$ faster than state-of-the-art orchestrators for cross-cloud serverless workflows, while achieving competitive cost in representative scenarios and providing strong execution guarantees.

cs.DC↗

HOSL: Hybrid-Order Split Learning for Memory-Constrained Edge Training

Split learning (SL) enables collaborative training of large language models (LLMs) between resource-constrained edge devices and compute-rich servers by partitioning model computation across the network boundary. However, existing SL systems predominantly rely on first-order (FO) optimization, which requires clients to store intermediate quantities such as activations for backpropagation. This results in substantial memory overhead, largely negating benefits of model partitioning. In contrast, zeroth-order (ZO) optimization eliminates backpropagation and significantly reduces memory usage, but often suffers from slow convergence and degraded performance. In this work, we propose HOSL, a novel Hybrid-Order Split Learning framework that addresses this fundamental trade-off between memory efficiency and optimization effectiveness by strategically integrating ZO optimization on the client side with FO optimization on the server side. By employing memory-efficient ZO gradient estimation at the client, HOSL eliminates backpropagation and activation storage, reducing client memory consumption. Meanwhile, server-side FO optimization ensures fast convergence and competitive performance. Theoretically, we show that HOSL achieves an $\mathcal{O}(\sqrt{d_c/TQ})$ rate, which depends on client-side model dimension $d_c$ rather than the full model dimension $d$, demonstrating that convergence improves as more computation is offloaded to the server. Extensive experiments on OPT models (125M and 1.3B parameters) across 6 tasks demonstrate that HOSL reduces client GPU memory by up to 3.7$\times$ compared to the FO method while achieving accuracy within 0.20%-4.23% of this baseline. Furthermore, HOSL outperforms the ZO baseline by up to 15.55%, validating the effectiveness of our hybrid strategy for memory-efficient training on edge devices.

cs.LG↗

Developing and characterizing a new-generation regolith simulant "IGCAS-AST01" for the Tianwen-2 target asteroid (469219) Kamo'oalewa

China plans to return samples from the near-Earth asteroid (469219) Kamo'oalewa, which we previously identified as an LL-chondrite-compositional, highly space-weathered object with fine-grained regolith. In this study, we developed 10 mL of Kamo'oalewa regolith simulant, designated "IGCAS-AST01", by irradiating LL5/6 chondrite (Kheneg Ljou^ad) powder with a high-energy pulsed laser. We then analyzed the composition, grain size distribution, density, porosity, visible to near-infrared reflectance spectrum, thermal emission spectrum, thermal diffusivity, specific heat capacity, and microstructural features of both the fresh (unirradiated) powder and IGCAS-AST01. IGCAS-AST01 is composed of 57.8 vol.% olivine, 19.9 vol.% orthopyroxene, 5.6 vol.% diopside, 12.2 vol.% plagioclase, 2.6 vol.% troilite, and minor amounts of other phases. It has a mean size of 26.99 um, a median size of 23.19 um, a density of 700 kg m^-3, and a porosity of 79.1%. Additionally, IGCAS-AST01 exhibits a low reflectance of 0.1 at 0.55 um and an extremely steep spectral slope. In the temperature range of 253.15-473.15 K, its thermal diffusivity and specific heat capacity range from 3.6-4.7 x 10^-6 m^2 s^-1 and 718.43-890.20 J kg^-1 K^-1, respectively. Furthermore, thick amorphous rims and abundant nanophase metallic iron particles are observed in olivine and pyroxene grains of IGCAS-AST01. These results could support the Tianwen-2 mission's payload calibration, sampling operations, on-orbit scientific data interpretation, and future sample analysis.

astro-ph.EP↗

Expectation Error Bounds for Transfer Learning in Linear Regression and Linear Neural Networks

In transfer learning, the learner leverages auxiliary data to improve generalization on a main task. However, the precise theoretical understanding of when and how auxiliary data help remains incomplete. We provide new insights on this issue in two canonical linear settings: ordinary least squares regression and under-parameterized linear neural networks. For linear regression, we derive exact closed-form expressions for the expected generalization error with bias-variance decomposition, yielding necessary and sufficient conditions for auxiliary tasks to improve generalization on the main task. We also derive globally optimal task weights as outputs of solvable optimization programs, with consistency guarantees for empirical estimates. For linear neural networks with shared representations of width $q \leq K$, where $K$ is the number of auxiliary tasks, we derive a non-asymptotic expectation bound on the generalization error, yielding the first non-vacuous sufficient condition for beneficial auxiliary learning in this setting, as well as principled directions for task weight curation. We achieve this by proving a new column-wise low-rank perturbation bound for random matrices, which improves upon existing bounds by preserving fine-grained column structures. Our results are verified on synthetic data simulated with controlled parameters.

cs.LG↗

Implementation of the multigrid Gaussian-Plane-Wave algorithm with GPU acceleration in PySCF

We introduce a GPU-accelerated multigrid Gaussian-Plane-Wave density fitting (FFTDF) approach for efficient Fock builds and nuclear gradient evaluations within Kohn-Sham density functional theory, as implemented in the GPU4PySCF module of PySCF. Our CUDA kernels employ a grid-based parallelization strategy for contracting Gaussian basis function pairs and achieve up to 80% of the FP64 peak performance on NVIDIA GPUs, with no loss of efficiency for high angular momentum (up to f-shell) functions. Benchmark calculations on molecules and solids with up to 1536 atoms and 20480 basis functions show up to 25x speedup on an H100 GPU relative to the CPU implementation on a 28-core shared memory node. For a 256-water cluster, the ground-state energy and nuclear gradients can be computed in ~30 seconds on a single H100 GPU. This implementation serves as an open-source foundation for many applications, such as ab initio molecular dynamics and high-throughput calculations.

physics.chem-ph↗

Band inversion transition in HgTe nanowire grown along the [001] direction

The low-energy effective Hamiltonian of a cylindrical HgTe nanowire grown along the [001] crystallographic direction is constructed by using the perturbation theory. Both the anisotropic term and the bulk inversion asymmetry term of the Kane model are taken into account. Although the anisotropic term has converted the crossing between the $E_{1}$ and $H_{1}$ subbands into an anticrossing at $k_{z}R\!=\!0$, the gap-closing-and-reopening transition in the subband structure can still occur at finite wave vectors $k_{z}R\!\approx\!\pm0.24$ for critical nanowire radius $R\!\approx\!3.45$ nm. The bulk inversion asymmetry does not contribute to the low-energy effective Hamiltonian, i.e., there is no spin splitting in the $E_{1}$, $H_{1}$, and $H_{2}$ subbands for a [001] oriented cylindrical nanowire.

cond-mat.mes-hall↗

Robust Atom Interferometry with Double Bragg Diffraction

This thesis develops a general theoretical and numerical framework for achieving high-contrast atom interferometry based on double Bragg diffraction (DBD). While DBD offers intrinsic symmetry, reduced sensitivity to internal-state systematics, and suitability for microgravity experiments, its performance has long been limited by imperfect diffraction and contrast loss. This work overcomes these limitations by constructing an analytic Hamiltonian description of DBD -- including Doppler effects and polarization imperfections -- and by deriving reduced two- and five-level models via a truncated Magnus-expansion approach. These models clarify the origin of AC-Stark shifts, polarization-induced errors, and Doppler selectivity, and they provide accurate predictions for realistic input momentum distributions. Building on this theoretical foundation, the thesis introduces a tri-frequency laser scheme with dynamically tunable detuning and evaluates different detuning-control strategies using a five-level S-matrix formalism. Linear detuning sweeps and optimal-control pulses are shown to provide near-ideal beam-splitter and mirror performance, respectively, ensuring robust contrast across a wide range of experimental imperfections. Complementary full three-dimensional simulations using the GPU-accelerated Universal Atom Interferometer Simulator (UATIS) incorporate interacting Bose-Einstein condensates and realistic optical potentials, revealing transverse effects and polarization-induced distortions that extend the predictions of the one-dimensional non-interacting models. Taken together, this thesis establishes a coherent theoretical and numerical framework demonstrating that, with appropriate detuning control, double-Bragg atom interferometers can achieve the robustness required for precision inertial sensing and future space-based quantum tests of fundamental physics.

quant-ph↗

Advances on two spectral conjectures regarding booksize of graphs

The booksize $ \mathrm{bk}(G) $ of a graph $ G $, introduced by Erdős, refers to the maximum integer $ r $ for which $G$ contains the book $ B_r $ as a subgraph. This paper investigates two open problems in spectral graph theory related to the booksize of graphs. First, we prove that for any positive integer $r$ and any $ B_{r+1} $-free graph $ G $ with $ m \geq (9r)^2 $ edges, the spectral radius satisfies $ ρ(G) \leq \sqrt{m} $. Equality holds if and only if $ G $ is a complete bipartite graph. This result improves the lower bound on the booksize of Nosal graphs (i.e., graphs with $ ρ(G) > \sqrt{m} $) from the previously established $ \mathrm{bk}(G) > \frac{1}{144}\sqrt{m} $ to $ \mathrm{bk}(G) > \frac{1}{9}\sqrt{m} $, presenting a significant advancement in the booksize conjecture proposed Li, Liu, and Zhang. Second, we show that for any positive integer $r$ and any non-bipartite $ B_{r+1} $-free graph $ G $ with $ m \geq (240r)^2 $ edges, the spectral radius $ρ$ satisfies $ρ^2<m-1+\frac{2}{ρ-1}$, unless $G$ is isomorphic to $S^+_{m,s}$ for some $s\in\{1,\ldots,r\}$. This resolves Liu and Miao's conjecture and further reveals an interesting phenomenon: even with a weaker spectral condition, $ρ^2\geq m-1+\frac2{ρ-1}$, we can still derive the supersaturation of the booksize for non-bipartite graphs.

math.CO↗