SearcharxivSearch

arXiv subjects

Ke Wei

Publications and source records attributed to Ke Wei.

At least 19 recordsLinked to original sources

On the Policy Convergence of Policy Mirror Descent Methods

We study the policy convergence of unregularized policy mirror descent (PMD) with arbitrary constant step sizes for finite discounted Markov decision processes. We focus on decomposable mirror maps of the form $h(p)=\sum_a \psi(p(a))$, where $\psi$ satisfies standard Legendre-type assumptions. Under these conditions, we prove that the policy sequence generated by PMD converges in the policy domain to a limiting optimal policy, even when the optimal policy set is not a singleton. This result covers a broad class of commonly used mirror maps, including the squared Euclidean mirror map underlying projected Q-ascent, the negative Shannon entropy underlying softmax natural policy gradient, Tsallis entropy, the Hellinger mapping, and the Fermi-Dirac entropy. Although policy convergence has been established previously for specific PMD instances or for regularized variants such as homotopic PMD, to the best of our knowledge, this is the first systematic and unified policy convergence theory for unregularized PMD under general decomposable mirror maps and arbitrary constant step sizes. Our analysis further reveals that the convergence behavior is governed by the differentiability of $\psi$ at $0$ and $1$, leading to different behaviors, including finite-time termination, asymptotic convergence, and an MDP-dependent dichotomy. When $\psi$ is twice continuously differentiable with strictly positive finite curvature, we further establish local policy convergence rates for the asymptotic convergence cases, covering the standard mirror maps mentioned above.

math.OC

Depletion-mode N-polar AlN-based high electron mobility transistors with improved on/off ratios

We report N-polar AlN-based high-electron mobility transistors (HEMTs) with a GaN channel thickness of 5.2 nm on N-polar AlN on sapphire. The threshold voltage is around -2.4 to -3.0 V with saturation currents over 240 mA/mm and on/off ratios as high as 10,000, much higher than previously reported N-polar AlN-based HEMTs. The high on/off ratio is attributed to the use of an abrupt AlN/GaN heterostructure with a dedicated AlN transition layer, together with improved gate leakage. The high frequency properties as well as the on-resistance of ~20 Ohm mm are all limited by the 2000 Ohm/square sheet resistance of the channel layer.

cond-mat.mtrl-sci

Downlink Channel Matrix Estimation from PMI-Only Feedback in FDD Systems: Maximum Likelihood and Sharp Excess Risk Bound

We study downlink channel estimation in a frequency-division duplex (FDD) massive MIMO system from PMI-only feedback under a 5G NR-type limited-feedback architecture. In this architecture, the user selects a preferred codeword from a shared codebook based on the reduced-dimensional channel and only reports its index (known as the precoding matrix indicator, PMI) back to the base station. Therefore, the channel must be estimated from these highly quantized, nonlinear PMI observations. Based on a probabilistic perturbation model, a constrained maximum likelihood estimator (MLE) is proposed for this estimation problem, whose objective can also be interpreted as a relaxation of the hard empirical decision error. The Cram\'er--Rao bound is derived for the complex-valued model, with the global phase ambiguity handled via gauge-fixing. For the real-valued setting, a global excess-risk bound of order $O(1/\sqrt{T})$ is established, which is then refined to a sharp local rate of order $O(1/T)$ under suitable identifiability conditions. Numerical results show that the MLE asymptotically attains the Cram\'er--Rao bound and outperforms several baseline methods on both synthetic data and realistic FDD channels.

cs.IT

Structure-Aware Distributed Backdoor Attacks in Federated Learning

While federated learning protects data privacy, it also makes the model update process vulnerable to long-term stealthy perturbations. Existing studies on backdoor attacks in federated learning mainly focus on trigger design or poisoning strategies, typically assuming that identical perturbations behave similarly across different model architectures. This assumption overlooks the impact of model structure on perturbation effectiveness. From a structure-aware perspective, this paper analyzes the coupling relationship between model architectures and backdoor perturbations. We introduce two metrics, Structural Responsiveness Score (SRS) and Structural Compatibility Coefficient (SCC), to measure a model's sensitivity to perturbations and its preference for fractal perturbations. Based on these metrics, we develop a structure-aware fractal perturbation injection framework (TFI) to study the role of architectural properties in the backdoor injection process. Experimental results show that model architecture significantly influences the propagation and aggregation of perturbations. Networks with multi-path feature fusion can amplify and retain fractal perturbations even under low poisoning ratios, while models with low structural compatibility constrain their effectiveness. Further analysis reveals a strong correlation between SCC and attack success rate, suggesting that SCC can predict perturbation survivability. These findings highlight that backdoor behaviors in federated learning depend not only on perturbation design or poisoning intensity but also on the interaction between model architecture and aggregation mechanisms, offering new insights for structure-aware defense design.

cs.LG

Decentralized Non-convex Stochastic Optimization with Heterogeneous Variance

Decentralized optimization is critical for solving large-scale machine learning problems over distributed networks, where multiple nodes collaborate through local communication. In practice, the variances of stochastic gradient estimators often differ across nodes, yet their impact on algorithm design and complexity remains unclear. To address this issue, we propose D-NSS, a decentralized algorithm with node-specific sampling, and establish its sample complexity depending on the arithmetic mean of local standard deviations, achieving tighter bounds than existing methods that rely on the worst-case or quadratic mean. We further derive a matching sample complexity lower bound under heterogeneous variance, thereby proving the optimality of this dependence. Moreover, we extend the framework with a variance reduction technique and develop D-NSS-VR, which under the mean-squared smoothness assumption attains an improved sample complexity bound while preserving the arithmetic-mean dependence. Finally, numerical experiments validate the theoretical results and demonstrate the effectiveness of the proposed algorithms.

math.OC

Stability and Generalization of Nonconvex Optimization with Heavy-Tailed Noise

The empirical evidence indicates that stochastic optimization with heavy-tailed gradient noise is more appropriate to characterize the training of machine learning models than that with standard bounded gradient variance noise. Most existing works on this phenomenon focus on the convergence of optimization errors, while the analysis for generalization bounds under the heavy-tailed gradient noise remains limited. In this paper, we develop a general framework for establishing generalization bounds under heavy-tailed noise. Specifically, we introduce a truncation argument to achieve the generalization error bound based on the algorithmic stability under the assumption of bounded $p$th centered moment with $p\in(1,2]$. Building on this framework, we further provide the stability and generalization analysis for several popular stochastic algorithms under heavy-tailed noise, including clipped and normalized stochastic gradient descent, as well as their mini-batch and momentum variants.

cs.LG

Ultralow-noise microwave oscillator via optical frequency division with a co-self-injection-locked miniature Fabry-Perot reference

Optical frequency division (OFD) provides the purest microwaves by down-converting the stability of optical cavity references. State-of-the-art references typically rely on electronic co-Pound-Drever-Hall locking to ultrahigh-Q microresonators-a complex approach that introduces servo bumps and increases footprint. Alternatively, optical co-self-injection-locking (co-SIL) offers inherent simplicity but is limited by the large thermo-refractive noise and confined mode volumes of integrated cavities. Here, we demonstrate a two-point OFD-based microwave oscillator that combines an ultrahigh-Q miniature Fabry-Perot cavity with optical co-SIL. Leveraging its low relative phase noise optical reference and combing with an integrated soliton microcomb, the system generates a microwave with phase noise of -147 dBc/Hz at 4 kHz offset (scaled to 10 GHz)-performance rivalling most electronically stabilized systems. This work marries the superior noise floor of ultrahigh-Q cavities with the simplicity of optical locking, providing a compact, cost-effective, and field-deployable path to pure microwaves for next-generation communications, radar and metrology.

physics.optics

Fast and Provable Nonconvex Robust Matrix Completion

We study the robust matrix completion (RMC) problem subject to both sparse outliers and stochastic noise. A non-convex method termed Accelerated Robust Matrix Completion (ARMC) is proposed, which accelerates a prior non-convex approach by incorporating an explicit subspace projection step into the low-rank update, leading to significantly improved computational efficiency. Through a delicate analysis based on the leave-one-out technique, the entrywise linear convergence guarantee of ARMC has been established. Notably, the derived bounds for sample complexity and outlier sparsity improve upon existing guarantees of the convex relaxation approach that also accounts for both sparse outliers and stochastic noise. Moreover, numerical experiments on synthetic and real data show that ARMC is superior to existing non-convex RMC methods.

cs.IT

Policy Mirror Descent with Temporal Difference Learning: Sample Complexity under Online Markov Data

This paper studies the policy mirror descent (PMD) method, which is a general policy optimization framework in reinforcement learning and can cover a wide range of policy gradient methods by specifying difference mirror maps. Existing sample complexity analysis for policy mirror descent either focuses on the generative sampling model, or the Markovian sampling model but with the action values being explicitly approximated to certain pre-specified accuracy. In contrast, we consider the sample complexity of policy mirror descent with temporal difference (TD) learning under the Markovian sampling model. Two algorithms called Expected TD-PMD and Approximate TD-PMD have been presented, which are off-policy and mixed policy algorithms respectively. Under a small enough constant policy update step size, the $\tilde{O}(\varepsilon^{-2})$ (a logarithm factor about $\varepsilon$ is hidden in $\tilde{O}(\cdot)$) sample complexity can be established for them to achieve average-time $\varepsilon$-optimality. The sample complexity is further improved to $O(\varepsilon^{-2})$ (without the hidden logarithm factor) to achieve the last-iterate $\varepsilon$-optimality based on adaptive policy update step sizes.

math.OC

On the Convergence of Policy Mirror Descent with Temporal Difference Evaluation

Policy mirror descent (PMD) is a general policy optimization framework in reinforcement learning, which can cover a wide range of typical policy optimization methods by specifying different mirror maps. Existing analysis of PMD requires exact or approximate evaluation (for example unbiased estimation via Monte Carlo simulation) of action values solely based on policy. In this paper, we consider policy mirror descent with temporal difference evaluation (TD-PMD). It is shown that, given the access to exact policy evaluations, the dimension-free $O(1/T)$ sublinear convergence still holds for TD-PMD with any constant step size and any initialization. In order to achieve this result, new monotonicity and shift invariance arguments have been developed. The dimension free $\gamma$-rate linear convergence of TD-PMD is also established provided the step size is selected adaptively. For the two common instances of TD-PMD (i.e., TD-PQA and TD-NPG), it is further shown that they enjoy the convergence in the policy domain. Additionally, we investigate TD-PMD in the inexact setting and give the sample complexity for it to achieve the last iterate $\varepsilon$-optimality under a generative model, which improves the last iterate sample complexity for PMD over the dependence on $1/(1-\gamma)$.

math.OC

Discovering heuristics in a complex SAT solver with large language models

The Satisfiability problem (SAT) is fundamental in computational complexity theory and has a wide range of industrial applications. Optimizing modern SAT solvers in real-world settings is quite challenging due to their intricate architectures. While automatic configuration frameworks have been developed, they rely on manually constrained search spaces. Here we develop AutoModSAT, a framework that uses large language models (LLMs) to automatically optimize SAT solvers. AutoModSAT combines an LLM-compatible modular solver design, unsupervised prompt optimization to diversify generated functions, and an efficient search procedure based on presearch strategy and a $(1+\lambda)$ evolutionary algorithm. Extensive experiments across a wide range of datasets demonstrate that AutoModSAT achieves $40\%$ performance improvement over the baseline solver and $30\%$ improvement over the state-of-the-art solvers. Moreover, AutoModSAT also attains a notable speedup compared to the parameter-tuned alternatives of the state-of-the-art solvers over most of the test datasets. These results demonstrate the potential of LLM-guided heuristic discovery for optimizing complex SAT solvers.

cs.AI

Strong Molecule-Light Entanglement with Molecular Cavity Optomechanics

We propose a molecular optomechanical platform to generate robust entanglement among bosonic modes-photons, phonons, and plasmons-under ambient conditions. The system integrates an ultrahigh-Q whispering-gallery-mode (WGM) optical resonator with a plasmonic nanocavity formed by a metallic nanoparticle and a single molecule. This hybrid architecture offers two critical advantages over standalone plasmonic systems: (i) Efficient redirection of Stokes photons from the lossy plasmonic mode into the long-lived WGM resonator, and (ii) Suppression of molecular absorption and approaching vibrational ground states via plasmon-WGM interactions. These features enable entanglement to transfer from the fragile plasmon-phonon subsystem to a photon-phonon bipartition in the blue-detuned regime, yielding robust stationary entanglement resilient to environmental noise. Remarkably, the achieved entanglement surpasses the theoretical bound for conventional two-mode squeezing in certain parameter regimes. Our scheme establishes a universal approach to safeguard entanglement in open quantum systems and opens avenues for noise-resilient quantum information technologies.

quant-ph

Molecular optomechanically-induced transparency

Molecular cavity optomechanics (COM), characterized by remarkably efficient optomechanical coupling enabled by a highly localized light field and ultra-small effective mode volume, holds significant promise for advancing applications in quantum science and technology. Here, we study optomechanically induced transparency and the associated group delay in a hybrid molecular COM system. We find that even with an extremely low optical quality factor, an obvious transparency window can appear, which is otherwise unattainable in a conventional COM system. Furthermore, by varying the ports of the probe light, the optomechanically induced transparency or absorption can be achieved, along with corresponding slowing or advancing of optical signals. These results indicate that our scheme provides a new method for adjusting the storage and retrieval of optical signals in such a molecular COM device.

physics.optics

MSV-Mamba: A Multiscale Vision Mamba Network for Echocardiography Segmentation

Ultrasound imaging frequently encounters challenges, such as those related to elevated noise levels, diminished spatiotemporal resolution, and the complexity of anatomical structures. These factors significantly hinder the model's ability to accurately capture and analyze structural relationships and dynamic patterns across various regions of the heart. Mamba, an emerging model, is one of the most cutting-edge approaches that is widely applied to diverse vision and language tasks. To this end, this paper introduces a U-shaped deep learning model incorporating a large-window Mamba scale (LMS) module and a hierarchical feature fusion approach for echocardiographic segmentation. First, a cascaded residual block serves as an encoder and is employed to incrementally extract multiscale detailed features. Second, a large-window multiscale mamba module is integrated into the decoder to capture global dependencies across regions and enhance the segmentation capability for complex anatomical structures. Furthermore, our model introduces auxiliary losses at each decoder layer and employs a dual attention mechanism to fuse multilayer features both spatially and across channels. This approach enhances segmentation performance and accuracy in delineating complex anatomical structures. Finally, the experimental results using the EchoNet-Dynamic and CAMUS datasets demonstrate that the model outperforms other methods in terms of both accuracy and robustness. For the segmentation of the left ventricular endocardium (${LV}_{endo}$), the model achieved optimal values of 95.01 and 93.36, respectively, while for the left ventricular epicardium (${LV}_{epi}$), values of 87.35 and 87.80, respectively, were achieved. This represents an improvement ranging between 0.54 and 1.11 compared with the best-performing model.

eess.IV

Leave-One-Out Analysis for Nonconvex Robust Matrix Completion with General Thresholding Functions

We study the problem of robust matrix completion (RMC), where the partially observed entries of an underlying low-rank matrix is corrupted by sparse noise. Existing analysis of the non-convex methods for this problem either requires the explicit but empirically redundant regularization in the algorithm or requires sample splitting in the analysis. In this paper, we consider a simple yet efficient nonconvex method which alternates between a projected gradient step for the low-rank part and a thresholding step for the sparse noise part. Inspired by leave-one out analysis for low rank matrix completion, it is established that the method can achieve linear convergence for a general class of thresholding functions, including for example soft-thresholding and SCAD. To the best of our knowledge, this is the first leave-one-out analysis on a nonconvex method for RMC. Additionally, when applying our result to low rank matrix completion, it improves the sampling complexity of existing result for the singular value projection method.

cs.IT

Elementary Analysis of Policy Gradient Methods

Projected policy gradient under the simplex parameterization, policy gradient and natural policy gradient under the softmax parameterization, are fundamental algorithms in reinforcement learning. There have been a flurry of recent activities in studying these algorithms from the theoretical aspect. Despite this, their convergence behavior is still not fully understood, even given the access to exact policy evaluations. In this paper, we focus on the discounted MDP setting and conduct a systematic study of the aforementioned policy optimization methods. Several novel results are presented, including 1) global linear convergence of projected policy gradient for any constant step size, 2) sublinear convergence of softmax policy gradient for any constant step size, 3) global linear convergence of softmax natural policy gradient for any constant step size, 4) global linear convergence of entropy regularized softmax policy gradient for a wider range of constant step sizes than existing result, 5) tight local linear convergence rate of entropy regularized natural policy gradient, and 6) a new and concise local quadratic convergence rate of soft policy iteration without the assumption on the stationary distribution under the optimal policy. New and elementary analysis techniques have been developed to establish these results.

math.OC

AutoSAT: Automatically Optimize SAT Solvers via Large Language Models

Conflict-Driven Clause Learning (CDCL) is the mainstream framework for solving the Satisfiability problem (SAT), and CDCL solvers typically rely on various heuristics, which have a significant impact on their performance. Modern CDCL solvers, such as MiniSat and Kissat, commonly incorporate several heuristics and select one to use according to simple rules, requiring significant time and expert effort to fine-tune in practice. The pervasion of Large Language Models (LLMs) provides a potential solution to address this issue. However, generating a CDCL solver from scratch is not effective due to the complexity and context volume of SAT solvers. Instead, we propose AutoSAT, a framework that automatically optimizes heuristics in a pre-defined modular search space based on existing CDCL solvers. Unlike existing automated algorithm design approaches focusing on hyperparameter tuning and operator selection, AutoSAT can generate new efficient heuristics. In this first attempt at optimizing SAT solvers using LLMs, several strategies including the greedy hill climber and (1+1) Evolutionary Algorithm are employed to guide LLMs to search for better heuristics. Experimental results demonstrate that LLMs can generally enhance the performance of CDCL solvers. A realization of AutoSAT outperforms MiniSat on 9 out of 12 datasets and even surpasses the state-of-the-art hybrid solver Kissat on 4 datasets.

cs.AI

Global Convergence of Natural Policy Gradient with Hessian-aided Momentum Variance Reduction

Natural policy gradient (NPG) and its variants are widely-used policy search methods in reinforcement learning. Inspired by prior work, a new NPG variant coined NPG-HM is developed in this paper, which utilizes the Hessian-aided momentum technique for variance reduction, while the sub-problem is solved via the stochastic gradient descent method. It is shown that NPG-HM can achieve the global last iterate $\epsilon$-optimality with a sample complexity of $\mathcal{O}(\epsilon^{-2})$, which is the best known result for natural policy gradient type methods under the generic Fisher non-degenerate policy parameterizations. The convergence analysis is built upon a relaxed weak gradient dominance property tailored for NPG under the compatible function approximation framework, as well as a neat way to decompose the error when handling the sub-problem. Moreover, numerical experiments on Mujoco-based environments demonstrate the superior performance of NPG-HM over other state-of-the-art policy gradient methods.

cs.LG