SearcharxivSearch

arXiv subjects

Ying Wang

Publications and source records attributed to Ying Wang.

At least 19 recordsLinked to original sources

QROB: Quantifying Realization Overhead in Quantum Compilation via Reverse Construction

Quantum compilation reconciles a program's idealized interaction topology with hardware locality constraints, yet evaluations at scale lack calibrated references for realization overhead. We present QROB, a scalable reverse-construction methodology that generates compilation instances backward from directly realizable configurations, retaining the inverse paths as feasible, compiler-independent references. QROB provides a common evaluation substrate for NISQ SWAP routing and fault-tolerant lattice-surgery scheduling, while extending its reference-preserving principle to capacity-constrained quantum memory-access scheduling. Across systems ranging from 9 to 156 qubits, evaluations highlight QROB's utility as both a diagnostic benchmark and a data source. First, for compiler characterization, QROB reveals substantial realization gaps in existing tools, with NISQ compilers incurring up to 24.1x the reference SWAP cost and fault-tolerant compilers requiring up to 7.0x the reference makespan. Second, as a supervision source for data-driven compilation, a router trained on QROB references outperforms Qiskit SABRE on 84.8% of real-world application circuits. Finally, on real hardware, QROB reference realizations achieve a median mirror-circuit survival rate 1.65x that of full Qiskit O3 compilations across three 156-qubit IBM Heron-r2 processors, demonstrating that closing algorithmic compilation gaps translates directly into physical fidelity gains.

quant-ph

Movable Antennas Enabled Wireless Powered Networks: Principles and Technologies

As an emerging framework, movable antenna (MA)-enabled wireless powered networks (WPNs) have attracted growing attention. WPNs integrate wireless communication and energy transfer. MA can dynamically adjust the position of antenna units by introducing additional spatial degrees of freedom, so as to make full use of channel gain, optimize the effect of energy beamforming, and further improve the performance of WPNs. In this article, we first classify the implementations of MA, and review the fundamental principles of WPNs. We then highlight the key advantages of MA-enabled WPNs in enhancing wireless power transfer efficiency, realizing flexible and adaptive beamforming, and improving system robustness and interference resilience. Furthermore, four representative application scenarios and three key enabling technologies are discussed. A case study is also presented to show the improvement of energy harvesting performance brought by MA for WPNs. Finally, we discuss the challenges and future directions of MA-enabled WPNs, aiming to provide reference for future research and practice.

cs.NI

Bayes Estimators with Performance Comparable to Empirical Bayes Estimators and Improved Local Robustness

Bayes estimation has been extensively studied and widely used in statistics, decision theory, signal processing, machine learning, and system identification. Among its variants, empirical Bayes (EB) estimation has attracted considerable attention due to its favorable estimation performance and computational tractability. However, the direct plug-in dependence of an EB estimator on hyperparameters can make it locally sensitive to hyper-parameter perturbations. This paper considers the linear regression model and focuses on the EB estimator by employing the marginal maximum likelihood hyper-parameter estimator. For conciseness, this estimator is simply referred to as the EB estimator. Given a family of EB weighting functions, a generalized Bayes estimator is constructed with the same excess mean squared error (XMSE) as the corresponding EB estimator. Here, the XMSE is a second-order asymptotic measure of the mean squared error difference between the estimator of interest and the maximum likelihood estimator. Furthermore, the EB estimator is shown to be at most firstorder sensitive to hyper-parameter perturbations, whereas the constructed Bayes estimator is at most second-order sensitive, making it locally more robust. The computational complexities of these two estimators are also analyzed. In some cases, the constructed Bayes estimator can be computationally comparable to, or more efficient than, the EB estimator. These theoretical results are further supported by numerical simulations.

stat.ME

Adaptive Beam Hopping and Power Control for Dual-Layer Over-the-Air Online Federated Learning in LEO Satellite Networks

This paper investigates over-the-air (OTA) computation enabled online federated learning (FL) in low-Earth orbit (LEO) satellite networks. Specifically, we consider a dual-layer OTA aggregation architecture, where ground devices upload analog model updates to serving satellites via uplink OTA aggregation, and satellites forward the aggregated signals to a data processing center through the second round OTA aggregation. Then, we formulate a long-term data-utilization maximization problem in which devices continuously collect new data and untrained samples gradually lose freshness. The problem is subject to the satellite beam budget, transmit-power limit, and global mean squared error (MSE) constraint that governs end-to-end aggregation distortion. This yields a coupled mixed-integer nonlinear programming (MINLP) problem, involving tightly coupled discrete beam-hopping decisions and continuous power control. Due to the combinatorial action space and nonconvex constraints, the problem is NP-hard and computationally intractable. Furthermore, the time-varying satellite topology and dynamic data generation render it a sequential decision-making problem, necessitating adaptive online scheduling. To address these issues, we cast the problem as a Markov decision process and develop a proximal policy optimization (PPO)-based deep reinforcement learning framework that jointly optimizes adaptive beam hopping and power control, using an MSE-aware reward to balance data utilization and aggregation accuracy. Numerical simulation results verify that the proposed algorithm consistently outperforms other benchmark schemes, achieving superior long-term data utilization and faster FL convergence while satisfying the MSE requirement.

cs.IT

Contrastive Branch Policy Optimization

Reinforcement learning with verifiable rewards (RLVR) enables language models to learn multi-turn interaction with external tools, yet its sparse outcome rewards provide no signal for identifying which intermediate decisions are responsible for success. Branch sampling induces local comparisons among alternative continuations, but existing methods tend to conflate two distinct problems: allocating a fixed rollout budget and translating branch outcomes into token-level credit. We introduce Contrastive Branch Policy Optimization (CBPO), which disentangles these two problems and assigns a dedicated mechanism to each. Generation entropy screens candidate branch positions across the entire response, while path-level and node-level decay distribute a fixed budget across trajectories and positions to prevent exploration from collapsing onto a few paths or adjacent tokens. A parent trajectory together with the branches that share an identical token prefix forms an exact-prefix group, and the reward variation within this controlled group defines the Contrastive Branch Value (CBV), an outcome-based estimate of local decision sensitivity that rescales continuation advantages without altering their sign. When multiple nodes are selected along the same trajectory, CBPO partitions it into non-overlapping credit segments, thereby avoiding duplicated gradients on shared tokens. Requiring only outcome rewards and no process-level annotation, CBPO provides a practical solution for fine-grained credit assignment in tool-integrated agent training. Extensive experiments on ten benchmarks, including five for mathematical reasoning and five for knowledge-intensive search, show that CBPO consistently outperforms state-of-the-art policy-optimization and branch-based methods, attaining the highest macro-average accuracy in both domains and across two model scales.

cs.LG

Polymer Genome in the Age of Artificial Intelligence

Artificial intelligence (AI) is redefining the landscape of polymer science. Although numerous AI applications have been introduced in this field, the roles of polymer encoding strategies and different applications of AI models in polymer design remain insufficiently understood. Here, we build upon the foundation of polymer databases to critically examine the performance and applicability of current encoding strategies across different use cases. We then focus on two major AI application domains, property prediction and inverse design, to evaluate the strengths, weaknesses and suitable scenarios for various model architectures. Finally, we emphasize the significance of the online platforms for promoting data accessibility and accelerating the migration from experience-based discovery toward AI-driven innovation in polymer science and engineering. Through these discussions, we aim to provide practical guidance for future research and development in AI-assisted polymer design.

cond-mat.soft

Beyond Legal Spacing: A Residual-Aware Characterization of Entangling-Zone Spacing in Neutral-Atom Compilation

Neutral-atom processors rely on spatially arranged qubit arrays and parallel Rydberg entangling gates for scalable execution. Their compilers enforce geometric spacing rules for simultaneous gates, yet legal separation does not make residual van der Waals coupling disappear. This paper studies that gap between geometric legality and residual noise by treating entangling-zone spacing as a cross-layer reliability-parallelism variable anchored to experimental neutral-atom geometry. We combine fixed-schedule residual replay, surface-code simulation with matched correlated decoding, and fresh recompilation to connect spacing to physical residual exposure, logical reliability, and makespan cost. The results show that near-floor spacing can produce structured correlated exposure that is visible both at the physical layer and, in the tightest case, after quantum error correction (QEC). Modest geometric slack strongly suppresses this residual contribution, but the timing cost of looser spacing is mediated by placement and scheduling rather than by a simple monotonic slowdown. These findings distinguish hardware legality from residual-noise safety and motivate spacing-aware compiler evaluations that report physical geometry, QEC absorption, and scheduling cost together.

quant-ph

Where Atom Loss Lands Matters: Decoder-Aware Risk Deposition in Neutral-Atom QEC

Neutral-atom arrays are emerging as a leading platform for scalable quantum error correction (QEC). Qubits are routed and reused across the array, while detected loss is reported to the decoder as erasure information. Existing neutral-atom compilers optimize this movement, including routing, shuttling, and reuse, and often model loss through scalar exposure costs. Yet total exposure is an incomplete statistic for erasure-corrected QEC. It captures how much loss occurs, but not where it lands on the code, which we call its deposition. Under the same expected atom-loss budget, different deposition patterns over a code patch induce substantially different logical error rates (LER). We formalize this as decoder-aware risk deposition and present CAST, a compiler-side optimization pass that overlays a code-topology sensitivity map on a role-indexed exposure ledger and minimizes a decoder-weighted harm objective under a comparable-exposure constraint, using only local route, role, and seam-cooling actions. Across surface-code memory, physical-scale architecture models, lattice surgery, and decoder-mismatch checks, CAST lowers LER relative to topology-blind exposure minimization, improving on it in 35 of 48 physical-scale settings and by as much as 5.3x where exposure is heterogeneous and routing has slack. The largest gains occur when high exposure and high decoder sensitivity are initially misaligned, giving CAST room to redirect risk toward lower-impact code roles. CAST shows that decoder-aware atom-loss risk deposition can be optimized as a compiler-side pass in neutral-atom QEC.

quant-ph

Dual-Layer Over-the-Air Federated Learning in LEO Satellite Networks: Architecture, Key Technologies and Applications

Low Earth orbit (LEO) satellite networks are emerging as a pivotal infrastructure for global edge intelligence. In this context, integrating over-the-air (OTA) computation with adaptive beam hopping (BH) provides an innovative framework that seamlessly merges physical-layer analog aggregation with dynamic resource orchestration. This effectively overcomes the stringent bandwidth and power constraints of space platforms while extending federated learning (FL) to pervasive Internet-of-things (IoT) deployments. In this article, we first outline the fundamental principles of the dual-layer OTA model and introduce the adaptive BH mechanism designed for time-varying topologies. Then, we summarize the distinct advantages of this learning-centric architecture, which include decoupling aggregation latency from device density, optimizing spatio-temporal resource efficiency, and balancing data freshness with channel quality. Several application scenarios are explored to highlight the framework's potential across diverse vertical industries. Furthermore, a specific case is studied to demonstrate the practical efficacy of the proposed scheduling policy. The results reveal substantial performance gains in terms of model convergence speed and data utilization for satellite-based FL systems. Finally, we discuss the implementation challenges and outline future research directions, aiming to provide insights for the evolution of ubiquitous non-terrestrial intelligence.

eess.SP

Antenna Positioning and Beamforming Optimization in MA Enabled Secure ISAC Systems: A Gradient-Based Meta Learning Approach

Integrated sensing and communications (ISAC) significantly improves spectral efficiency but introduces security risks regarding the interception of embedded communication signals. This paper proposes an movable antenna (MA)-enabled secure ISAC system that utilizes the spatial degrees of freedom of MA to mitigate these risks. Then, a problem is formulated to maximize the system secrecy rate by jointly optimizing antenna positioning, transmit beamforming, and artificial noise. However, the principal challenge arises from the non-convexity of the optimization problem and the strong coupling of the optimization variables. Generally, traditional optimization methods for this problem suffer from complex mathematical derivations, while existing deep learning approaches rely heavily on the training data distribution. To address these issues, we introduce a gradient-based meta learning (GML) algorithm, which works without pre-training and demonstrates favorable performance. Specifically, the algorithm establishes a neural network for each optimization variable, where the gradient of the objective function with respect to the variable serves as the input, and the output of the network determines the variable's update step. By handling the constraints and constructing penalty terms, the global loss function is used to guide the optimization process. Extensive numerical simulations confirm that the proposed algorithm achieves satisfactory performance in terms of both communication security and sensing capabilities.

eess.SP

GML-Based Optimization for Movable Antenna Wireless Networks: Challenges and Opportunities

Movable antenna (MA) is proposed as an emerging technology for future wireless networks. By leveraging the additional spatial degrees of freedom, MA can proactively reshape the wireless propagation environment, thereby enhancing network performance.However, fully unlocking the potential of MA networks necessitates the joint optimization of MA antenna positioning and beamforming. For this non-convex and highly coupled problem, existing solutions exhibit significant limitations. Therefore, this paper proposes a gradient-based meta learning (GML) optimization framework. Specifically, we first elaborate on the hardware architecture and channel characteristics of MA, based on which we analyze the primary challenges in optimizing MA wireless networks. Subsequently, we introduce the fundamental logic of the GML framework and compare it with existing methods. Furthermore, we discuss the constraint handling strategies for applying the proposed optimization framework to MA networks. A specific case is studied to show the performance of proposed framework based on numerical simulation. Finally, this paper outlines future research directions for both the GML framework and MA wireless networks.

eess.SP

Scattering Theory For 3D Cubic Damped Magnetic Schr\"odinger Equation

We consider the three-dimensional defocusing cubic nonlinear Schr\"odinger equation with variable coefficients, a magnetic potential, and a non-negative localized damping term, \[ i\partial_tu+(\nabla-iA)\cdot G(\nabla-iA)u+ia(x)u=|u|^2u, \qquad t>0,\quad x\in\mathbb R^3. \] No non-trapping condition is imposed on the metric $G$. Instead, the variable-coefficient region is assumed to be contained in the effective damping region. Under a one-centre condition on the tangential magnetic field, we prove global well-posedness for initial data in $H^{1+\varepsilon}$, uniform mass and energy bounds, and show the local energy decay. To obtain scattering, we impose a support condition on the full magnetic field inside the damping region. Under these stronger assumptions, the solution scatters to a free Schr\"odinger evolution in $H^s$ for every $0\le s<1$. The appendix discusses a separate constant-damping framework for abstract Hamiltonians.

math.AP

Hankel Transform and $(\alpha,\beta)$ Somos-4 Sequences

An $(\alpha,\beta)$ Somos-$4$ sequence $S_n$ is defined by the recurrence $S_nS_{n-4}=\alpha S_{n-1}S_{n-3}+\beta S_{n-2}^2$ ($n\geq 4$), with suitable initial values, where $\alpha$ and $\beta$ are constant parameters. A widely studied question is the following: When does the Hankel transform of a generating function become an $(\alpha,\beta)$ Somos-4 sequence? In particular, how can $\alpha$ and $\beta$ be derived for such a function? A sufficient condition for this problem has been established by Wang and Zhang. In this paper, we obtain the following three main results. (i): We extend the Wang--Zhang sufficient condition by working over the rational function field. Then we combine this result with the Sulanke--Xin quadratic transformation to resolve all of Barry's currently unsolved $(\alpha,\beta)$ Somos-4 conjectures, which arise in diverse contexts, including generalized Catalan recurrences, Riordan arrays, generalized Bernstein arrays, and elliptic curves. (ii): We show that the odd and even subsequences of an $(\alpha,\beta)$ Somos-4 sequence are again $(\alpha,\beta)$ Somos-4 sequences with transformed parameters. This is employed to establish Barry's Hurwitz transform conjecture. (iii): Using the theory of orthogonal polynomials, we prove a Hankel determinant formula and thereby prove a conjecture related to the $(\alpha,\beta)$ Somos-4 sequence. In addition, we prove some conjectures on formulas for periodic Hankel determinants.

math.CO

Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM

Heterogeneous architectures that combine neural processing unit (NPU) and processing-in-memory (PIM) are increasingly adopted to accelerate LLM inference. Prior work focuses on building a unified memory that allows NPUs and PIM to share data without duplication. However, these designs implicitly assume that each tensor is bound to a fixed execution device, and therefore rely on static, device-biased data mappings. We observe that this assumption does not hold in modern LLM workloads. Due to phase changes (e.g., prefill vs. decode) and dynamic behaviors such as MoE routing, the optimal execution device for the same tensor can change at runtime. Under such dynamic execution, device-biased mappings become mismatched to access patterns, leading to substantial bandwidth underutilization and performance loss. This paper presents PFM (PIM-as-Flexible-Memory), a dual-view memory system that decouples physical data layout from accessor-visible logical views. PFM stores data in a jointly optimized physical layout and exposes different logical interpretations to NPUs and PIM, enabling efficient access across devices without data duplication or relayout. We further design accessor-aware address translation and runtime scheduling mechanisms to support dynamic execution when LLM workloads fluctuate and the optimal execution device dynamically changes. Our evaluation across LLMs shows that PFM improves end-to-end throughput by up to 2.32$\times$, demonstrating its effectiveness and broad applicability as a unified memory management solution for NPU-PIM systems.

cs.AR

MicroEvo: Knowledge-Guided LLM Sampling for Efficient Microarchitecture Design Space Exploration

Microarchitecture design space exploration suffers from expansive search spaces and expensive PPA evaluation, leaving only a small simulation budget for design decision-making. Existing methods perform blind search without considering microarchitectural dependencies and fail to learn from the iterative search effectively, leading to wasted evaluations and weak Pareto convergence. In this paper, we propose MicroEvo, a knowledge-guided framework that couples off-the-shelf LLMs with Monte Carlo Tree Search (MCTS) for multi-objective microarchitecture optimization. MicroEvo combines LLM-driven evolutionary operators, a Pareto-aware tree policy that balances Pareto contribution and diversity, an active knowledge accumulation mechanism that extracts and reuses optimization insights, and state-aware directives that adapt the search behavior online. Experiments show that MicroEvo improves Pareto-front quality by up to 36.2% over NSGA-II and achieves 10.6x higher search efficiency, and also demonstrates strong scalability to a complex industrial-scale core. The code repository is available at: https://github.com/GEAR-SEU/MicroEvo-ICCAD-26.

cs.AI

Pattern Zooming: Near-Field Wideband Beam Training with Wavenumber-Domain Codebook

Near-field beam training is essential for harvesting the high-gain potential of extremely large-scale multiple input multiple output systems. To reduce training overhead, existing works have mostly leveraged the beam squint effect by utilizing time-delay (TD) beamforming with polar-domain codebook, which enables the simultaneous sweeping of multiple angles at a specific distance. However, such methods still suffer from high overhead due to the exhaustive distance searching. To address this challenge, we propose a pattern zooming based near-field wideband beam training with wavenumber-domain codebook. Specifically, we first establish a refined wideband Fourier planewave channel representation, based on which we reveal a pattern zooming effect, where the wavenumber-domain patterns across different subcarriers exhibit frequency-dependent scaling relative to the center frequency. Exploiting this property, we develop a TD-assisted beam sweeping strategy that simultaneously probes multiple wavenumber directions to rapidly acquire the complete wavenumber-domain pattern. Based on the acquired pattern, we further derive exact, approximation-free closed-form expressions that establish a rigorous mapping between the receiver coordinates and the wavenumber-domain pattern, enabling accurate user localization with only a few pilots. Finally, numerical results validate the superiority of our proposed scheme in terms of both beamforming gain and training overhead.

eess.SP

AP Association for RHS-Enabled Cell-Free Uplink MIMO in Industrial Indoor UAV Networks

Indoor industrial UAV uplink networks face serious blockage and shadowing from shelves, metal equipment, and production facilities. UAVs are also often clustered and fly along similar straight inspection routes at fixed heights. These features make traditional small-cell deployment less suitable, especially when high reliability, continuous coverage, and good service for weak UAVs are required. Cell-free networks can improve robustness through distributed access points (APs) and UAV?centric communications. Reconfigurable holographic surface (RHS)-enabled APs provide programmable analog receive beams and generate scalar post-RHS observations, which are jointly processed at the CPU for distributed uplink MIMO detection at relatively low hardware cost. Conventional AP association relies on distance, large-scale fading, or post-combining SINR obtained with user-specific digital combiners. Here, however, all UAVs served by a single-feed RHS AP share one amplitude?constrained receive pattern and one scalar AP output. We therefore derive an SINR-like score from this physical output and, under weak inter-AP disturbance correlation, approximate the CPU-side log-det objective by an additive per-AP surrogate, yielding a low-complexity ranking rule. The results show that the nearest AP is not always the best choice, the AP-UAV height difference may have an optimal value, and larger serving clusters bring diminishing returns. Simulations show that the proposed method improves the minimum UAV data rate, average spectral efficiency, fairness, and energy efficiency compared with benchmark schemes.

cs.NI

RTLCurator: Label-Efficient Data Curation for RTL Generation

Training large language models (LLMs) to write register-transfer level (RTL) requires large corpora of paired specifications and code, and such data is scarce enough that most public corpora are now synthesized. Synthesis provides scale but not correctness, and in two widely used RTL datasets only 24.4% and 53.5% of pairs pass generated functional tests. This raises the question of how much of such a corpus to keep and which part of it. Correctness alone is a poor answer. A pair that misbehaves in one corner case still shows valid syntax and interface conventions, and complex sequential designs are both harder to generate and harder to validate, so filtering by correctness leaves a corpus of short and simple modules. Correctness is also hard to obtain, since behavior leaves little trace on the surface in RTL, and validating an entire corpus only sorts pairs into passed and failed. We present RTLCurator, which learns a behavior-aware compatibility prior by contrasting each specification with implementations that fail simulation, and calibrates it to a new corpus using a small number of validated pairs. It then constructs the retained subset by balancing alignment, representation coverage, and RTL structural richness. On CodeV and RTLCoder, keeping 80% of the corpus this way improves on training with the full corpus across all reported metrics while validating only 10% of the pool, whereas ranking by the score alone falls below random selection and filtering the whole pool by simulation does no better.

cs.AR