Searcharxiv⌕ Search

arXiv subjects

Peng Yang

Publications and source records attributed to Peng Yang.

At least 37 records · Page 2Linked to original sources

FreCast: Refining Radar Echo Intensity via Phase-Preserving Amplitude Residual Diffusion for Precipitation Nowcasting

Precipitation nowcasting predicts the spatiotemporal evolution of future radar echoes from historical radar echo sequences, thereby estimating the occurrence, development, and movement of precipitation over the near term. In recent years, deep learning has become an important approach to precipitation nowcasting. Although state-of-the-art models can generally capture the overall spatial distribution of future precipitation, their predictions still exhibit substantial biases in radar echo intensity at individual locations. This observation motivates a more targeted strategy for reducing forecast errors. Instead of regenerating an entire radar echo sequence without spatial constraints, the predicted precipitation structure can be used to guide the refinement of echo intensities at individual locations. This structure-guided refinement directly targets echo intensity biases. Accordingly, we propose FreCast, a two-stage framework for radar echo prediction. The first stage generates an initial forecast of future radar echoes. The second stage uses the spatial structure of the initial forecast as a constraint to further correct intensity biases at individual locations in the first-stage prediction. Experiments on three datasets demonstrate that FreCast achieves consistent improvements across forecast skill metrics. Qualitative results further show that FreCast better preserves rainband continuity and intense precipitation structures at longer lead times.

cs.CV↗

The semileptonic decays of $\mathcal{B}_{Q_{1}Q_{2}}(\frac{1}{2}^{+})\rightarrow\mathcal{B}_{Q_{1}}^{*}(\frac{3}{2}^{+})$ in QCD sum rules

In the framework of QCD sum rules, we systematically analyze the weak transition process $\mathcal{B}_{Q_{1}Q_{2}}(\frac{1}{2}^{+})\rightarrow\mathcal{B}_{Q_{1}}^{*}(\frac{3}{2}^{+})$. When doing the operator product expansion in the QCD side, we consider the contributions of perturbative part and vacuum condensate terms up to dimension 6. In the phenomenological side, we eliminate the interferences of the low spin states and negative parity states by employing 16 different dirac structures. As an application, these form factors are finally used to analyze the semileptonic decays of $\mathcal{B}_{Q_{1}Q_{2}}(\frac{1}{2}^{+})\rightarrow\mathcal{B}_{Q_{1}}^{*}(\frac{3}{2}^{+})lν$, where these decays are driven by the transition processes $c\rightarrow d/s+l^{+}+ν_{l}$ and $b\rightarrow u+l^{-}+\overlineν_{l}$. The predicted physical quantities include not only the partial widths, ratios of $Γ_{L}/Γ_{T}$ and the branching fractions, but also some observables such as the forward-backward asymmetry parameter $A_{FB}^{l}$ of lepton, the $P_z^{F}$ component of the polarization vector for daughter baryon and the longitudinal polarization of the lepton $P_z^{l}$. We hope all of these theoretical predictions about the weak decays will be helpful for studying the properties of doubly heavy baryons in experiments in the future.

hep-ph↗

Quasinormal Modes of Extremal Reissner-Nordstrom Black Holes via Seiberg-Witten Quantization

We study the scalar perturbations of asymptotically flat extremal Reissner-Nordström black holes via the quantum Seiberg-Witten geometry of $\mathcal{N}=2$ SU(2) gauge theory with $N_f=2$ flavors. The radial master equation, governed by a double confluent Heun equation, is exactly mapped to the quantum Seiberg-Witten curve, providing an exact quantization condition derived from the non-perturbative Nekrasov-Shatashvili free energy. Analytically, this exact dictionary unveils precise gauge-theoretic interpretations for critical physical thresholds, demonstrating that the superradiance and mass decoupling limits naturally reduce the master equation to the Whittaker equation and the reduced doubly confluent Heun equation (the latter corresponds to the SW geometry of the $\mathcal{N}=2$ SU(2) gauge theory with $N_f=1$), respectively. At the strict extremal limit, the coalescence of horizons induces a topological singularity that complicates the spectral analysis. By accommodating this irregular singularity, our geometric framework resolves the singularity coalescence and enables the extraction of the discrete global quasinormal mode. As our main contribution, we provide the first non-perturbative evaluation of the quasinormal modes spectrum for simultaneously charged and massive scalar fields directly at strict extremity. Furthermore, our analytical results reproduce numerical benchmarks for both neutral and charged massless probes, and naturally capture quasi-resonance behaviors.

hep-th↗

Rare Events Govern Defect Formation under Weak Symmetry Breaking

Crossing a continuous phase transition out of equilibrium typically generates topological defects whose density obeys a universal power-law scaling predicted by the Kibble-Zurek mechanism. Recent numerical studies have revealed systematic deviations from this scaling in the presence of weak explicit symmetry breaking, manifested as an additional exponential suppression of defect formation. However, the origin of this correction and a general theoretical framework to describe it have remained elusive. Here, using large-deviation theory, we show that defect formation under weak symmetry breaking is controlled by rare fluctuations that drive local regions into the disfavored symmetry-broken state. This mechanism yields a closed-form expression for the defect density in arbitrary dimensions, valid in the weak-field and weak-noise limits. These theoretical predictions are verified through direct simulations of stochastic Ginzburg-Landau models in one and two spatial dimensions.

cond-mat.stat-mech↗

Decoupling Limit of Quiver Theories and the Angular Spectra of Extreme C-metrics

We investigate the angular eigenvalue problem of the extreme charged C-metric. In the extreme limit ($Q \to M$), the governing differential equation degenerates from a Fuchsian equation with five regular singular points into a Confluent Extended Heun Equation. To evaluate the angular spectrum analytically, we formulate a decoupling limit within the dual four-dimensional $\mathcal{N}=2$, $\mathrm{SU(2)}\times \mathrm{SU(2)}$ linear quiver gauge theory. Within this framework, we derive the parameter dictionary and renormalized Matone relations, which absorb the macroscopic residue shifts induced by the singularity fusion. Based on the regular boundary conditions of the angular equation, we utilize the instanton counting method to establish an algebraic quantization condition, yielding angular eigenvalues consistent with numerical results.

hep-th↗

RiverONE: Generating Knowledge-Intensive VLM by Simulated Quantum Machines

Quantum computing provides a powerful paradigm for representing and transforming high-dimensional information through superposition, entanglement, and measurement-induced nonlinear features. While current quantum hardware is not yet practical for direct large-scale vision-language model (VLM) inference, simulated quantum computation can be used during model construction to generate structured parameters for compact classical AI systems. We build RiverONE, a lightweight vision-language model for quantum calibration plot understanding, using simulated quantum computation. It employs a specialized visual encoder and an InternVL-based language backbone. To compensate for compression-induced information loss, we introduce quantum-generated parameters, which are materialized as classical tensors after training. This allows RiverONE to run entirely on classical GPUs at inference time, with no quantum hardware or runtime quantum simulation. With approximately 1.9 billion parameters, RiverONE achieves at least 95\% of the performance of NVIDIA Ising Calibration 1 on quantum calibration plot understanding tasks while using less than 10\% of its parameter count. These results suggest that simulated quantum computation can serve as a practical construction-stage mechanism for building lightweight, knowledge-intensive scientific VLMs. Our code is available at https://github.com/THeWakeSystems/RiverOne.

quant-ph↗

Adaptive Plug-and-Play Channel Estimation with Consistency Models for MIMO Systems

This paper proposes a consistency-model-based channel estimation algorithm for multiple-input multiple-output (MIMO) systems. The proposed algorithm employs a consistency model (CM) to learn the angle-domain channel distribution and uses the trained CM as a plug-and-play (PnP) generative prior for MIMO channel estimation. The proposed algorithm alternates between a pilot-observation-based data-consistency update and a CM-prior-based denoising update. In addition, the proposed algorithm adaptively selects the penalty parameter according to residual energy and residual whiteness, and adjusts the CM denoising level according to the observed signal-to-noise ratio (SNR), thereby avoiding the performance degradation caused by fixed inference schedules under varying observation conditions. Simulation results show that the proposed algorithm not only reduces the number of inference steps by 50%--90, but also achieves high estimation accuracy and favorable cross-dataset performance.

eess.SP↗

From Flat to Hierarchical: Evolving Tree-structured Thoughts for Fine-grained Alpha Mining

Alpha mining, aimed at discovering predictive return signals, is typically formulated as symbolic regression. Traditional symbolic methods suffer from search inefficiency and biased prior knowledge. Recently, Large Language Models (LLMs) have emerged as a promising alternative, automatically generating textual thoughts and executable codes to achieve both efficient and interpretable alpha mining. However, existing approaches mostly focus on leveraging LLM's reasoning and reflection capabilities, yet largely neglect the positional bias due to the flat thought representation which restricts efficiency and diversity of the search process. This paper introduces Tree-structured thought Evolution (TreEvo), which evolves hierarchically decomposed thoughts to expand the effective search space. In addition, we propose a set of evolutionary operators tailored to structured thoughts. Experiments on four real-market datasets demonstrate that TreEvo not only obtains competitive alphas with traditional methods in up to 200 times fewer evaluations, but also consistently outperforms LLM-driven EAs across all datasets by $14.31\%$ on average.

cs.CE↗

StepAudio 2.5 Technical Report

Unified audio-language modeling has emerged as a prominent trend in modern speech systems, promising to bring the reasoning capabilities of large language models to auditory tasks. However, existing unified foundations often struggle to match the depth of specialized systems across automatic speech recognition (ASR), text-to-speech synthesis (TTS), and realtime spoken interaction. Bridging this gap remains an open challenge. This report presents StepAudio 2.5, a unified audio-language foundation model that matches or exceeds specialized systems across all three capabilities. Rather than treating these tasks as architecturally distinct, we operate on the premise that once text and audio share a multimodal representational space, task specialization becomes a matter of operational regimes: data construction, optimization targets, and decoding constraints. Guided by this insight, we advance the post-training paradigm from standard supervised learning to task-tailored Reinforcement Learning from Human Feedback (RLHF), using it as the primary mechanism to define complex optimization targets. We leverage this RLHF-centric alignment, alongside specialized decoding, to shape a shared backbone into three distinct operational modes. Concretely, the ASR branch advances transcription efficiency via verifiable multi-token decoding; the TTS branch achieves controllable, expressive synthesis through preference-based RLHF and context-rich supervision; and the Realtime branch realizes low-latency, persona-consistent dialogue via generative reward modeling within an RLHF framework. On standard benchmarks, StepAudio 2.5 achieves state-of-the-art results across ASR, TTS, and Realtime, demonstrating that a singular audio-language foundation can successfully internalize the distinct deployment objectives of speech understanding, generation, and live interaction.

eess.AS↗

Omni-scale Learning-based Sequential Decision Framework for Order Fulfillment of Tote-handling Robotic Systems

Driven by the rapid expansion of e-commerce and small-batch production, the size of the intralogistics load unit of finished goods, semi-finished goods and raw materials is steadily shrinking. Totes are gradually replacing pallets as the primary handling and storage container. This shift has propelled tote-handling robotic systems to the forefront of automation order fulfillment centers. The order-fulfillment decisions of tote-handling robotic systems share a common order-tote-robot sequential decision-making nature. Existing studies primarily focus on decision mechanisms tailored to particular systems, making it difficult to generalize or transfer them to other contexts. We propose an Omni-scale Learning-based Sequential Decision Framework for Order Fulfillment of Tote-handling Robotic Systems (OLSF-TRS), a generalized and scalable sequential decision framework that combines structured combinatorial optimization with multi-agent reinforcement learning to coordinate order,tote, and robot decisions. On small-scale tote-handling robotic systems, OLSF-TRS achieves near-optimal performance with average optimality gaps below 3.5% across two distinct system configurations. In large-scale scenarios, OLSF-TRS consistently outperforms heuristic baselines across two different system types, reducing total tote movements by 8-12% and over 30% compared to SOTA rule-based approaches, while maintaining real-time responsiveness. These improvements translate into tangible operational benefits, including cost reduction, lower energy consumption, and enhanced throughput stability. The proposed framework delivers an efficient and unified order fulfillment decision-making framework for widely deployed tote-handling robotic systems,supporting high-quality order fulfillment in both e-commerce and industrial logistics sectors.

cs.RO↗

SynerDiff: Synergetic Continuous Batching for Fast and Parallel Diffusion Model Inference

The expansion of Artificial Intelligence-generated content service requires diffusion model serving to simultaneously achieve high throughput and low task end-to-end (E2E) latency. However, existing continuous batching methods suffer from severe resource contention during UNet-VAE concurrency, leading to latency spikes. Furthermore, concurrent multi-task scheduling entails a trade-off between UNet throughput and VAE latency across varying scheduling strategies. To address these, we propose SynerDiff, an efficient continuous batching system built on intra-inter level synergy. At the intra-concurrency level, SynerDiff alleviates resource contention by pruning component-specific resource bottlenecks via VAE Chunking and Adaptive Skip-CFG. At the inter-concurrency level, leveraging components' differential sensitivity to scheduling granularities, a threshold-aware scheduler plans concurrent sequences and tunes intra-concurrency decisions to minimize VAE latency while maintaining UNet within high-throughput threshold. Additionally, a feedback controller dynamically adjusts this threshold based on queue loads to boost system capacity ceiling. Experimental results show that, SynerDiff improves throughput by 1.6$\times$ and decreases both average E2E and P99 tail latencies by up to 78.7\%, compared to benchmarks while guaranteeing high image fidelity.

cs.AI↗

Accelerating Multi-Condition T2I Generation via Adaptive Condition Offloading and Pruning

Text-to-image (T2I) generation using multiple conditions enables fine-grained user control on the generated image. Yet, incorporating multi-condition inputs incurs substantial computation and communication overhead, due to additional preprocessing subtasks and control optimizations. It hence leads to unacceptable generation latency. In this paper, we propose an end-edge collaborative system design to accelerate multi-condition T2I generation through adaptive condition offloading and pruning. Extensive offline profiling reveal that, different conditions exhibit significant diversity in computation and communication costs. To this end, we propose a \textit{Subtask Manager} that jointly optimizes condition inference offloading and bandwidth allocation using a heuristic algorithm, balancing local and edge execution delays to minimize overall preprocessing latency. Then, we design a lightweight feature-driven \textit{Conditioning Scale Estimator} that evaluates the contribution of each condition by analyzing its feature activation strength and overlap with other conditions. This allows adaptive conditioning scale selection and pruning of insignificant conditions, thereby accelerating the denoising process. Extensive experimental results show that our system reduces latency by nearly 25\% and improves 6\% average generation quality, outperforming other benchmarks.

cs.MM↗

Externally Controlled Trials: A Review of Design and Borrowing Through a Causal Lens

Externally controlled trials (ECTs) are increasingly used when randomized controls are infeasible, unethical, or insufficient, including applications in rare diseases, oncology, pediatrics, and post-approval effectiveness research. Although methodological work has expanded rapidly across causal inference, Bayesian dynamic borrowing, and hybrid trial designs, the literature remains fragmented. We adopt a six-step scientific roadmap to organize modern ECT methodology in two primary settings: (i) single-arm trials that evaluate efficacy through comparison with external controls, and (ii) hybrid controlled trials that augment the internal control arm with external controls drawn from real-world data or historical studies. The roadmap clarifies causal estimands, identifiability assumptions, and how statistical parameters arise from identification, and shows how modeling and borrowing strategies trade off efficiency and robustness, especially under covariate shift and outcome drift. Within this framework, we synthesize and evaluate recent Bayesian and frequentist developments, compare their strengths, limitations, operating characteristics, and available software, and emphasize the role of sensitivity analysis. By re-framing ECT methodology through a causal lens, this work establishes a coherent foundation for integrating external data into regulatory and clinical decision-making and highlights core challenges and opportunities for future research.

stat.ME↗

EvoMarket: A High-Fidelity and Scalable Financial Market Simulator

High-fidelity, scalable market simulation is a key instrument for mechanism evaluation, stress testing, and counterfactual policy analysis. Yet existing simulators rarely achieve \emph{mechanism fidelity} beyond single-asset intraday settings, \emph{microstructure fidelity} against historical limit order books (LOB), and \emph{computational tractability} at market scale in a single system. This paper presents \textit{EvoMarket}, a discrete-event, multi-agent financial market simulator designed for intervention-oriented experiments in multi-asset and cross-day environments. EvoMarket couples a high-throughput execution core (optimized LOB data structures, hierarchical scheduling under propagation delays, and asynchronous per-asset matching) with explicit institutional mechanisms (market calendars, opening call auctions, price limits, and T+1 settlement). To avoid expensive black-box calibration, EvoMarket introduces an Oracle-guided in-run self-calibration mechanism that interprets microstructure discrepancy as missing order flow and synthesizes corrective orders at recording checkpoints. Experiments on China A-share order-flow and LOB data show close replay alignment over five trading days, fidelity gains from budgeted in-run calibration across depth levels, broad agent order-space coverage, and scalable performance under increasing input order rates and market breadth. We further demonstrate cross-asset linkage and event-study style intervention evaluation that produces structured dependence and interpretable event-time responses.

cs.CE↗

Data-Driven Distributed Stability Certification for Power Systems via Input-State Trajectories

This article proposes a data-driven framework to verify the distributed conditions that guarantee the system-wide stability for interconnected power systems. To guarantee system wide stability, the dynamics of each bus are required to satisfy an output differential passivity (ODP) condition with a sufficient index. These ODP indices uniformly quantify the impacts on the system-wide stability of individual bus dynamics and the coupling strength from the power network. To obtain these indices without explicit physical models, we derive a data-driven linear matrix inequality (LMI) criterion based exclusively on measured input-state trajectories. Furthermore, extracting the optimal ODP index is formulated as a convex semi-definite programming (SDP) problem. Simulations verify the effectiveness of the proposed method under both single-device offline evaluation and system-wide online certification scenarios.

eess.SY↗

Multimodal Large Language Model Enabled Robust Beamforming for HAP Downlink Communications

Small changes in high altitude platform (HAP) attitude can cause significant deviations in HAP downlink beam directions, thereby severely degrading HAP downlink communication performance. In this paper, we develop a multimodal large language model (LLM) enabled beamforming framework to achieve robust HAP downlink communications.Specifically, we design a vision-language LLM (VL-LLM) that learns from multivariate flight telemetry to forecast short-term HAP attitudes under platform shaking and support delay-aware proactive beam steering.We design an offline forecast-error calibration procedure to obtain upper bounds on forecast errors and improve the reliability of proactive analog beam steering.Based on the attitude forecasts, we proactively update the analog beamformer and propose a QoS-driven beamforming and admission method with a lightweight feasibility-enforcement step to satisfy instantaneous transmit-power and QoS requirements.Simulation results indicate that the designed VL-LLM can accurately capture changes in the HAP attitude and the proposed beamforming method achieves a 22.1% higher user service ratio and a 12.5% higher sum-rate than representative baselines.The measured mean and p99 total latencies are 36.24 ms and 40.13 ms, respectively, supporting practical delay-aware deployment.

cs.NI↗

Generative AI Agent Empowered Power Allocation for HAP Propulsion and Communication Systems

High altitude platforms (HAPs) are emerging as a key enabler for 6G coverage, yet limited energy must be split between propulsion and communications. Most prior HAP studies ignore propulsion power or rely on surrogates that miss hull-propeller interference, leading to misestimated communication power budgets and degraded beamforming. More importantly, HAP power allocation is intrinsically a multi-system, multidisciplinary problem in which aerodynamics, propulsion-system efficiency, and communication-system performance (quality of service (QoS) and energy efficiency (EE)) are tightly coupled.To address these challenges, this paper designs an interactive generative artificial intelligence (AI)-empowered HAP power allocation agent.By interacting with the AI agent, we develop an accurate propulsion power consumption model that takes into account the aerodynamic interference between the HAP's hull and the propeller. Assisted by the AI agent, we further formulate a HAP beamforming problem to improve user QoS and enhance the EE of the HAP communication system.This paper also proposes a QoS-enhanced energy-efficient (Q3E) beamforming algorithm to solve the formulated problem.Simulation results demonstrate the accuracy of the propulsion-power model and the effectiveness of the Q3E algorithm.

cs.NI↗

Detecting Malicious Intents in Smart Contracts with Pre-trained Programming Language Models

Malicious developer intents in smart contracts constitute significant security threats to decentralized applications, leading to substantial economic losses. Prior work introduced SmartIntentNN, a deep learning model for detecting unsafe developer intents. By combining the Universal Sentence Encoder, a K-means clustering-based intent highlighting mechanism, and a Bidirectional Long Short-Term Memory (BiLSTM) network, the model achieved an F1 score of 0.8633 on an evaluation set of 10,000 real-world smart contracts across ten distinct intent categories. This paper presents SmartIntentV2 (Smart Contract Intent Neural Network Version 2). The primary enhancement is the integration of a BERT-based pre-trained programming language model, which we domain-adaptively pre-train on a dataset of 16,000 real-world smart contracts using a Masked Language Modeling objective. SmartIntentV2 retains the BiLSTM-based multi-label classification network for intent detection. On the same evaluation set of 10,000 smart contracts, it achieves superior performance with an accuracy of 0.9789, precision of 0.9090, recall of 0.9476, and an F1 score of 0.9279, substantially outperforming its predecessor and other baseline models. Notably, SmartIntentV2 also delivers a 65.5% relative improvement in F1 score over GPT-4.1 on this specialized task. These results establish SmartIntentV2 as a new state-of-the-art model for smart contract intent detection.

cs.SE↗