SearcharxivSearch

arXiv subjects

Shuhan Zhang

Publications and source records attributed to Shuhan Zhang.

At least 19 recordsLinked to original sources

Uniform-in-time strong convergence rates of fully discrete approximations for stochastic Cahn--Hilliard equations with multiplicative noise

This paper investigates the uniform-in-time strong convergence rates of a fully discrete approximation for the stochastic Cahn--Hilliard equation driven by multiplicative noise in spatial dimensions $d\in\{1,2,3\}$. The proposed scheme combines a spectral Galerkin method in space with a backward Euler scheme in time. The main analytical difficulties arise from the state-dependent stochastic perturbation, the absence of a global monotonicity structure for the nonlinear term, and the fourth-order nature of the Cahn--Hilliard operator. In particular, these features make the derivation of uniform $L^{\infty}$-moment estimates highly nontrivial in three dimensions. For the continuous equation, by utilizing the It\^{o} formula to $\|u\|^p$ and introducing the energy functional $\mathcal{E}(u(t))$, we derive the uniform moment boundedness of the solution. At the fully discrete level, we develop discrete energy estimates and close the required high-order moment bounds through an induction argument. Based on these regularity estimates, we deduce uniform-in-time strong convergence rates for the fully discrete scheme. Moreover, we prove the existence and uniqueness of invariant measures for both the exact dynamics and the fully discrete numerical dynamics. Numerical experiments are provided to confirm the theoretical findings.

math.NA

WeSep: A Modular and Cue-Composable Framework for Target Speaker Extraction

The study of Target Speaker Extraction (TSE) aims to isolate a desired speaker from overlapping speech mixture given auxiliary cues. Existing systems are typically designed for specific cue types, limiting flexibility when cue availability varies across scenarios. We present WeSep, a unified framework that reformulates TSE as a heterogeneous cue-conditioned learning problem. In WeSep, cue modules and separator backbones are decoupled through standardized interfaces, enabling configurable cue injection and flexible integration of diverse modalities. The design enables systematic study of cue structure, intra- and cross-modal interaction, and dynamic cue availability within a shared optimization framework, facilitating adaptation to real-world conditions. Experiments across enrollment, spatial, visual, and textual cues reveal modality-dependent characteristics and demonstrate stable optimization under heterogeneous cue availability. The toolkit will be publicly available.

eess.AS

The Western Jet of SS 433/W50: Hard X-ray Emission, Spectral Evolution, and Comparison to the Eastern Jet

The W50 nebula powered by the microquasar SS 433 is a unique laboratory for exploring several fundamental astrophysical phenomena. This study presents observations from NuSTAR and XMM-Newton, concentrating on the western lobe of W50. Detection of hard non-thermal X-ray emission is reported, extending up to approximately 30 keV. This emission originates from a compact, knotty area referred to as the "Head", located at approximately 17 arcmin (equivalent to 26.5 pc at an assumed distance of 5.5 kpc) to the west of SS 433, and characterized by a power-law spectrum with a hard photon index of 1.55 +/- 0.07 (0.5-30 keV). Moving westward from SS 433, the photon index gradually steepens, ultimately reaching a photon index of 2.10 +/- 0.05 in the "w2" region centered at approximately 35 arcmin or approximately 56 pc from SS 433. The distinct hard X-ray knots observed serve as clear markers for sites of particle acceleration within the western jet. The synchrotron radiation from the "Head" region implies equipartition magnetic field strength B of approximately 15 microG. Notably, these properties (western "Head" location, unusually hard spectral index, inferred magnetic field, and spectral evolution away from SS 433) are very similar to what has been observed in the eastern lobe, supporting a symmetric jet-driven origin. Finally, the broadband spectral energy distribution (SED) and X-ray morphology are modeled using semi-analytic jet models, exploring different jet velocity and magnetic field configurations. The results favor a scenario in which in-situ particle acceleration and synchrotron emission dominate, with implications for understanding particle transport, jet dynamics, and W50's role as a Galactic PeVatron.

astro-ph.HE

Quantum Process Realization of LDPC Code Dualities and Product Constructions

We realize a broad class of code constructions, including Kramers-Wannier duality, tensor product, and check product, as quantum processes consisting of ancilla initialization, local unitaries, and projective measurements. Using ZX-calculus, we represent these transformations diagrammatically and provide a systematic algorithm for extracting quantum circuits. Central to our framework is the observation that the physical content of a classical LDPC code is captured by the operator algebra associated with its Tanner graph, and that code transformations correspond to maps between such algebras. Kramers-Wannier duality then admits a natural interpretation as gauging, while tensor and check products correspond to coupled-layer constructions in which interlayer coupling and projection implement a quotient on stacked operator algebras. Together, these results establish a unified framework connecting code transformations, quantum circuits, and mappings between distinct quantum phases of matter.

quant-ph

AlphaFlowTSE: One-Step Generative Target Speaker Extraction via Conditional AlphaFlow

In target speaker extraction (TSE), we aim to recover target speech from a multi-talker mixture using a short enrollment utterance as reference. Recent studies on diffusion and flow-matching generators have improved target-speech fidelity. However, multi-step sampling increases latency, and one-step solutions often rely on a mixture-dependent time coordinate that can be unreliable for real-world conversations. We present AlphaFlowTSE, a one-step conditional generative model trained with a Jacobian-vector product (JVP)-free AlphaFlow objective. AlphaFlowTSE learns mean-velocity transport along a mixture-to-target trajectory starting from the observed mixture, eliminating auxiliary mixing-ratio prediction, and stabilizes training by combining flow matching with an interval-consistency teacher-student target. Experiments on Libri2Mix and REAL-T confirm that AlphaFlowTSE improves target-speaker similarity and real-mixture generalization for downstream automatic speech recognition (ASR).

cs.SD

Wasserstein Proximal Policy Gradient

We study policy gradient methods for continuous-action, entropy-regularized reinforcement learning through the lens of Wasserstein geometry. Starting from a Wasserstein proximal update, we derive Wasserstein Proximal Policy Gradient (WPPG) via an operator-splitting scheme that alternates an optimal transport update with a heat step implemented by Gaussian convolution. This formulation avoids evaluating the policy's log density or its gradient, making the method directly applicable to expressive implicit stochastic policies specified as pushforward maps. We establish a global linear convergence rate for WPPG, covering both exact policy evaluation and actor-critic implementations with controlled approximation error. Empirically, WPPG is simple to implement and attains competitive performance on standard continuous-control benchmarks.

cs.LG

DeepHalo: A Neural Choice Model with Controllable Context Effects

Modeling human decision-making is central to applications such as recommendation, preference learning, and human-AI alignment. While many classic models assume context-independent choice behavior, a large body of behavioral research shows that preferences are often influenced by the composition of the choice set itself -- a phenomenon known as the context effect or Halo effect. These effects can manifest as pairwise (first-order) or even higher-order interactions among the available alternatives. Recent models that attempt to capture such effects either focus on the featureless setting or, in the feature-based setting, rely on restrictive interaction structures or entangle interactions across all orders, which limits interpretability. In this work, we propose DeepHalo, a neural modeling framework that incorporates features while enabling explicit control over interaction order and principled interpretation of context effects. Our model enables systematic identification of interaction effects by order and serves as a universal approximator of context-dependent choice functions when specialized to a featureless setting. Experiments on synthetic and real-world datasets demonstrate strong predictive performance while providing greater transparency into the drivers of choice.

cs.LG

Relative Localization System Design for SnailBot: A Modular Self-reconfigurable Robot

This paper presents the design and implementation of a relative localization system for SnailBot, a modular self reconfigurable robot. The system integrates ArUco marker recognition, optical flow analysis, and IMU data processing into a unified fusion framework, enabling robust and accurate relative positioning for collaborative robotic tasks. Experimental validation demonstrates the effectiveness of the system in realtime operation, with a rule based fusion strategy ensuring reliability across dynamic scenarios. The results highlight the potential for scalable deployment in modular robotic systems.

cs.RO

CAD-Coder:Text-Guided CAD Files Code Generation

Computer-aided design (CAD) is a way to digitally create 2D drawings and 3D models of real-world products. Traditional CAD typically relies on hand-drawing by experts or modifications of existing library files, which doesn't allow for rapid personalization. With the emergence of generative artificial intelligence, convenient and efficient personalized CAD generation has become possible. However, existing generative methods typically produce outputs that lack interactive editability and geometric annotations, limiting their practical applications in manufacturing. To enable interactive generative CAD, we propose CAD-Coder, a framework that transforms natural language instructions into CAD script codes, which can be executed in Python environments to generate human-editable CAD files (.Dxf). To facilitate the generation of editable CAD sketches with annotation information, we construct a comprehensive dataset comprising 29,130 Dxf files with their corresponding script codes, where each sketch preserves both editability and geometric annotations. We evaluate CAD-Coder on various 2D/3D CAD generation tasks against existing methods, demonstrating superior interactive capabilities while uniquely providing editable sketches with geometric annotations.

cs.GR

Holographic Classical Shadow Tomography

We introduce "holographic shadows", a new class of randomized measurement schemes for classical shadow tomography that achieves the optimal scaling of sample complexity for learning geometrically local Pauli operators at any length scale, without the need for fine-tuning protocol parameters such as circuit depth or measurement rate. Our approach utilizes hierarchical quantum circuits, such as tree quantum circuits or holographic random tensor networks. Measurements within the holographic bulk correspond to measurements at different scales on the boundary (i.e. the physical system of interests), facilitating efficient quantum state estimation across observable at all scales. Considering the task of estimating string-like Pauli observables supported on contiguous intervals of $k$ sites in a 1D system, our method achieves an optimal sample complexity scaling of $\sim d^k\mathrm{poly}(k)$, with $d$ the local Hilbert space dimension. We present a holographic minimal cut framework to demonstrate the universality of this sample complexity scaling and validate it with numerical simulations, illustrating the efficacy of holographic shadows in enhancing quantum state learning capabilities.

quant-ph

Hard X-ray emission from the eastern jet of SS 433 powering the W50 `Manatee' nebula: Evidence for particle re-acceleration

We present a broadband X-ray study of W50 (`the Manatee nebula'), the complex region powered by the microquasar SS 433, that provides a test-bed for several important astrophysical processes. The W50 nebula, a Galactic PeVatron candidate, is classified as a supernova remnant but has an unusual double-lobed morphology likely associated with the jets from SS 433. Using NuSTAR, XMM-Newton, and Chandra observations of the inner eastern lobe of W50, we have detected hard non-thermal X-ray emission up to $\sim$30 keV, originating from a few-arcminute size knotty region (`Head') located $\lesssim$ 18$^{\prime}$ (29 pc for a distance of 5.5 kpc) east of SS 433, and constrain its photon index to 1.58$\pm$0.05 (0.5-30 keV band). The index gradually steepens eastward out to the radio `ear' where thermal soft X-ray emission with a temperature $kT$$\sim$0.2 keV dominates. The hard X-ray knots mark the location of acceleration sites within the jet and require an equipartition magnetic field of the order of $\gtrsim$12$μ$G. The unusually hard spectral index from the `Head' region challenges classical particle acceleration processes and points to particle injection and re-acceleration in the sub-relativistic SS 433 jet, as seen in blazars and pulsar wind nebulae.

astro-ph.HE

ADEPT: Automatic Differentiable DEsign of Photonic Tensor Cores

Photonic tensor cores (PTCs) are essential building blocks for optical artificial intelligence (AI) accelerators based on programmable photonic integrated circuits. PTCs can achieve ultra-fast and efficient tensor operations for neural network (NN) acceleration. Current PTC designs are either manually constructed or based on matrix decomposition theory, which lacks the adaptability to meet various hardware constraints and device specifications. To our best knowledge, automatic PTC design methodology is still unexplored. It will be promising to move beyond the manual design paradigm and "nurture" photonic neurocomputing with AI and design automation. Therefore, in this work, for the first time, we propose a fully differentiable framework, dubbed ADEPT, that can efficiently search PTC designs adaptive to various circuit footprint constraints and foundry PDKs. Extensive experiments show superior flexibility and effectiveness of the proposed ADEPT framework to explore a large PTC design space. On various NN models and benchmarks, our searched PTC topology outperforms prior manually-designed structures with competitive matrix representability, 2-30x higher footprint compactness, and better noise robustness, demonstrating a new paradigm in photonic neural chip design. The code of ADEPT is available at https://github.com/JeremieMelo/ADEPT using the https://github.com/JeremieMelo/pytorch-onn (TorchONN) library.

cs.ET

Ultra Light OCR Competition Technical Report

Ultra Light OCR Competition is a Chinese scene text recognition competition jointly organized by CSIG (China Society of Image and Graphics) and Baidu, Inc. In addition to focusing on common problems in Chinese scene text recognition, such as long text length and massive characters, we need to balance the trade-off of model scale and accuracy since the model size limitation in the competition is 10M. From experiments in aspects of data, model, training, etc, we proposed a general and effective method for Chinese scene text recognition, which got us second place among over 100 teams with accuracy 0.817 in TestB dataset. The code is available at https://aistudio.baidu.com/aistudio/projectdetail/2159102.

cs.CV

LinEasyBO: Scalable Bayesian Optimization Approach for Analog Circuit Synthesis via One-Dimensional Subspaces

A large body of literature has proved that the Bayesian optimization framework is especially efficient and effective in analog circuit synthesis. However, most of the previous research works only focus on designing informative surrogate models or efficient acquisition functions. Even if searching for the global optimum over the acquisition function surface is itself a difficult task, it has been largely ignored. In this paper, we propose a fast and robust Bayesian optimization approach via one-dimensional subspaces for analog circuit synthesis. By solely focusing on optimizing one-dimension subspaces at each iteration, we greatly reduce the computational overhead of the Bayesian optimization framework while safely maximizing the acquisition function. By combining the benefits of different dimension selection strategies, we adaptively balancing between searching globally and locally. By leveraging the batch Bayesian optimization framework, we further accelerate the optimization procedure by making full use of the hardware resources. Experimental results quantitatively show that our proposed algorithm can accelerate the optimization procedure by up to 9x and 38x compared to LP-EI and REMBOpBO respectively when the batch size is 15.

eess.SY

An Efficient Asynchronous Batch Bayesian Optimization Approach for Analog Circuit Synthesis

In this paper, we propose EasyBO, an Efficient ASYnchronous Batch Bayesian Optimization approach for analog circuit synthesis. In this proposed approach, instead of waiting for the slowest simulations in the batch to finish, we accelerate the optimization procedure by asynchronously issuing the next query points whenever there is an idle worker. We introduce a new acquisition function that can better explore the design space for asynchronous batch Bayesian optimization. A new strategy is proposed to better balance the exploration and exploitation and guarantee the diversity of the query points. And a penalization scheme is proposed to further avoid redundant queries during the asynchronous batch optimization. The efficiency of optimization can thus be further improved. Compared with the state-of-the-art batch Bayesian optimization algorithm, EasyBO achieves up to 7.35 times speed-up without sacrificing the optimization results.

eess.SY

An Efficient Batch Constrained Bayesian Optimization Approach for Analog Circuit Synthesis via Multi-objective Acquisition Ensemble

Bayesian optimization is a promising methodology for analog circuit synthesis. However, the sequential nature of the Bayesian optimization framework significantly limits its ability to fully utilize real-world computational resources. In this paper, we propose an efficient parallelizable Bayesian optimization algorithm via Multi-objective ACquisition function Ensemble (MACE) to further accelerate the optimization procedure. By sampling query points from the Pareto front of the probability of improvement (PI), expected improvement (EI) and lower confidence bound (LCB), we combine the benefits of state-of-the-art acquisition functions to achieve a delicate tradeoff between exploration and exploitation for the unconstrained optimization problem. Based on this batch design, we further adjust the algorithm for the constrained optimization problem. By dividing the optimization procedure into two stages and first focusing on finding an initial feasible point, we manage to gain more information about the valid region and can better avoid sampling around the infeasible area. After achieving the first feasible point, we favor the feasible region by adopting a specially designed penalization term to the acquisition function ensemble. The experimental results quantitatively demonstrate that our proposed algorithm can reduce the overall simulation time by up to 74 times compared to differential evolution (DE) for the unconstrained optimization problem when the batch size is 15. For the constrained optimization problem, our proposed algorithm can speed up the optimization process by up to 15 times compared to the weighted expected improvement based Bayesian optimization (WEIBO) approach, when the batch size is 15.

cs.LG

BigCarl: Mining frequent subnets from a single large Petri net

While there have been lots of work studying frequent subgraph mining, very rare publications have discussed frequent subnet mining from more complicated data structures such as Petri nets. This paper studies frequent subnets mining from a single large Petri net. We follow the idea of transforming a Petri net in net graph form and to mine frequent sub-net graphs to avoid high complexity. Technically, we take a minimal traversal approach to produce a canonical label of the big net graph. We adapted the maximal independent embedding set approach to the net graph representation and proposed an incremental pattern growth (independent embedding set reduction) way for discovering frequent sub-net graphs from the single large net graph, which are finally transformed back to frequent subnets. Extensive performance studies made on a single large Petri net, which contains 10K events, 40K conditions and 30 K arcs, showed that our approach is correct and the complexity is reasonable.

cs.DB

PSpan:Mining Frequent Subnets of Petri Nets

This paper proposes for the first time an algorithm PSpan for mining frequent complete subnets from a set of Petri nets. We introduced the concept of complete subnets and the net graph representation. PSpan transforms Petri nets in net graphs and performs sub-net graph mining on them, then transforms the results back to frequent subnets. PSpan follows the pattern growth approach and has similar complexity like gSpan in graph mining. Experiments have been done to confirm PSpan's reliability and complexity. Besides C/E nets, it applies also to a set of other Petri net subclasses.

cs.LG