SearcharxivSearch

arXiv subjects

Xinze Zhang

Publications and source records attributed to Xinze Zhang.

17 recordsLinked to original sources

Most Probable KAM Tori in Stochastic Hamiltonian Systems Driven by Multiplicative Noise

This paper investigates the effect of state-dependent multiplicative noise on the integrable structure of Hamiltonian systems. Under a local high-dimensional Lamperti-type condition, we derive the Onsager--Machlup functional in the corresponding flattening coordinates. When the diffusion vector fields are Hamiltonian, the resulting action is minimized uniquely by the deterministic Hamiltonian trajectory. Combining this characterization with a finitely differentiable KAM theorem, we prove that, under suitable assumptions, the invariant tori with Diophantine frequencies of an integrable Hamiltonian system persist, in the sense of most probable paths, under a small deterministic perturbation and multiplicative noise. Furthermore, as the intensity of the multiplicative noise tends to zero, we establish a localized large deviation principle that characterizes the asymptotic probability of solution trajectories deviating from these invariant tori.

math.DS

URDF Synthesis from RGB-D Sequences via Differentiable Joint Inference and Energy-Consistent Verification

Reconstructing simulation-ready digital twins of articulated objects from sensor observations remains constrained by two persistent gaps: (i) part-level geometric reconstruction is decoupled from kinematic-parameter estimation, and (ii) the recovered models often violate basic dynamic invariants such as energy conservation, leading to drift when the URDF is replayed in physics simulators. We present KinemaForge, a constraint-driven pipeline that jointly infers part-level shape, joint topology, and joint parameters from short RGB-D sequences and validates the result against an energy-consistent verifier built on differentiable rigid-body dynamics. The pipeline introduces three components: a kinematic constraint graph that encodes joint-part incidences as soft edges; a differentiable screw-axis solver that backpropagates from rendered observations through Featherstone's articulated-body algorithm to joint parameters; and an energy residual loss that penalises non-physical free responses of the reconstructed model. Across five PartNet-Mobility categories and an internal RGB-D benchmark, KinemaForge reduces the average joint-axis error from 4.52 degrees to 2.83 degrees (-37.4%) over the strongest geometric baseline (PARIS) and from 5.30 degrees to 2.83 degrees (-46.6%) over the interaction-based Ditto baseline, lowers long-horizon simulation drift by 64% (vs. PARIS) over 50 s rollouts, and yields URDFs whose closed-loop manipulation success rate improves by 14.6 percentage points over Ditto in our preliminary evaluation. Code and reconstruction data will be released upon acceptance.

cs.CV

Bridging Geographic Bias in Urban Streetscape Inference via Lifelong Learning with Visual-Semantic Pivoting

Visual perception of urban streetscapes underpins evidence-based decisions in landscape planning, public health, and place-making. Yet models trained on a few well-photographed metropolises systematically misjudge underrepresented districts, propagating geographic bias into downstream policy. We address this gap with HVSP-LL, a lifelong learning framework that couples a stratified visual-semantic pivoting module with an equity-aware rehearsal mechanism. The pivoting module organises landscape concepts along a three-tier ontology (macro structure, meso composition, micro element) and aligns image features to learnable semantic anchors at each tier, providing transferable representations that resist distributional drift. The lifelong adaptation component sequentially absorbs new urban regions while constraining inter-region perception gaps through a worst-region sample-reweighting objective and a structurally-aware exemplar buffer. We evaluate HVSP-LL on a panoramic streetscape benchmark assembled from twelve cities across four continents and seven perceptual dimensions. The framework attains 0.834 Spearman correlation on the held-out city sequence, an absolute 6.1 point improvement over the strongest continual baseline, and shrinks the inter-city perception gap to 0.094 -- a 38% reduction relative to the strongest continual baseline (0.151) and a 57% reduction relative to a representative regularisation baseline (0.218). Ablations confirm that each tier of the pivoting hierarchy contributes monotonically, and the equity-aware rehearsal converts mean backward transfer from -0.038 (without retention) to +0.013, eliminating catastrophic forgetting on the held-out sequence. Our results indicate that hierarchical anchoring is a practical pathway toward geographically equitable streetscape inference at city scale.

cs.CV

Conformal Risk Prediction for Non-Alcoholic Fatty Liver Disease Using Gradient Boosting with Distribution-Free Coverages

Non-alcoholic fatty liver disease (NAFLD) affects roughly 25% of global adults, posing substantial hepatic and cardiovascular risks. Yet, population-level screening tools remain inadequate. We present Method, a machine-learning framework for NAFLD risk prediction coupling gradient-boosted decision trees with conformal prediction to yield calibrated, distribution-free coverage guarantees on individual risk estimates. It integrates a mutual-information-based stability selection procedure to identify a compact, clinically interpretable feature subset via bootstrap resampling, constructing prediction sets whose marginal coverage provably exceeds a user-specified confidence level. We evaluated Method on a multicenter cohort from Guangzhou, China (primary n=2,187; external validation n=412) using 78 candidate features across demographics, metabolic biomarkers, and lifestyle factors. Method achieves an AUROC of 0.912 internally and 0.891 externally, outperforming deep neural networks, TabNet, support vector machines, and logistic regression. Conformal prediction sets achieve 91.3% empirical coverage at the 90% nominal level. A three-tier risk stratification derived from these scores separates the population into distinct groups, with the high-risk subgroup showing a 12-month progression rate 4.7 times that of the low-risk tier. The selected features -- notably waist circumference, ALT, GGT, triglycerides, fasting glucose, and BMI -- align with established metabolic risk factors, providing biological plausibility.

cs.LG

VigilFormer: Deformable Attention for Video Anomaly Detection with Causal Risk Inference

Video anomaly detection in surveillance settings must balance detection accuracy against real-time throughput, a tension that existing methods address either through stronger feature extractors or more efficient architectures, but rarely both. We present VigilFormer, a unified framework that combines deformable spatio-temporal attention with causal temporal modeling to detect anomalies in untrimmed surveillance video. The proposed Deformable Spatio-Temporal Encoder (DSTE) attends to a sparse set of informative locations across frames, avoiding the quadratic cost of dense attention while retaining the ability to capture irregular motion patterns. A Causal Anomaly Classifier (CAC) applies dilated causal convolutions over snippet-level features and optimizes a contrastive multiple-instance learning objective that separates anomalous and normal representations without frame-level labels. To meet deployment constraints, an Adaptive Confidence Scheduler (ACS) dynamically skips low-information frames at inference time, reducing redundant computation in static scenes. Evaluated on UCF-Crime, ShanghaiTech, and CUHK Avenue, VigilFormer achieves AUC scores of 87.83%, 97.21%, and 89.74% respectively, at 41.5 FPS on a single GPU, outperforming recent weakly-supervised methods in both accuracy and speed.

cs.CV

Cross-Modal Hierarchical Fusion for from Multi-Sensor Ground Observation

Dense volumetric reconstruction of cloud microphysical fields from sparse ground-based instruments remains an open problem, largely because the available measurements are heterogeneous in both modality and spatial coverage. We present AtmoFuseNet, a framework that fuses multi-view sky camera imagery with millimeter-wave cloud radar and ceilometer observations to produce 4D (three spatial dimensions plus time) estimates of cloud state and wind. The method operates in three stages: a cross-modal hierarchical aggregation module that combines image feature pyramids with instrument-derived vertical profiles through layer-wise cross-attention; a conditional variational refinement module that maps the resulting volume to physically consistent microphysical fields under differentiable radar and image forward models; and a correlation-based motion estimator that recovers per-voxel 3D wind vectors from consecutive volumetric reconstructions. On collocated observations from a semi-arid site, AtmoFuseNet reaches 0.026 g m^-3 liquid water content MAE and 1.18 m s^-1 wind speed MAE, improving over existing retrieval baselines. Ablation experiments isolate the contribution of each module.

cs.CV

Most Probable KAM Tori in Stochastic Hamiltonian Systems

This paper investigates in depth how stochastic perturbations affect the integrable structure of Hamiltonian systems and develops a KAM theory for stochastic Hamiltonian dynamics, in the sense of the most probable path. We first derive the Onsager-Machlup functional for stochastic Hamiltonian systems driven by time-dependent noise coefficients and identify the most probable path of the system trajectories. Building on this, we establish a large deviation principle and obtain an explicit rate function that quantitatively characterizes trajectory deviations, in particular for rare events. The main contribution of this work is to prove that, under stochastic noise, the original quasi-periodic invariant tori persist in the sense of the most probable path, thereby demonstrating the stability of KAM structures in random environments. Moreover, we show that the Onsager-Machlup functional coincides exactly with the large deviation rate function, thereby providing a quantitative characterization of both the structural persistence of quasi-periodic motions and the geometry of fluctuations in stochastic Hamiltonian systems. Overall, our results extend the classical KAM framework to stochastic settings and offer new insight into the behavior of complex dynamical systems under noise.

math.DS

Temporally Unified Adversarial Perturbations for Time Series Forecasting

While deep learning models have achieved remarkable success in time series forecasting, their vulnerability to adversarial examples remains a critical security concern. However, existing attack methods in the forecasting field typically ignore the temporal consistency inherent in time series data, leading to divergent and contradictory perturbation values for the same timestamp across overlapping samples. This temporally inconsistent perturbations problem renders adversarial attacks impractical for real-world data manipulation. To address this, we introduce Temporally Unified Adversarial Perturbations (TUAPs), which enforce a temporal unification constraint to ensure identical perturbations for each timestamp across all overlapping samples. Moreover, we propose a novel Timestamp-wise Gradient Accumulation Method (TGAM) that provides a modular and efficient approach to effectively generate TUAPs by aggregating local gradient information from overlapping samples. By integrating TGAM with momentum-based attack algorithms, we ensure strict temporal consistency while fully utilizing series-level gradient information to explore the adversarial perturbation space. Comprehensive experiments on three benchmark datasets and four representative state-of-the-art models demonstrate that our proposed method significantly outperforms baselines in both white-box and black-box transfer attack scenarios under TUAP constraints. Moreover, our method also exhibits superior transfer attack performance even without TUAP constraints, demonstrating its effectiveness and superiority in generating adversarial perturbations for time series forecasting models.

cs.LG

A Policy Gradient-Based Sequence-to-Sequence Method for Time Series Prediction

Sequence-to-sequence architectures built upon recurrent neural networks have become a standard choice for multi-step-ahead time series prediction. In these models, the decoder produces future values conditioned on contextual inputs, typically either actual historical observations (ground truth) or previously generated predictions. During training, feeding ground-truth values helps stabilize learning but creates a mismatch between training and inference conditions, known as exposure bias, since such true values are inaccessible during real-world deployment. On the other hand, using the model's own outputs as inputs at test time often causes errors to compound rapidly across prediction steps. To mitigate these limitations, we introduce a new training paradigm grounded in reinforcement learning: a policy gradient-based method to learn an adaptive input selection strategy for sequence-to-sequence prediction models. Auxiliary models first synthesize plausible input candidates for the decoder, and a trainable policy network optimized via policy gradients dynamically chooses the most beneficial inputs to maximize long-term prediction performance. Empirical evaluations on diverse time series datasets confirm that our approach enhances both accuracy and stability in multi-step forecasting compared to conventional methods.

cs.LG

Rethinking Recurrent Neural Networks for Time Series Forecasting: A Reinforced Recurrent Encoder with Prediction-Oriented Proximal Policy Optimization

Time series forecasting plays a crucial role in contemporary engineering information systems for supporting decision-making across various industries, where Recurrent Neural Networks (RNNs) have been widely adopted due to their capability in modeling sequential data. Conventional RNN-based predictors adopt an encoder-only strategy with sliding historical windows as inputs to forecast future values. However, this approach treats all time steps and hidden states equally without considering their distinct contributions to forecasting, leading to suboptimal performance. To address this limitation, we propose a novel Reinforced Recurrent Encoder with Prediction-oriented Proximal Policy Optimization, RRE-PPO4Pred, which significantly improves time series modeling capacity and forecasting accuracy of the RNN models. The core innovations of this method are: (1) A novel Reinforced Recurrent Encoder (RRE) framework that enhances RNNs by formulating their internal adaptation as a Markov Decision Process, creating a unified decision environment capable of learning input feature selection, hidden skip connection, and output target selection; (2) An improved Prediction-oriented Proximal Policy Optimization algorithm, termed PPO4Pred, which is equipped with a Transformer-based agent for temporal reasoning and develops a dynamic transition sampling strategy to enhance sampling efficiency; (3) A co-evolutionary optimization paradigm to facilitate the learning of the RNN predictor and the policy agent, providing adaptive and interactive time series modeling. Comprehensive evaluations on five real-world datasets indicate that our method consistently outperforms existing baselines, and attains accuracy better than state-of-the-art Transformer models, thus providing an advanced time series predictor in engineering informatics.

cs.LG

The LLN and CLT for the statistical ensembles of discrete integrable Hamiltonian systems

This paper investigates the behavior of statistical ensembles under iteration map induced by discrete integrable Hamiltonian systems in deterministic case and stochastic case, addressing the problem from two perspectives: the Law of Large Numbers and the Central Limit Theorem. In deterministic case, the Law of Large Numbers simplifies the convergence conditions to the extent that the Riemann-Lebesgue lemma is no longer required. In the stochastic setting, we extend the results to general stochastic processes, beginning with the perturbation term represented by standard Brownian motion. Moreover, we establish a Central Limit Theorem for the statistical ensemble. A numerical example is also included.

math.PR

Persistence of Invariant Tori for Stochastic Nonlinear Schrödinger in the Sense of Most Probable Paths

This paper investigates the application of KAM theory to the stochastic nonlinear Schrödinger equation on infinite lattices, focusing on the stability of low-dimensional invariant tori in the sense of most probable paths. For generality, we provide an abstract proof within the framework of stochastic Hamiltonian systems on infinite lattices. We begin by constructing the Onsager-Machlup functional for these systems in a weighted infinite sequence space. Using the Euler-Lagrange equation, we identify the most probable transition path of the system's trajectory under stochastic perturbations. Additionally, we establish a large deviation principle for the system and derive a rate function that quantifies the deviation of the system's trajectory from the most probable path, especially in rare events. Combining this with classical KAM theory for the nonlinear Schrödinger equation, we demonstrate the persistence of low-dimensional invariant tori under small deterministic and stochastic perturbations. Furthermore, we prove that the probability of the system's trajectory deviating from these tori can be described by the derived rate function, providing a new probabilistic framework for understanding the stability of stochastic Hamiltonian systems on infinite lattices.

math.DS

Onsager-Machlup functional for stochastic differential equations with time-varying noise

This paper is devoted to studying the Onsager-Machlup functional for stochastic differential equations with time-varying noise of the α-Hölder, 0<α<1/4, dXt =f(t,Xt)dt+g(t)dWt. Our study focuses on scenarios where the diffusion coefficient g(t) exhibits temporal variability, starkly contrasting the conventional assumption of a constant diffusion coefficient in the existing literature. This variance brings some complexity to the analysis. Through this investigation, we derive the Onsager-Machlup functional, which acts as the Lagrangian for mapping the most probable transition path between metastable states in stochastic processes affected by time-varying noise. This is done by introducing new measurable norms and applying an appropriate version of the Girsanov transformation. To illustrate our theoretical advancements, we provide numerical simulations, including cases of a one-dimensional SDE and a fast-slow SDE system, which demonstrate the application to multiscale stochastic volatility models, thereby highlighting the significant impact of time-varying diffusion coefficients.

math.PR

Poisson stability of solutions for stochastic evolution equations driven by fractional Brownian motion

In this paper, we study the problem of Poisson stability of solutions for stochastic semi-linear evolution equation driven by fractional Brownian motion \mathrm{d} X(t)= \left( AX(t) + f(t, X(t)) \right) \mathrm{d}t + g\left(t, X(t)\right)\mathrm{d}B^H_{Q}(t), where A is an exponentially stable linear operator acting on a separable Hilbert space \mathbb{H}, coefficients f and g are Poisson stable in time, and B^H_Q (t) is a Q-cylindrical fBm with Hurst index H. First, we establish the existence and uniqueness of the solution for this equation. Then, we prove that under the condition where the functions f and g are sufficiently "small", the equation admits a solution that exhibits the same character of recurrence as f and g. The discussion is further extended to the asymptotic stability of these Poisson stable solutions. Finally, we include an example to validate our results.

math.DS

Feedback-based Modal Mutual Search for Attacking Vision-Language Pre-training Models

Although vision-language pre-training (VLP) models have achieved remarkable progress on cross-modal tasks, they remain vulnerable to adversarial attacks. Using data augmentation and cross-modal interactions to generate transferable adversarial examples on surrogate models, transfer-based black-box attacks have become the mainstream methods in attacking VLP models, as they are more practical in real-world scenarios. However, their transferability may be limited due to the differences on feature representation across different models. To this end, we propose a new attack paradigm called Feedback-based Modal Mutual Search (FMMS). FMMS introduces a novel modal mutual loss (MML), aiming to push away the matched image-text pairs while randomly drawing mismatched pairs closer in feature space, guiding the update directions of the adversarial examples. Additionally, FMMS leverages the target model feedback to iteratively refine adversarial examples, driving them into the adversarial region. To our knowledge, this is the first work to exploit target model feedback to explore multi-modality adversarial boundaries. Extensive empirical evaluations on Flickr30K and MSCOCO datasets for image-text matching tasks show that FMMS significantly outperforms the state-of-the-art baselines.

cs.CV

Onsager-Machlup functional for stochastic lattice dynamical systems driven by time-varying noise

This paper investigates the Onsager-Machlup functional of stochastic lattice dynamical systems (SLDSs) driven by time-varying noise. We extend the Onsager-Machlup functional from finite-dimensional to infinite-dimensional systems, and from constant to time-varying diffusion coefficients. We first verify the existence and uniqueness of general SLDS solutions in the infinite sequence weighted space $l^2_ρ$. Building on this foundation, we employ techniques such as the infinite-dimensional Girsanov transform, Karhunen-Loève expansion, and probability estimation of Brownian motion balls to derive the Onsager-Machlup functionals for SLDSs in $l^2_ρ$ space. Additionally, we use a numerical example to illustrate our theoretical findings, based on the Euler Lagrange equation corresponding to the Onsage Machup functional.

math.PR

Error-feedback stochastic modeling strategy for time series forecasting with convolutional neural networks

Despite the superiority of convolutional neural networks demonstrated in time series modeling and forecasting, it has not been fully explored on the design of the neural network architecture and the tuning of the hyper-parameters. Inspired by the incremental construction strategy for building a random multilayer perceptron, we propose a novel Error-feedback Stochastic Modeling (ESM) strategy to construct a random Convolutional Neural Network (ESM-CNN) for time series forecasting task, which builds the network architecture adaptively. The ESM strategy suggests that random filters and neurons of the error-feedback fully connected layer are incrementally added to steadily compensate the prediction error during the construction process, and then a filter selection strategy is introduced to enable ESM-CNN to extract the different size of temporal features, providing helpful information at each iterative process for the prediction. The performance of ESM-CNN is justified on its prediction accuracy of one-step-ahead and multi-step-ahead forecasting tasks respectively. Comprehensive experiments on both the synthetic and real-world datasets show that the proposed ESM-CNN not only outperforms the state-of-art random neural networks, but also exhibits stronger predictive power and less computing overhead in comparison to trained state-of-art deep neural network models.

cs.LG