SearcharxivSearch

arXiv subjects

Angxiu Ni

Publications and source records attributed to Angxiu Ni.

At least 19 recordsLinked to original sources

Full-window branch discovery and loss-selected EnKF continuation for data assimilation

We develop a framework for offline full-window branch discovery, optionally followed by online continuation with an ensemble Kalman filter (EnKF). Three mechanisms drive the branch search: adjoint path-kernel (APK) differentiation balances kernel differentiation and correction-stabilized path perturbation, shifting the optimization from exploration to exploitation; an optimized Gaussian initial law broadens the search over initial-state basins; and loss-weighted mixing across independent runs recombines successful path components. We may then select an interior state using a local loss and continue online with an EnKF. In 40-dimensional Lorenz-96 experiments, the mean offline path RMSE of APK is 4.3 times smaller than that of population weak-$\mathrm{4D\text{-}Var}_x$. The resulting APK-EnKF method has a mean online RMSE 64 times smaller than that of ordinary EnKF.

math.NA

Optimal response for stochastic differential equations in $\mathbb{T}^d$ with perturbations on the drift term

We study stochastic differential equations on the $d$-dimensional flat torus $\mathbb{T}^d$ with drift and perturbation coefficients in $L^{\infty}(\mathbb{T}^d;\mathbb{R}^d)$ and additive non-degenerate noise. For the associated transfer operators, we analyse the dependence of the stationary measure and of the expectation of a given observable on small perturbations of the drift. In this framework, we prove a linear response formula for the invariant density and for the expectation of a given observable. We then address an optimal response problem, namely the determination of admissible perturbations that maximise the first-order variation of a prescribed observable. We establish existence of optimal perturbations and, in a Hilbert space framework, prove uniqueness and provide an explicit characterisation of the optimiser. This yields a practical Fourier-based numerical method, which we implement in several numerical examples, including both low and high-dimensional settings.

math.DS

Divergence-kernel method for linear responses of densities and generative models

We derive the divergence-kernel formula for the linear response of random dynamical systems. Specifically, the pathwise expression is for the parameter-derivative of the marginal or stationary density, not an averaged observable. Our formula works for multiplicative and parameterized noise over any period of time; it does not require hyperbolicity. Then we derive a Monte-Carlo algorithm for linear responses. We develop a new framework of generative models, DK-SDE, where the model is a parameterized SDE, that (1) directly uses the KL divergence between the empirical data distribution and the marginal density of the SDE as the training objective, and (2) accommodates parametrizations in both drift and diffusion over a long time span, allowing prior structural knowledge to be incorporated explicitly. The optimization is done by gradient-descent enabled by the divergence-kernel method, which involves only forward processes and therefore substantially reduces memory cost. We demonstrate the new model on a 20-dimensional Lorenz system.

math.DS

Adjoint path-kernel method for backpropagation and data assimilation in unstable diffusions

We derive the adjoint path-kernel method for computing parameter-gradients (linear responses) of SDEs. Its cost is almost independent of the number of parameters, and it works for non-hyperbolic systems with parameter-controlled multiplicative noise. With this new formula, we extend the conventional backpropagation method to settings with gradient explosion, and demonstrate it on the 40-dimensional Lorenz 96 system. Moreover, we consider a difficult version of the 4D-Var data assimilation problem where (1) the deterministic part of the model is chaotic, (2) the loss is a single long-time functional accounting for discrepancies in both the observations and the dynamics, (3) some parameters in the dynamics are unknown, and (4) some coordinates of the states cannot be observed, and cannot be reasonably inferred from other coordinates within a short time. We model the correction term at each time-step separately as a parameterized function of the random state. With our new tool, we can run stochastic gradient descent to find the path and parameters that best match the low-dimensional observation data. We demonstrate this on the 10D Lorenz-96 system with 8D observations.

math.PR

Divergence-Kernel method for scores of random systems

We derive the divergence-kernel formula for the scores of random dynamical systems, then formally pass to the continuous-time limit of SDEs. Our formula works for multiplicative noise systems over any period of time; it does not require hyperbolicity. We also consider several special cases: (1) for additive noise, we give a pure kernel formula; (2) for short-time, we give a pure divergence formula; (3) we give a formula which does not involve scores of the initial distribution. Based on the new formula, we derive a pathwise Monte-Carlo algorithm for scores, and demonstrate it on the 40-dimensional Lorenz 96 system with multiplicative noise.

math.PR

Path-Kernel Method for Differentiating Unstable Diffusions

We derive and prove the path-kernel formula for the linear response (parameter-derivative of averaged statistics) of SDEs. The parameter may affect the drift coefficient, the diffusion coefficient, and the initial condition. The formula tempers the unstableness by gradually moving the derivative from path-perturbation to kernel-differentiation, without assuming hyperbolicity. We prove it by direct comparison of bundles of paths across different parameter values. We also derive a pathwise Monte Carlo algorithm for estimating linear responses and demonstrate it on the 40-dimensional noisy Lorenz--96 system. Our result provides a new computational tool for optimization, and has already led to a follow-up application to data assimilation.

math.PR

Optimal Response for Hyperbolic Systems by the fast adjoint response method

In a uniformly hyperbolic system, we consider the problem of finding the optimal infinitesimal perturbation to apply to the system, from a certain set $P$ of feasible ones, to maximally increase the expectation of a given observation function. We perturb the system both by composing with a diffeomorphism near the identity or by adding a deterministic perturbation to the dynamics. In both cases, using the fast adjoint response formula, we show that the linear response operator, which associates the response of the expectation to the perturbation on the dynamics, is bounded in terms of the $C^{1,\alpha}$ norm of the perturbation. Under the assumption that $P$ is a strictly convex, closed subset of a Hilbert space $\cH$ that can be continuously mapped in the space of $C^3$ vector fields on our phase space, we show that there is a unique optimal perturbation in $P$ that maximizes the increase of the given observation function. Furthermore since the response operator is represented by a certain element $v$ of $\cH$, when the feasible set $P$ is the unit ball of $\cH$, the optimal perturbation is $v/||v||_{\cH}$. We also show how to compute the Fourier expansion $v$ in different cases. Our approach can work even on high dimensional systems. We demonstrate our method on numerical examples in dimensions 2, 3, and 21.

math.DS

Ergodic and foliated kernel-differentiation method for linear responses of random systems

We extend the kernel-differentiation method for the linear response (parameter-derivative of averaged observables) of random dynamical systems. First, for the linear response of physical (or stationary) measures, we extend the method to an ergodic version, which is sampled by an infinitely long sample path, so it is more efficient than previous results. This is achieved by combining the likelihood ratio trick, decay of correlations, and ergodic theorem. Second, when the noise and perturbation are along a given foliation, we show that the method is still valid for both finite and infinite time. These results are derived using basic calculus via a microscopic view of transfer operators. We use the ergodic formula to numerically compute the linear response of a tent map with additive noise. We use the foliated formula to compute the linear response of an unstable neural network with 51 layers $\times$ 9 neurons, with respect to the bias parameter. We show that adding foliated noise incurs a smaller error than adding noise in all directions, in terms of approximating the trend between the averaged observable and the parameter. We give a rough error analysis of the kernel-differentiation method and show that it can be expensive for small noise. Finally, we derive the three basic linear response methods (path-perturbation, divergence, and kernel-differentiation methods) in simplified settings, and propose a potential future program unifying them.

math.PR

Equivariant divergence formula for chaotic flows

We prove the equivariant divergence formula for the axiom A flow attractors, which is a recursive formula for perturbation of transfer operators of physical measures along center-unstable manifolds. Hence the linear response acquires an `ergodic theorem', which means that it can be sampled by recursively computing only $2u$ many vectors on one orbit, where $u$ is the unstable dimension.

math.DS

No-propagate algorithm for linear responses of random chaotic systems

We develop the no-propagate algorithm for sampling the linear response of random dynamical systems, which are non-uniform hyperbolic deterministic systems perturbed by noise with smooth density. We first derive a Monte-Carlo type formula and then the algorithm, which is different from the ensemble (stochastic gradient) algorithms, finite-element algorithms, and fast-response algorithms; it does not involve the propagation of vectors or covectors, and only the density of the noise is differentiated, so the formula is not cursed by gradient explosion, dimensionality, or non-hyperbolicity. We demonstrate our algorithm on a tent map perturbed by noise and a chaotic neural network with 51 layers $\times$ 9 neurons. By itself, this algorithm approximates the linear response of non-hyperbolic deterministic systems, with an additional error proportional to the noise. We also discuss the potential of using this algorithm as a part of a bigger algorithm with smaller error.

math.DS

Backpropagation in hyperbolic chaos via adjoint shadowing

To generalize the backpropagation method to both discrete-time and continuous-time hyperbolic chaos, we introduce the adjoint shadowing operator $\mathcal{S}$ acting on covector fields. We show that $\mathcal{S}$ can be equivalently defined as: (a) $\mathcal{S}$ is the adjoint of the linear shadowing operator $S$; (b) $\mathcal{S}$ is given by a `split then propagate' expansion formula; (c) $\mathcal{S}(\omega)$ is the only bounded inhomogeneous adjoint solution of $\omega$. By (a), $\mathcal{S}$ adjointly expresses the shadowing contribution, a significant part of the linear response, where the linear response is the derivative of the long-time statistics with respect to system parameters. By (b), $\mathcal{S}$ also expresses the other part of the linear response, the unstable contribution. By (c), $\mathcal{S}$ can be efficiently computed by the nonintrusive shadowing algorithm in Ni and Talnikar (2019 J. Comput. Phys. 395 690-709), which is similar to the conventional backpropagation algorithm. For continuous-time cases, we additionally show that the linear response admits a well-defined decomposition into shadowing and unstable contributions.

math.DS

Fast adjoint differentiation of chaos via computing unstable perturbations of transfer operators

We devise the fast adjoint response algorithm for the gradient of physical measures (long-time-average statistics) of discrete-time hyperbolic chaos with respect to many system parameters. Its cost is independent of the number of parameters. The algorithm transforms our new theoretical tools, the adjoint shadowing lemma and the equivariant divergence formula, into the form of progressively computing $u$ many bounded vectors on one orbit. Here $u$ is the unstable dimension. We demonstrate our algorithm on an example difficult for previous methods, a system with random noise, and a system of a discontinuous map. We also give a short formal proof of the equivariant divergence formula. Compared to the better-known finite-element method, our algorithm is not cursed by dimensionality of the phase space (typical real-life systems have very high dimensions), since it samples by one orbit. Compared to the ensemble/stochastic method, our algorithm is not cursed by the butterfly effect, since the recursive relations in our algorithm is bounded.

math.DS

Recursive divergence formulas for perturbing unstable transfer operators and physical measures

We show that the derivative of the (measure) transfer operator with respect to the parameter of the map is a divergence. Then, for physical measures of discrete-time hyperbolic chaotic systems, we derive an equivariant divergence formula for the unstable perturbation of transfer operators along unstable manifolds. This formula and hence the linear response, the parameter-derivative of physical measures, can be sampled by recursively computing only $2u$ many vectors on one orbit, where $u$ is the unstable dimension. The numerical implementation of this formula in \cite{far} is neither cursed by dimensionality nor the sensitive dependence on initial conditions.

math.NA

Fast differentiation of hyperbolic chaos

We derive and prove the `fast response' formula for the linear response, the parameter derivatives of long-time-averaged statistics, of hyperbolic deterministic chaotic systems. The expression is pointwisely defined so we can compute the linear response in high-dimensions via Monte-Carlo-type algorithms. It has two parts, where the shadowing contribution is computed by the nonintrusive shadowing algorithm. The unstable contribution is expressed by renormalized second-order tangent equations; importantly, it does not contain any distributional derivatives. The algorithm's cost is solving $u$, the unstable dimension, many first-order and second-order tangent equations along a long orbit; the main error is the sampling error of the orbit. We numerically demonstrate the algorithm on a 21-dimensional example, which is difficult for previous methods.

math.DS

Approximating linear response by nonintrusive shadowing algorithms

Nonintrusive shadowing algorithms efficiently compute $v$, the difference between shadowing trajectories, then use $v$ to compute derivatives of averaged objectives of chaos with respect to parameters of the dynamical system. However, previous proofs of shadowing methods wrongly assume that shadowing trajectories are representative. In contrast, the linear response formula is proved rigorously, but is more difficult to compute. We prove that $v$ gives only a part, called the shadowing contribution, of the linear response; hence, the other part, the unstable contribution, is the systematic error of shadowing methods. For systems with a small ratio of unstable dimensions, with some further statistical assumptions, we show that the unstable contribution is small. We also briefly describe an algorithm for the unstable contribution, which is simpler to derive but less efficient than the fast linear response algorithm. Moreover, we prove the convergence of the nonintrusive shadowing algorithm, the fastest shadowing algorithm, to $v$ and to the shadowing contribution.

math.DS

Linear Range in Gradient Descent

This paper defines linear range as the range of parameter perturbations which lead to approximately linear perturbations in the states of a network. We compute linear range from the difference between actual perturbations in states and the tangent solution. Linear range is a new criterion for estimating the effectivenss of gradients and thus having many possible applications. In particular, we propose that the optimal learning rate at the initial stages of training is such that parameter changes on all minibatches are within linear range. We demonstrate our algorithm on two shallow neural networks and a ResNet.

math.OC

Adjoint shadowing directions in hyperbolic systems for sensitivity analysis

For hyperbolic diffeomorphisms, we define adjoint shadowing directions as a bounded inhomogeneous adjoint solution whose initial condition has zero component in the unstable adjoint direction. For hyperbolic flows, we define adjoint shadowing directions similarly, with the additional requirement that the average of its inner-product with the trajectory direction is zero. In both cases, we show unique existence of adjoint shadowing directions, and how they can be used for adjoint sensitivity analysis. Our work set a theoretical foundation for efficient adjoint sensitivity methods for long-time-averaged objectives such as NILSAS.

math.DS

Adjoint sensitivity analysis on chaotic dynamical systems by Non-Intrusive Least Squares Adjoint Shadowing (NILSAS)

We develop the NILSAS algorithm, which performs adjoint sensitivity analysis of chaotic systems via computing the adjoint shadowing direction. NILSAS constrains its minimization to the adjoint unstable subspace, and can be implemented with little modification to existing adjoint solvers. The computational cost of NILSAS is independent of the number of parameters. We demonstrate NILSAS on the Lorenz 63 system and a weakly turbulent three-dimensional flow over a cylinder.

physics.comp-ph