SearcharxivSearch

arXiv subjects

Jochen Bröcker

Publications and source records attributed to Jochen Bröcker.

18 recordsLinked to original sources

Learning Probabilistic Filters with Strictly Proper Scoring Rules

Bayesian filtering of partially and noisily observed dynamical systems seeks to infer the evolving conditional distribution of the state of a dynamical system, given observations, in an online fashion. This Bayesian filtering distribution is the natural object for uncertainty quantification, but it is rarely available as a supervised learning target. However, one can often use the forecast model to generate synthetic system trajectories, along with synthetic observations. We introduce the proper scoring ensemble filter (PSEF), an ensemble data assimilation method based on training an analysis map to approximate the filtering distribution using only synthetic state--observation trajectories. The analysis step is represented as a permutation-invariant, transformer-based map that takes as input a forecast ensemble and observations, producing an analysis ensemble. Training is based on strictly proper scoring rules -- with the energy score used in our implementation -- so that probabilistic accuracy is rewarded over the whole probability distribution. We prove that, under a realizability assumption, the population objective is minimized by the true Bayesian filtering distribution. We also derive the finite-ensemble empirical objective used in training and relate its single state--observation trajectory form to the population objective, using a mean-field consistency argument. Numerical experiments show that the learned filter accurately approximates challenging filtering distributions, including nonlinear, non-Gaussian, and multi-modal posteriors, and achieves stronger performance in data assimilation tasks than classical methods or learning-based methods with mean-squared-error objectives. For close-to-Gaussian problems, learning a correction to the EnKF is the best approach, while for highly non-Gaussian problems an end-to-end approach that discards this inductive bias is superior.

cs.LG

Continuous Data Assimilation for Semilinear Parabolic Equations with Multiplicative Observation Noise

The problem of continuous data assimilation for semilinear parabolic equations based on partial observations corrupted by noise is investigated. The noise is allowed to be multiplicative, with additive noise arising as a special case. In a general Gelfand triple framework, an abstract theory for the nudging equation is developed that covers both weak and strong formulations. Mean square convergence of the assimilation error is proved under suitable assumptions, and, under additional integrability conditions on the noise, a uniform almost sure convergence result is established. Finally, the framework is applied to several PDE models, including the 2D Navier-Stokes, 2D magnetohydrodynamics, 2D quasi-geostrophic, and 1D Allen-Cahn equations.

math.AP

A generalisation of the signal-to-noise ratio using proper scoring rules

A generalised concept of the signal-to-noise ratio (or equivalently the ratio of predictable components, or RPC) is provided, based on proper scoring rules. This definition is the natural generalisation of the classical RPC, yet it allows one to define and analyse the signal-to-noise properties of any type of forecast that is amenable to scoring, thus drastically widening the applicability of these concepts. The methodology is illustrated through numerical examples of ensemble forecasts, scored using the continuous ranked probability score (CRPS), and of probability forecasts of a binary event, scored using the logarithmic score. Numerical examples are carried out using both synthetic data with prescribed signal-to-noise ratios as well as seasonal ensemble hindcasts of the North Atlantic Oscillation (NAO) index. The latter have previously been interpreted as having a signal-to-noise "paradox", or anomalous signal-to-noise ratio, using the RPC statistic. For the synthetic data, the RPC statistic as well as the scoring rule-based ones agree regarding which data sets exhibit anomalous signal-to-noise ratios, but exhibit different variance, indicating different statistical properties. For the NAO data, on the other hand, the different statistics are more equivocal on whether the signal-to-noise ratio is anomalous.

stat.AP

Reconstruction of wide spectrum forcing in transport-diffusion and Navier-Stokes equations

This article considers the problem of reconstructing unknown driving forces based on incomplete knowledge of the system and its state. This is studied in both a linear and nonlinear setting that is paradigmatic in geophysical fluid dynamics and various applications. Two algorithms are proposed to address this problem: one that iteratively reconstructs forcing and another that provides a continuous-time reconstruction. Convergence is shown to be guaranteed provided that observational resolution is sufficiently high and algorithmic parameters are properly tuned according to the prior information; these conditions are quantified precisely. The class of reconstructable forces identified here include those which are time-dependent and potentially inject energy at all length scales. This significantly expands upon the class of forces in previous studies, which could only accommodate those with band-limited spectra. The second algorithm moreover provides a conceptually streamlined approach that allows for a more straightforward analysis and simplified practical implementation.

math.OC

The evolution of a non-autonomous chaotic system under non-periodic forcing: a climate change example

Complex Earth System Models are widely utilised to make conditional statements about the future climate under some assumptions about changes in future atmospheric greenhouse gas concentrations; these statements are often referred to as climate projections. The models themselves are high-dimensional nonlinear systems and it is common to discuss their behaviour in terms of attractors and low-dimensional nonlinear systems such as the canonical Lorenz `63 system. In a non-autonomous situation, for instance due to anthropogenic climate change, the relevant object is sometimes considered to be the pullback or snapshot attractor. The pullback attractor, however, is a collection of {\em all} plausible states of the system at a given time and therefore does not take into consideration our knowledge of the current state of the Earth System when making climate projections, and are therefore not very informative regarding annual to multi-decadal climate projections. In this article, we approach the problem of measuring and interpreting the mid-term climate of a model by using a low-dimensional, climate-like, nonlinear system with three timescales of variability, and non-periodic forcing. We introduce the concept of an {\em evolution set} which is dependent on the starting state of the system, and explore its links to different types of initial condition uncertainty and the rate of external forcing. We define the {\em convergence time} as the time that it takes for the distribution of one of the dependent variables to lose memory of its initial conditions. We suspect a connection between convergence times and the classical concept of mixing times but the precise nature of this connection needs to be explored. These results have implications for the design of influential climate and Earth System Model ensembles, and raise a number of issues of mathematical interest.

physics.geo-ph

Linear and fractional response for nonlinear dissipative SPDEs

A framework to establish response theory for a class of nonlinear stochastic partial differential equations (SPDEs) is provided. More specifically, it is shown that for a certain class of observables, the averages of those observables against the stationary measure of the SPDE are differentiable (linear response) or, under weaker conditions, locally Hölder continuous (fractional response) as functions of a deterministic additive forcing. The method allows to consider observables that are not necessarily differentiable. For such observables, spectral gap results for the Markov semigroup associated with the SPDE have recently been established that are fairly accessible. This is important here as spectral gaps are a major ingredient for establishing linear response. The results are applied to the 2D stochastic Navier-Stokes equation and the stochastic two-layer quasi-geostrophic model, an intermediate complexity model popular in the geosciences to study atmosphere and ocean dynamics. The physical motivation for studying the response to perturbations in the forcings for models in geophysical fluid dynamics comes from climate change and relate to the question as to whether statistical properties of the dynamics derived under current conditions will be valid under different forcing scenarios.

math-ph

Exponential ergodicity for a stochastic two-layer quasi-geostrophic model

Ergodic properties of a stochastic medium complexity model for atmosphere and ocean dynamics are analysed. More specifically, a two-layer quasi-geostrophic model for geophysical flows is studied, with the upper layer being perturbed by additive noise. This model is popular in the geosciences, for instance to study the effects of a stochastic wind forcing on the ocean. A rigorous mathematical analysis however meets with the challenge that in the model under study, the noise configuration is spatially degenerate as the stochastic forcing acts only on the top layer. Exponential convergence of solutions laws to the invariant measure is established, implying a spectral gap of the associated Markov semigroup on a space of Hölder continuous functions. The approach provides a general framework for generalised coupling techniques suitable for applications to dissipative SPDEs. In case of the two-layer quasi-geostrophic model, the results require the second layer to obey a certain passivity condition.

math.PR

Exponential stability and asymptotic properties of the optimal filter for signals with deterministic hyperbolic dynamics

The problem of stability of the optimal filter is revisited. The optimal filter (or filtering process) is the conditional probability of the current state of some stochastic process (the signal process), given both present and past values of another process (the observation process). Typically the filtering process satisfies a dynamical equation, and the question investigated here concerns the stability of this dynamics. In contrast to previous work, signal processes given by the iterations of a deterministic mapping $f$ are considered, with only the initial condition being random. While the stability of the filter may emerge from strong randomness of the signal processes, different and more dynamical effects will be exploited in the present work. More specifically, we consider uniformly hyperbolic $f$ with strong instabilities providing the necessary mixing. This however requires that the filtering process is initialised with densities exhibiting already a certain level of smoothness. Furthermore, $f$ may also have stable directions along which the filtering process will eventually not have a density, a major technical difficulty. Further results show that the filtering process is asymptotically concentrated on the attractor and furthermore will have densities with respect to the invariant (SRB)~measure along instable manifolds of $f$.

math.PR

Uniform reliability tests for forecasting systems with small lead time

A long noted difficulty when assessing the reliability (or calibration) of forecasting systems is that reliability, in general, is a hypothesis not about a finite dimensional parameter but about an entire functional relationship. A calibrated probability forecast for binary events for instance should equal the conditional probability of the event given the forecast for {\em any} value of the forecast. Attempts to estimate deviations from calibration at a specific forecast value meet with the difficulty that the probability of the forecast assuming that value is typically zero. Considering the estimated {\em cumulative} deviations from reliability instead however, tests are presented for which the asymptotic distribution of the test statistic can be established rigorously. The distribution turns out to be universal, provided the forecasts "look one step ahead" only, or in other words, verify at the next time step in the future. Furthermore, the tests develop power against a wide class of alternatives. Numerical experiments for both artificial data as well as operational weather forecasting systems are also presented, as are possible extensions to forecasts with longer lead times.

physics.data-an

Existence and Uniqueness For Variational Data Assimilation in Continuous Time

A variant of the optimal control problem is considered which is nonstandard in that the performance index contains "stochastic" integrals, that is, integrals against very irregular functions. The motivation for considering such performance indices comes from dynamical estimation problems where observed time series need to be "fitted" with trajectories of dynamical models. The observations may be contaminated with white noise, which gives rise to the nonstandard performance indices. Problems of this kind appear in engineering, physics, and the geosciences where this is referred to as data assimilation. Pathwise existence of minimisers is obtained, along with a maximum principle as well as preliminary results in dynamic programming. The results extend previous results on the maximum aposteriori estimator of trajectories of diffusion processes. To obtain these results, classical concepts from optimal control need to be substantially modified due to the nonstandard nature of the performance index, as well the fact that typical models in the geosciences do not satisfy linear growth nor monotonicity conditions.

math.OC

Almost sure error bounds for data assimilation in dissipative systems with unbounded observation noise

Data assimilation is uniquely challenging in weather forecasting due to the high dimensionality of the employed models and the nonlinearity of the governing equations. Although current operational schemes are used successfully, our understanding of their long-term error behaviour is still incomplete. In this work, we study the error of some simple data assimilation schemes in the presence of unbounded (e.g. Gaussian) noise on a wide class of dissipative dynamical systems with certain properties, including the Lorenz models and the 2D incompressible Navier-Stokes equations. We exploit the properties of the dynamics to derive analytic bounds on the long-term error for individual realisations of the noise in time. These bounds are proportional to the amplitude of the noise. Furthermore, we find that the error exhibits a form of stationary behaviour, and in particular an accumulation of error does not occur. This improves on previous results in which either the noise was bounded or the error was considered in expectation only.

physics.ao-ph

What is the correct cost functional for variational data assimilation?

Variational approaches to data assimilation, and weakly constrained four dimensional variation (WC-4DVar) in particular, are important in the geosciences but also in other communities (often under different names). The cost functions and the resulting optimal trajectories may have a probabilistic interpretation, for instance by linking data assimilation with Maximum Aposteriori (MAP) estimation. This is possible in particular if the unknown trajectory is modelled as the solution of a stochastic differential equation (SDE), as is increasingly the case in weather forecasting and climate modelling. In this case, the MAP estimator (or "most probable path" of the SDE) is obtained by minimising the Onsager--Machlup functional. Although this fact is well known, there seems to be some confusion in the literature, with the energy (or "least squares") functional sometimes been claimed to yield the most probable path. The first aim of this paper is to address this confusion and show that the energy functional does not, in general, provide the most probable path. The second aim is to discuss the implications in practice. Although the mentioned results pertain to stochastic models in continuous time, they do have consequences in practice where SDE's are approximated by discrete time schemes. It turns out that using an approximation to the SDE and calculating its most probable path does not necessarily yield a good approximation to the most probable path of the SDE proper. This suggest that even in discrete time, a version of the Onsager--Machlup functional should be used, rather than the energy functional, at least if the solution is to be interpreted as a MAP estimator.

physics.ao-ph

Assessing the reliability of ensemble forecasting systems under serial dependence

The problem of testing the reliability of ensemble forecasting systems is revisited. A popular tool to assess the reliability of ensemble forecasting systems (for scalar verifications) is the rank histogram, this histogram is expected to be more or less flat, since for a reliable ensemble, the ranks are uniformly distributed among their possible outcomes. Quantitative tests for flatness (e.g.\ Pearson's goodness--of--fit test) have been suggested, without exception though, these tests assume the ranks to be a sequence of independent random variables, which is not the case in general as can be demonstrated with simple toy examples. In this paper, tests are developed that take the temporal correlations between the ranks into account. A refined analysis shows that exploiting the reliability property, the ranks still exhibit strong decay of correlations. This property is key to the analysis, and the proposed tests are valid for general ensemble forecasting systems with minimal extraneous assumptions.

physics.ao-ph

Sensitivity And Out-Of-Sample Error in Continuous Time Data Assimilation

Data assimilation refers to the problem of finding trajectories of a prescribed dynamical model in such a way that the output of the model (usually some function of the model states) follows a given time series of observations. Typically though, these two requirements cannot both be met at the same time--tracking the observations is not possible without the trajectory deviating from the proposed model equations, while adherence to the model requires deviations from the observations. Thus, data assimilation faces a trade-off. In this contribution, the sensitivity of the data assimilation with respect to perturbations in the observations is identified as the parameter which controls the trade-off. A relation between the sensitivity and the out-of-sample error is established which allows to calculate the latter under operational conditions. A minimum out-of-sample error is proposed as a criterion to set an appropriate sensitivity and to settle the discussed trade-off. Two approaches to data assimilation are considered, namely variational data assimilation and Newtonian nudging, aka synchronisation. Numerical examples demonstrate the feasibility of the approach.

physics.ao-ph

On Variational Data Assimilation in Continuous Time

Variational data assimilation in continuous time is revisited. The central techniques applied in this paper are in part adopted from the theory of optimal nonlinear control. Alternatively, the investigated approach can be considered as a continuous time generalisation of what is known as weakly constrained four dimensional variational assimilation (WC--4DVAR) in the geosciences. The technique allows to assimilate trajectories in the case of partial observations and in the presence of model error. Several mathematical aspects of the approach are studied. Computationally, it amounts to solving a two point boundary value problem. For imperfect models, the trade off between small dynamical error (i.e. the trajectory obeys the model dynamics) and small observational error (i.e. the trajectory closely follows the observations) is investigated. For (nearly) perfect models, this trade off turns out to be (nearly) trivial in some sense, yet allowing for some dynamical error is shown to have positive effects even in this situation. The presented formalism is dynamical in character; no assumptions need to be made about the presence (or absence) of dynamical or observational noise, let alone about their statistics.

physics.ao-ph

A Lower Bound on Arbitrary $f$--Divergences in Terms of the Total Variation

An important tool to quantify the likeness of two probability measures are f-divergences, which have seen widespread application in statistics and information theory. An example is the total variation, which plays an exceptional role among the f-divergences. It is shown that every f-divergence is bounded from below by a monotonous function of the total variation. Under appropriate regularity conditions, this function is shown to be monotonous. Remark: The proof of the main proposition is relatively easy, whence it is highly likely that the result is known. The author would be very grateful for any information regarding references or related work.

math.PR

Generating Probabilities From Numerical Weather Forecasts by Logistic Regression

Logistic models are studied as a tool to convert output from numerical weather forecasting systems (deterministic and ensemble) into probability forecasts for binary events. A logistic model obtains by putting the logarithmic odds ratio equal to a linear combination of the inputs. As any statistical model, logistic models will suffer from over-fitting if the number of inputs is comparable to the number of forecast instances. Computational approaches to avoid over-fitting by regularisation are discussed, and efficient approaches for model assessment and selection are presented. A logit version of the so called lasso, which is originally a linear tool, is discussed. In lasso models, less important inputs are identified and discarded, thereby providing an efficient and automatic model reduction procedure. For this reason, lasso models are particularly appealing for diagnostic purposes.

physics.ao-ph

Reliability, Sufficiency, and the Decomposition of Proper Scores

Scoring rules are an important tool for evaluating the performance of probabilistic forecasting schemes. In the binary case, scoring rules (which are strictly proper) allow for a decomposition into terms related to the resolution and to the reliability of the forecast. This fact is particularly well known for the Brier Score. In this paper, this result is extended to forecasts for finite--valued targets. Both resolution and reliability are shown to have a positive effect on the score. It is demonstrated that resolution and reliability are directly related to forecast attributes which are desirable on grounds independent of the notion of scores. This finding can be considered an epistemological justification of measuring forecast quality by proper scores. A link is provided to the original work of DeGroot et al (1982), extending their concepts of sufficiency and refinement. The relation to the conjectured sharpness principle of Gneiting et al (2005a) is elucidated.

physics.ao-ph