SearcharxivSearch

arXiv subjects

Jens Timmer

Publications and source records attributed to Jens Timmer.

13 recordsLinked to original sources

Detecting frequency modulation in stochastic time series data

We propose a new statistical test to identify non-stationary frequency-modulated stochastic processes from time series data. Our method uses the instantaneous phase as a discriminatory statistics with reliable critical values derived from surrogate data. We simulated an oscillatory second-order autoregressive process to evaluate the size and power of the test. We found that the test we propose is able to correctly identify more than 99% of non-stationary data when the frequency of simulated data is doubled after the first half of the time series. Our method is easily interpretable, computationally cheap and does not require choosing hyperparameters that are dependent on the data.

physics.data-an

Non-parametric model-based estimation of the effective reproduction number for SARS-CoV-2

Viral outbreaks, such as the current COVID-19 pandemic, are commonly described by compartmental models by means of ordinary differential equation (ODE) systems. The parameter values of these ODE models are typically unknown and need to be estimated based on accessible data. In order to describe realistic pandemic scenarios with strongly varying situations, these model parameters need to be assumed as time-dependent. While parameter estimation for the typical case of time-constant parameters does not pose larger issues, the determination of time-dependent parameters, e.g.~the transition rates of compartmental models, remains notoriously difficult, in particular since the function class of these time-dependent parameters is unknown. In this work, we present a novel method which utilizes the Augmented Kalman Smoother in combination with an Expectation-Maximization algorithm to simultaneously estimate all time-dependent parameters in an SIRD compartmental model. This approach only requires incidence data, but no prior knowledge on model parameters or any further assumptions on the function class of the time-dependencies. In contrast to other approaches for the estimation of the time-dependent reproduction number, no assumptions on the parameterization of the serial interval distribution are required. With this method, we are able to adequately describe COVID-19 data in Germany and to give non-parametric model-based time course estimates for the effective reproduction number.

q-bio.PE

On structural and practical identifiability

We discuss issues of structural and practical identifiability of partially observed differential equations which are often applied in systems biology. The development of mathematical methods to investigate structural non-identifiability has a long tradition. Computationally efficient methods to detect and cure it have been developed recently. Practical non-identifiability on the other hand has not been investigated at the same conceptually clear level. We argue that practical identifiability is more challenging than structural identifiability when it comes to modelling experimental data. We discuss that the classical approach based on the Fisher information matrix has severe shortcomings. As an alternative, we propose using the profile likelihood, which is a powerful approach to detect and resolve practical non-identifiability.

stat.ME

PEtab -- interoperable specification of parameter estimation problems in systems biology

Reproducibility and reusability of the results of data-based modeling studies are essential. Yet, there has been -- so far -- no broadly supported format for the specification of parameter estimation problems in systems biology. Here, we introduce PEtab, a format which facilitates the specification of parameter estimation problems using Systems Biology Markup Language (SBML) models and a set of tab-separated value files describing the observation model and experimental data as well as parameters to be estimated. We already implemented PEtab support into eight well-established model simulation and parameter estimation toolboxes with hundreds of users in total. We provide a Python library for validation and modification of a PEtab problem and currently 20 example parameter estimation problems based on recent studies. Specifications of PEtab, the PEtab Python library, as well as links to examples, and all supporting software tools are available at https://github.com/PEtab-dev/PEtab, a snapshot is available at https://doi.org/10.5281/zenodo.3732958. All original content is available under permissive licenses.

q-bio.QM

Analyzing effective models: An example from JAK/STAT5 signaling

In systems biology effective models are widely used due to the complexity of biological system. They result from a coarse-graining process which employs specific assumptions. Frequently one does not start with a model taking all details into account and then performs a coarse-graining process, but rather one starts right away with the effective equations and often the underlying assumptions remain hidden or unclear. We exemplify the analysis of an effective model by analyzing a time delay equation for the JAK/STAT5 signaling pathway and show how one can avoid wrong conclusions and obtain a deeper understanding of the biological system . By analyzing the assumptions leading to a coarse-grained model one might be able to gain new insight into the involved biological processes. Further, the compliance of the model with experimental data can be considered as a validation of the assumptions made in the derivation of the mathematical equations.

q-bio.MN

Cause and Cure of Sloppiness in Ordinary Differential Equation Models

Data-based mathematical modeling of biochemical reaction networks, e.g. by nonlinear ordinary differential equation (ODE) models, has been successfully applied. In this context, parameter estimation and uncertainty analysis is a major task in order to assess the quality of the description of the system by the model. Recently, a broadened eigenvalue spectrum of the Hessian matrix of the objective function covering orders of magnitudes was observed and has been termed as sloppiness. In this work, we investigate the origin of sloppiness from structures in the sensitivity matrix arising from the properties of the model topology and the experimental design. Furthermore, we present strategies using optimal experimental design methods in order to circumvent the sloppiness issue and present non-sloppy designs for a benchmark model.

q-bio.MN

A Unified Approach to Integration and Optimization of Parametric Ordinary Differential Equations

Parameter estimation in ordinary differential equations, although applied and refined in various fields of the quantitative sciences, is still confronted with a variety of difficulties. One major challenge is finding the global optimum of a log-likelihood function that has several local optima, e.g. in oscillatory systems. In this publication, we introduce a formulation based on continuation of the log-likelihood function that allows to restate the parameter estimation problem as a boundary value problem. By construction, the ordinary differential equations are solved and the parameters are estimated both in one step. The formulation as a boundary value problem enables an optimal transfer of information given by the measurement time courses to the solution of the estimation problem, thus favoring convergence to the global optimum. This is demonstrated explicitly for the fully as well as the partially observed Lotka-Volterra system.

q-bio.QM

A Variational Approach to Parameter Estimation in Ordinary Differential Equations

Ordinary differential equations are widely-used in the field of systems biology and chemical engineering to model chemical reaction networks. Numerous techniques have been developed to estimate parameters like rate constants, initial conditions or steady state concentrations from time-resolved data. In contrast to this countable set of parameters, the estimation of entire courses of network components corresponds to an innumerable set of parameters. The approach presented in this work is able to deal with course estimation for extrinsic system inputs or intrinsic reactants, both not being constrained by the reaction network itself. Our method is based on variational calculus which is carried out analytically to derive an augmented system of differential equations including the unconstrained components as ordinary state variables. Finally, conventional parameter estimation is applied to the augmented system resulting in a combined estimation of courses and parameters. The combined estimation approach takes the uncertainty in input courses correctly into account. This leads to precise parameter estimates and correct confidence intervals. In particular this implies that small motifs of large reaction networks can be analysed independently of the rest. By the use of variational methods, elements from control theory and statistics are combined allowing for future transfer of methods between the two fields.

q-bio.MN

Joining Forces of Bayesian and Frequentist Methodology: A Study for Inference in the Presence of Non-Identifiability

Increasingly complex applications involve large datasets in combination with non-linear and high dimensional mathematical models. In this context, statistical inference is a challenging issue that calls for pragmatic approaches that take advantage of both Bayesian and frequentist methods. The elegance of Bayesian methodology is founded in the propagation of information content provided by experimental data and prior assumptions to the posterior probability distribution of model predictions. However, for complex applications experimental data and prior assumptions potentially constrain the posterior probability distribution insufficiently. In these situations Bayesian Markov chain Monte Carlo sampling can be infeasible. From a frequentist point of view insufficient experimental data and prior assumptions can be interpreted as non-identifiability. The profile likelihood approach offers to detect and to resolve non-identifiability by experimental design iteratively. Therefore, it allows one to better constrain the posterior probability distribution until Markov chain Monte Carlo sampling can be used securely. Using an application from cell biology we compare both methods and show that a successive application of both methods facilitates a realistic assessment of uncertainty in model predictions.

physics.data-an

Likelihood based observability analysis and confidence intervals for predictions of dynamic models

Mechanistic dynamic models of biochemical networks such as Ordinary Differential Equations (ODEs) contain unknown parameters like the reaction rate constants and the initial concentrations of the compounds. The large number of parameters as well as their nonlinear impact on the model responses hamper the determination of confidence regions for parameter estimates. At the same time, classical approaches translating the uncertainty of the parameters into confidence intervals for model predictions are hardly feasible. In this article it is shown that a so-called prediction profile likelihood yields reliable confidence intervals for model predictions, despite arbitrarily complex and high-dimensional shapes of the confidence regions for the estimated parameters. Prediction confidence intervals of the dynamic states allow a data-based observability analysis. The approach renders the issue of sampling a high-dimensional parameter space into evaluating one-dimensional prediction spaces. The method is also applicable if there are non-identifiable parameters yielding to some insufficiently specified model predictions that can be interpreted as non-observability. Moreover, a validation profile likelihood is introduced that should be applied when noisy validation experiments are to be interpreted. The properties and applicability of the prediction and validation profile likelihood approaches are demonstrated by two examples, a small and instructive ODE model describing two consecutive reactions, and a realistic ODE model for the MAP kinase signal transduction pathway. The presented general approach constitutes a concept for observability analysis and for generating reliable confidence intervals of model predictions, not only, but especially suitable for mathematical models of biological systems.

physics.data-an

Induction level determines signature of gene expression noise in cellular systems

Noise in gene expression, either due to inherent stochasticity or to varying inter- and intracellular environment, can generate significant cell-to-cell variability of protein levels in clonal populations. We present a theoretical framework, based on stochastic processes, to quantify the different sources of gene expression noise taking cell division explicitly into account. Analytical, time-dependent solutions for the noise contributions arising from the major steps involved in protein synthesis are derived. The analysis shows that the induction level of the activator or transcription factor is crucial for the characteristic signature of the dominant source of gene expression noise and thus bridges the gap between seemingly contradictory experimental results. Furthermore, on the basis of experimentally measured cell distributions, our simulations suggest that transcription factor binding and promoter activation can be modelled independently of each other with sufficient accuracy.

q-bio.OT

Analyzing X-ray variability by Linear State Space Models

In recent years, autoregressive models have had a profound impact on the description of astronomical time series as the observation of a stochastic process. These methods have advantages compared with common Fourier techniques concerning their inherent stationarity and physical background. However, if autoregressive models are used, it has to be taken into account that real data always contain observational noise often obscuring the intrinsic time series of the object. We apply the technique of a Linear State Space Model which explicitly models the noise of astronomical data and allows to estimate the hidden autoregressive process. As an example, we have analysed the X-ray flux variability of the Active Galaxy NGC 5506 observed with EXOSAT.

astro-ph

The Seyfert Galaxy NGC 6814 -- a highly variable X-ray source

The Seyfert galaxy NGC 6814 is a highly variable X-ray source despite the fact that it has recently been shown not to be the source of periodic variability. The 1.5 year monitoring by ROSAT has revealed a long term downward trend of the X-ray flux and an episode of high and rapidly varying flux (e.g. by a factor of about 3 in 8 hours) during the October 1992 PSPC observation. Temporal analysis of this data using both Fourier and autoregressive techniques have shown that the variability timescales are larger than a few hundred seconds. The behavior at higher frequencies can be described by white noise.

astro-ph