SearcharxivSearch

arXiv subjects

Adeline Samson

Publications and source records attributed to Adeline Samson.

At least 19 recordsLinked to original sources

Strang splitting estimator for nonlinear multivariate stochastic differential equations with Pearson-type multiplicative noise

Multivariate Pearson diffusions are characterized by a linear drift and a diffusion matrix that is quadratic in the state variables. We derive closed-form expressions for the mean and covariance matrix of this class using matrix exponential integrals, and extend this framework to a broader class of nonlinear diffusions with Pearson-type multiplicative noise. The main contribution is a new parameter estimator for these nonlinear multiplicative models based on Strang splitting, which decomposes the stochastic system into a deterministic nonlinear ordinary differential equation and a multivariate Pearson diffusion. The estimator is constructed by composing their respective flows and applying a Gaussian transition approximation with exact moments from the Pearson component. We prove that the estimator is consistent and asymptotically efficient. We also introduce a new model within this class, the Student Kramers oscillator, and prove existence and uniqueness of the strong solution and of an invariant measure. We evaluate the estimator through simulation studies on this oscillator and on the multivariate Wright-Fisher diffusion from population genetics, where it outperforms the Euler-Maruyama, Gaussian approximation, and local linearization estimators. We conclude with an application to Greenland ice core data using the Student Kramers oscillator.

stat.ME

Hawkes process with a diffusion-driven baseline: long-run behavior, inference, statistical tests

Event-driven systems in fields such as neuroscience, social networks, and finance often exhibit dynamics influenced by continuously evolving external covariates. Motivated by these applications, we introduce a new class of multivariate Hawkes processes, in which the spontaneous rate of events is modulated by a diffusion process. This framework allows the point process to adapt dynamically to continuously evolving covariates, capturing both intrinsic self-excitation and external influences. In this article, we establish the probabilistic properties of the coupled process, proving stability and ergodicity under moderate assumptions. Classical functional results, including law of large numbers and mixing properties, are extended to this diffusion-driven setting. Building on these results, we study parametric inference for the Hawkes component: we derive consistency and asymptotic normality of the maximum likelihood estimator in the long-time regime, and derive stronger convergence results under additional assumptions on the covariate process. We further propose hypothesis testing procedures to assess the statistical relevance of the covariate. Simulation studies illustrate the validity of the asymptotic results and the effectiveness of the proposed inference methods. Overall, this work provides theoretical and practical foundations for diffusion-driven Hawkes models.

math.ST

Spatial constraints improve filtering of measurement noise from animal tracks

Advances in tracking technologies for animal movement require new statistical tools to better exploit the increasing amount of data. Animal positions are usually calculated using the GPS or Argos satellite system and include potentially non-Gaussian and heavy-tailed measurement error patterns. Errors are usually handled through a Kalman filter algorithm, which can be sensitive to non-Gaussian error distributions. We introduce a latent movement model through an underdamped Langevin stochastic differential equation (SDE) that includes an additional drift term to ensure that the animal remains in a known spatial domain of interest. This can be applied to aquatic animals moving in water or terrestrial animals moving in a restricted zone delimited by fences or natural barriers. We demonstrate that the incorporation of these spatial constraints into the latent movement model can improve the accuracy of filtering for noisy observations of the positions. We implement an Extended Kalman Filter as well as a particle filter adapted to non-Gaussian error distributions. Our filters are based on solving the SDE through splitting schemes to approximate the latent dynamic. We illustrate the approach on a real Argos telemetry track of a bowhead whale in Foxe Basin, Canada.

stat.AP

A time warping model for seasonal data with application to age estimation from narwhal tusks

Signals with varying periodicity frequently appear in real-world phenomena, necessitating the development of efficient modelling techniques to map the measured nonlinear timeline to linear time. Here we propose a regression model that allows for a representation of periodic and dynamic patterns observed in time series data. The model incorporates a hidden strictly positive stochastic process that represents the instantaneous frequency, allowing the model to adapt and accurately capture varying time scales. A case study focusing on age estimation of narwhal tusks is presented, where cyclic element signals associated with annual growth layer groups are analyzed. We apply the methodology to data from one such tusk collected in West Greenland and use the fitted model to estimate the age of the narwhal. The proposed method is validated using simulated signals with known cycle counts and practical considerations and modelling challenges are discussed in detail. This research contributes to the field of time series analysis, providing a tool and valuable insights for understanding and modeling complex cyclic patterns in diverse domains.

stat.AP

Varying coefficients correlated velocity models in complex landscapes with boundaries applied to narwhal responses to noise exposure

Narwhals in the Arctic are increasingly exposed to human activities that can temporarily or permanently threaten their survival by modifying their behavior. We examine GPS data from a population of narwhals exposed to ship and seismic airgun noise during a controlled experiment in 2018 in the Scoresby Sound fjord system in Southeast Greenland. The fjord system has a complex shore line, restricting the behavioral response options for the narwhals to escape the threats. We propose a new continuous-time correlated velocity model with varying coefficients that includes spatial constraints on movement. To assess the sound exposure effect we compare a baseline model for the movement before exposure to a response model for the movement during exposure. Our model, applied to the narwhal data, suggests increased tortuosity of the trajectories as a consequence of the spatial constraints, and further indicates that sound exposure can disturb narwhal motion up to a couple of tens of kilometers. Specifically, we found an increase in velocity and a decrease in the movement persistence.

stat.AP

Inference for the stochastic FitzHugh-Nagumo model from real action potential data via approximate Bayesian computation

The stochastic FitzHugh-Nagumo (FHN) model is a two-dimensional nonlinear stochastic differential equation with additive degenerate noise, whose first component, the only one observed, describes the membrane voltage evolution of a single neuron. Due to its low-dimensionality, its analytical and numerical tractability and its neuronal interpretation, it has been used as a case study to test the performance of different statistical methods in estimating the underlying model parameters. Existing methods, however, often require complete observations, non-degeneracy of the noise or a complex architecture (e.g., to estimate the transition density of the process, "recovering" the unobserved second component) and they may not (satisfactorily) estimate all model parameters simultaneously. Moreover, these studies lack real data applications for the stochastic FHN model. The proposed method tackles all challenges (non-globally Lipschitz drift, non-explicit solution, lack of available transition density, degeneracy of the noise and partial observations). It is an intuitive and easy-to-implement sequential Monte Carlo approximate Bayesian computation algorithm, which relies on a recent computationally efficient and structure-preserving numerical splitting scheme for synthetic data generation and on summary statistics exploiting the structural properties of the process. All model parameters are successfully estimated from simulated data and, more remarkably, real action potential data of rats. The presented novel real-data fit may broaden the scope and credibility of this classic and widely used neuronal model.

stat.CO

Strang Splitting for Parametric Inference in Second-order Stochastic Differential Equations

We address parameter estimation in second-order stochastic differential equations (SDEs), which are prevalent in physics, biology, and ecology. The second-order SDE is converted to a first-order system by introducing an auxiliary velocity variable, which raises two main challenges. First, the system is hypoelliptic since the noise affects only the velocity, making the Euler-Maruyama estimator ill-conditioned. We propose an estimator based on the Strang splitting scheme to overcome this. Second, since the velocity is rarely observed, we adapt the estimator to partial observations. We present four estimators for complete and partial observations, using the full pseudo-likelihood or only the velocity-based partial pseudo-likelihood. These estimators are intuitive, easy to implement, and computationally fast, and we prove their consistency and asymptotic normality. Our analysis demonstrates that using the full pseudo-likelihood with complete observations reduces the asymptotic variance of the diffusion estimator. With partial observations, the asymptotic variance increases as a result of information loss but remains unaffected by the likelihood choice. However, a numerical study on the Kramers oscillator reveals that using the partial pseudo-likelihood for partial observations yields less biased estimators. We apply our approach to paleoclimate data from the Greenland ice core by fitting the Kramers oscillator model, capturing transitions between metastable states reflecting observed climatic conditions during glacial eras.

stat.ME

Non-asymptotic statistical test of the diffusion coefficient of stochastic differential equations

We develop several statistical tests of the determinant of the diffusion coefficient of a stochastic differential equation, based on discrete observations on a time interval $[0,T]$ sampled with a time step $\Delta$. Our main contribution is to control the test Type I and Type II errors in a non asymptotic setting, i.e. when the number of observations and the time step are fixed. The test statistics are calculated from the process increments. In dimension 1, the density of the test statistic is explicit. In dimension 2, the test statistic has no explicit density but upper and lower bounds are proved. We also propose a multiple testing procedure in dimension greater than 2. Every test is proved to be of a given non-asymptotic level and separability conditions to control their power are also provided. A numerical study illustrates the properties of the tests for stochastic processes with known or estimated drifts.

math.ST

Parameter Estimation in Nonlinear Multivariate Stochastic Differential Equations Based on Splitting Schemes

The likelihood functions for discretely observed nonlinear continuous-time models based on stochastic differential equations are not available except for a few cases. Various parameter estimation techniques have been proposed, each with advantages, disadvantages, and limitations depending on the application. Most applications still use the Euler-Maruyama discretization, despite many proofs of its bias. More sophisticated methods, such as Kessler's Gaussian approximation, Ozaki's Local Linearization, A\"it-Sahalia's Hermite expansions, or MCMC methods, might be complex to implement, do not scale well with increasing model dimension, or can be numerically unstable. We propose two efficient and easy-to-implement likelihood-based estimators based on the Lie-Trotter (LT) and the Strang (S) splitting schemes. We prove that S has $L^p$ convergence rate of order 1, a property already known for LT. We show that the estimators are consistent and asymptotically efficient under the less restrictive one-sided Lipschitz assumption. A numerical study on the 3-dimensional stochastic Lorenz system complements our theoretical findings. The simulation shows that the S estimator performs the best when measured on precision and computational speed compared to the state-of-the-art.

stat.ME

Optimal control for parameter estimation in partially observed hypoelliptic stochastic differential equations

We deal with the problem of parameter estimation in stochastic differential equations (SDEs) in a partially observed framework. We aim to design a method working for both elliptic and hypoelliptic SDEs, the latters being characterized by degenerate diffusion coefficients. This feature often causes the failure of contrast estimators based on Euler Maruyama discretization scheme and dramatically impairs classic stochastic filtering methods used to reconstruct the unobserved states. All of theses issues make the estimation problem in hypoelliptic SDEs difficult to solve. To overcome this, we construct a well-defined cost function no matter the elliptic nature of the SDEs. We also bypass the filtering step by considering a control theory perspective. The unobserved states are estimated by solving deterministic optimal control problems using numerical methods which do not need strong assumptions on the diffusion coefficient conditioning. Numerical simulations made on different partially observed hypoelliptic SDEs reveal our method produces accurate estimate while dramatically reducing the computational price comparing to other methods.

math.OC

A splitting method for SDEs with locally Lipschitz drift: Illustration on the FitzHugh-Nagumo model

In this article, we construct and analyse an explicit numerical splitting method for a class of semi-linear stochastic differential equations (SDEs) with additive noise, where the drift is allowed to grow polynomially and satisfies a global one-sided Lipschitz condition. The method is proved to be mean-square convergent of order 1 and to preserve important structural properties of the SDE. First, it is hypoelliptic in every iteration step. Second, it is geometrically ergodic and has an asymptotically bounded second moment. Third, it preserves oscillatory dynamics, such as amplitudes, frequencies and phases of oscillations, even for large time steps. Our results are illustrated on the stochastic FitzHugh-Nagumo model and compared with known mean-square convergent tamed/truncated variants of the Euler-Maruyama method. The capability of the proposed splitting method to preserve the aforementioned properties may make it applicable within different statistical inference procedures. In contrast, known Euler-Maruyama type methods commonly fail in preserving such properties, yielding ill-conditioned likelihood-based estimation tools or computationally infeasible simulation-based inference algorithms.

math.NA

Parameter estimation and treatment optimization in a stochastic model for immunotherapy of cancer

Adoptive Cell Transfer therapy of cancer is currently in full development and mathematical modeling is playing a critical role in this area. We study a stochastic model developed by Baar et al. in 2015 for modeling immunotherapy against melanoma skin cancer. First, we estimate the parameters of the deterministic limit of the model based on biological data of tumor growth in mice. A Nonlinear Mixed Effects Model is estimated by the Stochastic Approximation Expectation Maximization algorithm. With the estimated parameters, we head back to the stochastic model and calculate the probability that the T cells all get exhausted during the treatment. We show that for some relevant parameter values, an early relapse is due to stochastic fluctuations (complete T cells exhaustion) with a non negligible probability. Then, focusing on the relapse related to the T cell exhaustion, we propose to optimize the treatment plan (treatment doses and restimulation times) by minimizing the T cell exhaustion probability in the parameter estimation ranges.

q-bio.PE

Hypoelliptic diffusions: discretization, filtering and inference from complete and partial observations

The statistical problem of parameter estimation in partially observed hypoelliptic diffusion processes is naturally occurring in many applications. However, due to the noise structure, where the noise components of the different coordinates of the multi-dimensional process operate on different time scales, standard inference tools are ill conditioned. In this paper, we propose to use a higher order scheme to approximate the likelihood, such that the different time scales are appropriately accounted for. We show consistency and asymptotic normality with non-typical convergence rates. When only partial observations are available, we embed the approximation into a filtering algorithm for the unobserved coordinates, and use this as a building block in a Stochastic Approximation Expectation Maximization algorithm. We illustrate on simulated data from three models; the Harmonic Oscillator, the FitzHugh-Nagumo model used to model the membrane potential evolution in neuroscience, and the Synaptic Inhibition and Excitation model used for determination of neuronal synaptic input.

stat.ME

Stochastic Proximal Gradient Algorithms for Penalized Mixed Models

Motivated by penalized likelihood maximization in complex models, we study optimization problems where neither the function to optimize nor its gradient have an explicit expression, but its gradient can be approximated by a Monte Carlo technique. We propose a new algorithm based on a stochastic approximation of the Proximal-Gradient (PG) algorithm. This new algorithm, named Stochastic Approximation PG (SAPG) is the combination of a stochastic gradient descent step which - roughly speaking - computes a smoothed approximation of the past gradient along the iterations, and a proximal step. The choice of the step size and the Monte Carlo batch size for the stochastic gradient descent step in SAPG are discussed. Our convergence results cover the cases of biased and unbiased Monte Carlo approximations. While the convergence analysis of the Monte Carlo-PG is already addressed in the literature (see Atchad\'e et al. [2016]), the convergence analysis of SAPG is new. The two algorithms are compared on a linear mixed effect model as a toy example. A more challenging application is proposed on non-linear mixed effect models in high dimension with a pharmacokinetic data set including genomic covariates. To our best knowledge, our work provides the first convergence result of a numerical method designed to solve penalized Maximum Likelihood in a non-linear mixed effect model.

stat.CO

Multivariate inhomogeneous diffusion models with covariates and mixed effects

Modeling of longitudinal data often requires diffusion models that incorporate overall time-dependent, nonlinear dynamics of multiple components and provide sufficient flexibility for subject-specific modeling. This complexity challenges parameter inference and approximations are inevitable. We propose a method for approximate maximum-likelihood parameter estimation in multivariate time-inhomogeneous diffusions, where subject-specific flexibility is accounted for by incorporation of multidimensional mixed effects and covariates. We consider $N$ multidimensional independent diffusions $X^i = (X^i_t)_{0\leq t\leq T^i}, 1\leq i\leq N$, with common overall model structure and unknown fixed-effects parameter $\mu$. Their dynamics differ by the subject-specific random effect $\phi^i$ in the drift and possibly by (known) covariate information, different initial conditions and observation times and duration. The distribution of $\phi^i$ is parametrized by an unknown $\vartheta$ and $\theta = (\mu, \vartheta)$ is the target of statistical inference. Its maximum likelihood estimator is derived from the continuous-time likelihood. We prove consistency and asymptotic normality of $\hat{\theta}_N$ when the number $N$ of subjects goes to infinity using standard techniques and consider the more general concept of local asymptotic normality for less regular models. The bias induced by time-discretization of sufficient statistics is investigated. We discuss verification of conditions and investigate parameter estimation and hypothesis testing in simulations.

stat.ME

Coupling stochastic EM and Approximate Bayesian Computation for parameter inference in state-space models

We study the class of state-space models and perform maximum likelihood estimation for the model parameters. We consider a stochastic approximation expectation-maximization (SAEM) algorithm to maximize the likelihood function with the novelty of using approximate Bayesian computation (ABC) within SAEM. The task is to provide each iteration of SAEM with a filtered state of the system, and this is achieved using an ABC sampler for the hidden state, based on sequential Monte Carlo (SMC) methodology. It is shown that the resulting SAEM-ABC algorithm can be calibrated to return accurate inference, and in some situations it can outperform a version of SAEM incorporating the bootstrap filter. Two simulation studies are presented, first a nonlinear Gaussian state-space model then a state-space model having dynamics expressed by a stochastic differential equation. Comparisons with iterated filtering for maximum likelihood inference, and Gibbs sampling and particle marginal methods for Bayesian inference are presented.

stat.CO

Parametric estimation of complex mixed models based on meta-model approach

Complex biological processes are usually experimented along time among a collection of individuals. Longitudinal data are then available and the statistical challenge is to better understand the underlying biological mechanisms. The standard statistical approach is mixed-effects model, with regression functions that are now highly-developed to describe precisely the biological processes (solutions of multi-dimensional ordinary differential equations or of partial differential equation). When there is no analytical solution, a classical estimation approach relies on the coupling of a stochastic version of the EM algorithm (SAEM) with a MCMC algorithm. This procedure needs many evaluations of the regression function which is clearly prohibitive when a time-consuming solver is used for computing it. In this work a meta-model relying on a Gaussian process emulator is proposed to replace this regression function. The new source of uncertainty due to this approximation can be incorporated in the model which leads to what is called a mixed meta-model. A control on the distance between the maximum likelihood estimates in this mixed meta-model and the maximum likelihood estimates obtained with the exact mixed model is guaranteed. Eventually, numerical simulations are performed to illustrate the efficiency of this approach.

math.ST

A SAEM Algorithm for Fused Lasso Penalized Non Linear Mixed Effect Models: Application to Group Comparison in Pharmacokinetic

Non linear mixed effect models are classical tools to analyze non linear longitudinal data in many fields such as population Pharmacokinetic. Groups of observations are usually compared by introducing the group affiliations as binary covariates with a reference group that is stated among the groups. This approach is relatively limited as it allows only the comparison of the reference group to the others. In this work, we propose to compare the groups using a penalized likelihood approach. Groups are described by the same structural model but with parameters that are group specific. The likelihood is penalized with a fused lasso penalty that induces sparsity on the differences between groups for both fixed effects and variances of random effects. A penalized Stochastic Approximation EM algorithm is proposed that is coupled to Alternating Direction Method Multipliers to solve the maximization step. An extensive simulation study illustrates the performance of this algorithm when comparing more than two groups. Then the approach is applied to real data from two pharmacokinetic drug-drug interaction trials.

stat.CO