SearcharxivSearch

arXiv subjects

Andrey Pepelyshev

Publications and source records attributed to Andrey Pepelyshev.

At least 19 recordsLinked to original sources

Multidimensional Dickman distribution and operator selfdecomposability

The one-dimensional Dickman distribution arises in various stochastic models across number theory, combinatorics, physics, and biology. Recently, a definition of the multidimensional Dickman distribution has appeared in the literature, together with its application to approximating the small jumps of multidimensional Lévy processes. In this paper, we extend this definition to a class of vector-valued random elements, which we characterise as fixed points of a specific affine transformation involving a random matrix obtained from the matrix exponential of a uniformly distributed random variable. We prove that these new distributions possess the key properties of infinite divisibility and operator selfdecomposability. Furthermore, we identify several cases where this new distribution arises as a limiting distribution.

math.PR

Numerical computation of the Rosenblatt distribution and applications

The Rosenblatt distribution plays a key role in the limit theorems for non-linear functionals of stationary Gaussian processes with long-range dependence. We derive new expressions for the characteristic function of the Rosenblatt distribution. Also we present a novel accurate approximation of all eigenvalues of the Riesz integral operator associated with the correlation function of the Gaussian process and propose an efficient algorithm for computation of the density of the Rosenblatt distribution. We perform Monte-Carlo simulation for small sample sizes to demonstrate the appearance of the Rosenblatt distribution for several functionals of stationary Gaussian processes with long-range dependence.

math.ST

Dickman type stochastic processes with short- and long- range dependence

We study properties of the (generalized) Dickman distribution with two parameters and the stationary solution of the Ornstein-Uhlenbeck stochastic differential equation driven by a Poisson process. In particular, we show that the marginal distribution of this solution is the Dickman distribution. Additionally, we investigate superpositions of Ornstein-Uhlenbeck processes which may have short- or long-range dependencies and marginal distribution of the form of the Dickman distribution. The numerical algorithm for simulation of these processes is presented.

math.PR

Statistical Modelling for Improving Efficiency of Online Advertising

Real-time bidding has transformed the digital advertising landscape, allowing companies to buy website advertising space in a matter of milliseconds in the time it takes a webpage to load. Joint research between Cardiff University and Crimtan has employed statistical modelling in conjunction with machine-learning techniques on big data to develop computer algorithms that can select the most appropriate person to which an ad should be shown. These algorithms have been used to identify suitable bidding strategies for that particular advert in order to make the whole process as profitable as possible for businesses. Crimtan's use of the algorithms have enabled them to improve the service that they offer to clients, save money, make significant efficiency gains and attract new business. This has had a knock-on effect with the clients themselves, who have reported an increase in conversion rates as a result of more targeted, accurate and informed advertising. We have also used mixed Poisson processes for modelling for analysing repeat-buying behaviour of online customers. To make numerical comparisons, we use real data collected by Crimtan in the process of running several recent ad campaigns.

stat.AP

Prediction in regression models with continuous observations

We consider the problem of predicting values of a random process or field satisfying a linear model $y(x)=θ^\top f(x) + \varepsilon(x)$, where errors $\varepsilon(x)$ are correlated. This is a common problem in kriging, where the case of discrete observations is standard. By focussing on the case of continuous observations, we derive expressions for the best linear unbiased predictors and their mean squared error. Our results are also applicable in the case where the derivatives of the process $y$ are available, and either a response or one of its derivatives need to be predicted. The theoretical results are illustrated by several examples in particular for the popular Matérn $3/2$ kernel.

math.ST

Analytic Evaluation of the Fractional Moments for the Quasi-Stationary Distribution of the Shiryaev Martingale on an Interval

We consider the quasi-stationary distribution of the classical Shiryaev diffusion restricted to the interval $[0,A]$ with absorption at a fixed $A>0$. We derive analytically a closed-form formula for the distribution's fractional moment of an {\em arbitrary} given order $s\in\mathbb{R}$; the formula is consistent with that previously found by Polunchenko and Pepelyshev (2018) for the case of $s\in\mathbb{N}$. We also show by virtue of the formula that, if $s<1$, then the $s$-th fractional moment of the quasi-stationary distribution becomes that of the exponential distribution (with mean $1/2$) in the limit as $A\to+\infty$; the limiting exponential distribution is the stationary distribution of the reciprocal of the Shiryaev diffusion.

stat.CO

Best linear unbiased estimators in continuous time regression models

In this paper the problem of best linear unbiased estimation is investigated for continuous-time regression models. We prove several general statements concerning the explicit form of the best linear unbiased estimator (BLUE), in particular when the error process is a smooth process with one or several derivatives of the response process available for construction of the estimators. We derive the explicit form of the BLUE for many specific models including the cases of continuous autoregressive errors of order two and integrated error processes (such as integrated Brownian motion). The results are illustrated by several examples.

stat.ME

Optimal designs for regression models with autoregressive errors structure

In the one-parameter regression model with AR(1) and AR(2) errors we find explicit expressions and a continuous approximation of the optimal discrete design for the signed least square estimator. The results are used to derive the optimal variance of the best linear estimator in the continuous time model and to construct efficient estimators and corresponding optimal designs for finite samples. The resulting procedure (estimator and design) provides nearly the same efficiency as the weighted least squares and its variance is close to the optimal variance in the continuous time model. The results are illustrated by several examples demonstrating the feasibility of our approach.

math.ST

Real-time financial surveillance via quickest change-point detection methods

We consider the problem of efficient financial surveillance aimed at "on-the-go" detection of structural breaks (anomalies) in "live"-monitored financial time series. With the problem approached statistically, viz. as that of multi-cyclic sequential (quickest) change-point detection, we propose a semi-parametric multi-cyclic change-point detection procedure to promptly spot anomalies as they occur in the time series under surveillance. The proposed procedure is a derivative of the likelihood ratio-based Shiryaev-Roberts (SR) procedure; the latter is a quasi-Bayesian surveillance method known to deliver the fastest (in the multi-cyclic sense) speed of detection, whatever be the false alarm frequency. We offer a case study where we first carry out, step by step, statistical analysis of a set of real-world financial data, and then set up and devise (a) the proposed SR-based anomaly-detection procedure and (b) the celebrated Cumulative Sum (CUSUM) chart to detect structural breaks in the data. While both procedures performed well, the proposed SR-derivative, conforming to the intuition, seemed slightly better.

stat.AP

Optimal designs in regression with correlated errors

This paper discusses the problem of determining optimal designs for regression models, when the observations are dependent and taken on an interval. A complete solution of this challenging optimal design problem is given for a broad class of regression models and covariance kernels. We propose a class of estimators which are only slightly more complicated than the ordinary least-squares estimators. We then demonstrate that we can design the experiments, such that asymptotically the new estimators achieve the same precision as the best linear unbiased estimator computed for the whole trajectory of the process. As a by-product we derive explicit expressions for the BLUE in the continuous time model and analytic expressions for the optimal designs in a wide class of regression models. We also demonstrate that for a finite number of observations the precision of the proposed procedure, which includes the estimator and design, is very close to the best achievable. The results are illustrated on a few numerical examples.

stat.ME

Optimal design for linear models with correlated observations

In the common linear regression model the problem of determining optimal designs for least squares estimation is considered in the case where the observations are correlated. A necessary condition for the optimality of a given design is provided, which extends the classical equivalence theory for optimal designs in models with uncorrelated errors to the case of dependent data. If the regression functions are eigenfunctions of an integral operator defined by the covariance kernel, it is shown that the corresponding measure defines a universally optimal design. For several models universally optimal designs can be identified explicitly. In particular, it is proved that the uniform distribution is universally optimal for a class of trigonometric regression models with a broad class of covariance kernels and that the arcsine distribution is universally optimal for the polynomial regression model with correlation structure defined by the logarithmic potential. To the best knowledge of the authors these findings provide the first explicit results on optimal designs for regression models with correlated observations, which are not restricted to the location scale model.

math.ST

Asymptotic optimal designs under long-range dependence error structure

We discuss the optimal design problem in regression models with long-range dependence error structure. Asymptotic optimal designs are derived and it is demonstrated that these designs depend only indirectly on the correlation function. Several examples are investigated to illustrate the theory. Finally, the optimal designs are compared with asymptotic optimal designs which were derived by Bickel and Herzberg [Ann. Statist. 7 (1979) 77--95] for regression models with short-range dependent error.

math.ST

Optimal designs for discriminating between dose-response models in toxicology studies

We consider design issues for toxicology studies when we have a continuous response and the true mean response is only known to be a member of a class of nested models. This class of non-linear models was proposed by toxicologists who were concerned only with estimation problems. We develop robust and efficient designs for model discrimination and for estimating parameters in the selected model at the same time. In particular, we propose designs that maximize the minimum of $D$- or $D_1$-efficiencies over all models in the given class. We show that our optimal designs are efficient for determining an appropriate model from the postulated class, quite efficient for estimating model parameters in the identified model and also robust with respect to model misspecification. To facilitate the use of optimal design ideas in practice, we have also constructed a website that freely enables practitioners to generate a variety of optimal designs for a range of models and also enables them to evaluate the efficiency of any design.

math.ST

Optimal designs for random effect models with correlated errors with applications in population pharmacokinetics

We consider the problem of constructing optimal designs for population pharmacokinetics which use random effect models. It is common practice in the design of experiments in such studies to assume uncorrelated errors for each subject. In the present paper a new approach is introduced to determine efficient designs for nonlinear least squares estimation which addresses the problem of correlation between observations corresponding to the same subject. We use asymptotic arguments to derive optimal design densities, and the designs for finite sample sizes are constructed from the quantiles of the corresponding optimal distribution function. It is demonstrated that compared to the optimal exact designs, whose determination is a hard numerical problem, these designs are very efficient. Alternatively, the designs derived from asymptotic theory could be used as starting designs for the numerical computation of exact optimal designs. Several examples of linear and nonlinear models are presented in order to illustrate the methodology. In particular, it is demonstrated that naively chosen equally spaced designs may lead to less accurate estimation.

stat.AP

Optimal designs for dose-finding experiments in toxicity studies

We construct optimal designs for estimating fetal malformation rate, prenatal death rate and an overall toxicity index in a toxicology study under a broad range of model assumptions. We use Weibull distributions to model these rates and assume that the number of implants depend on the dose level. We study properties of the optimal designs when the intra-litter correlation coefficient depends on the dose levels in different ways. Locally optimal designs are found, along with robustified versions of the designs that are less sensitive to misspecification in the initial values of the model parameters. We also report efficiencies of commonly used designs in toxicological experiments and efficiencies of the proposed optimal designs when the true rates have non-Weibull distributions. Optimal design strategies for finding multiple-objective designs in toxicology studies are outlined as well.

stat.ME