SearcharxivSearch

arXiv subjects

Rafal Kulik

Publications and source records attributed to Rafal Kulik.

At least 19 recordsLinked to original sources

Convergence of Stochastic Gradient Descent with mini-batching and infinite variance

Stochastic gradient descent (SGD) with mini-batching is a standard tool in large-scale optimization, yet its theoretical properties under heavy-tailed gradient noise remain largely unexplored. In this paper we study SGD with increasing batch sizes when the gradient noise belongs to the domain of attraction of an $α$-stable law with $α\in(1,2)$. Building on existing results for the finite-variance regime and for heavy-tailed SGD without batching, we establish three main results. First, we derive $L^p$ moment bounds for the SGD error and show that increasing batch sizes lead to faster convergence rates. In particular, batching enables convergence in probability even for a constant stepsize. Second, we prove that the properly normalized SGD iterates converge in distribution to the stationary law of an Ornstein-Uhlenbeck process driven by an $α$-stable Lévy process. Third, for Polyak-Ruppert averaging we obtain a stable limit theorem with a normalization that explicitly depends on the batch-size schedule.

math.PR

Limit theorems for unbounded cluster functionals of regularly varying time series

A blocks method is used to define clusters of extreme values in stationary time series. The cluster starts at the first large value in the block and ends at the last one. The block cluster measure (the point measure at clusters) encodes different aspects of extremal properties. Its limiting behaviour is handled by vague convergence, hence the set of test functions consists of bounded, shift-invariant functionals that vanish around zero. If unbounded or non shift-invariant functionals are considered, we may obtain convergence at a different rate, depending on the type of the functional and the block size (small vs. large blocks). There are two prominent examples of such functionals: the locations of large jumps and the cluster length. We obtain a comprehensive characterization of the limiting behaviour of the block cluster measure evaluated at such functionals for stationary, regularly varying time series. Once the convergence of the block cluster measure is established, we can proceed with consistency of the empirical cluster measure. Consistency holds in the small and moderate blocks scenario, while fails in the large blocks situation. Next, we continue with weak convergence of the empirical cluster processes. The starting point is the seminal paper by Drees and Rootzen (2010). Under the appropriate uniform integrability condition (related to small blocks) the results in the latter paper are still valid. In the moderate and large blocks scenario, the Drees and Rootzen empirical cluster process diverges, but converges weakly when re-normalized properly.

math.ST

Estimation of cluster functionals for regularly varying time series: runs estimators

Cluster indices describe extremal behaviour of stationary time series. We consider runs estimators of cluster indices. Using a modern theory of multivariate, regularly varying time series, we obtain central limit theorems under conditions that can be easily verified for a large class of models. In particular, we show that blocks and runs estimators have the same limiting variance.

math.ST

Long range dependence of heavy tailed random functions

We introduce a definition of long range dependence of random processes and fields on an (unbounded) index space $T\subseteq \R^d$ in terms of integrability of the covariance of indicators that a random function exceeds any given level. This definition is particularly designed to cover the case of random functions with infinite variance. We show the value of this new definition and its connection to limit theorems on some examples including subordinated Gaussian as well as random volatility fields and time series.

math.PR

Estimation of cluster functionals for regularly varying time series: sliding blocks estimators

Cluster indices describe extremal behaviour of stationary time series. We consider their sliding blocks estimators. Using a modern theory of multivariate, regularly varying time series, we obtain central limit theorems under conditions that can be easily verified for a large class of models. In particular, we show that in the Peak over Threshold framework, sliding and disjoint blocks estimators have the same limiting variance.

math.ST

The tail empirical process of regularly varying functions of geometrically ergodic Markov chains

We consider a stationary regularly varying time series which can be expressedas a function of a geometrically ergodic Markov chain. We obtain practical conditionsfor the weak convergence of the tail array sums and feasible estimators ofcluster statistics. These conditions include the so-called geometric drift or Foster-Lyapunovcondition and can be easily checked for most usual time series models witha Markovian structure. We illustrate these conditions on several models and statisticalapplications. A counterexample is given to show a different limiting behaviorwhen the geometric drift condition is not fulfilled.

math.ST

Statistical inference for heavy tailed series with extremal independence

We consider stationary time series $\{X_j, j \in Z\} whose finite dimensional distributions are regularly varying with extremal independence. We assume that for each $h \geq 1$, conditionally on $X_0$ to exceed a threshold tending to infinity, the conditional distribution of $X_h$ suitably normalized converges weakly to a non degenerate distribution. We consider in this paper the estimation of the normalization and of the limiting distribution.

math.ST

Heavy tailed time series with extremal independence

We consider strictly stationary heavy tailed time series whose finite-dimensional exponent measures are concentrated on axes, and hence their extremal properties cannot be tackled using classical multivariate regular variation that is suitable for time series with extremal dependence. We recover relevant information about limiting behavior of time series with extremal independence by introducing a sequence of scaling functions and conditional scaling exponent. Both quantities provide more information about joint extremes than a widely used tail dependence coefficient. We calculate the scaling functions and the scaling exponent for variety of models, including Markov chains, exponential autoregressive model, stochastic volatility with heavy tailed innovations or volatility.

math.ST

Multichannel Deconvolution with Long Range Dependence: Upper bounds on the $L^p$-risk $(1 \le p < \infty)$

We consider multichannel deconvolution in a periodic setting with long-memory errors under three different scenarios for the convolution operators, i.e., super-smooth, regular-smooth and box-car convolutions. We investigate global performances of linear and hard-thresholded non-linear wavelet estimators for functions over a wide range of Besov spaces and for a variety of loss functions defining the risk. In particular, we obtain upper bounds on convergence rates using the $L^p$-risk $(1 \le p < \infty)$. Contrary to the case where the errors follow independent Brownian motions, it is demonstrated that multichannel deconvolution with errors that follow independent fractional Brownian motions with different Hurst parameters results in a much more involved situation. An extensive finite-sample numerical study is performed to supplement the theoretical findings.

math.ST

Multichannel Deconvolution with Long-Range Dependence: A Minimax Study

We consider the problem of estimating the unknown response function in the multichannel deconvolution model with long-range dependent Gaussian errors. We do not limit our consideration to a specific type of long-range dependence rather we assume that the errors should satisfy a general assumption in terms of the smallest and larger eigenvalues of their covariance matrices. We derive minimax lower bounds for the quadratic risk in the proposed multichannel deconvolution model when the response function is assumed to belong to a Besov ball and the blurring function is assumed to possess some smoothness properties, including both regular-smooth and super-smooth convolutions. Furthermore, we propose an adaptive wavelet estimator of the response function that is asymptotically optimal (in the minimax sense), or near-optimal within a logarithmic factor, in a wide range of Besov balls. It is shown that the optimal convergence rates depend on the balance between the smoothness parameter of the response function, the kernel parameters of the blurring function, the long memory parameters of the errors, and how the total number of observations is distributed among the total number of channels. Some examples of inverse problems in mathematical physics where one needs to recover initial or boundary conditions on the basis of observations from a noisy solution of a partial differential equation are used to illustrate the application of the theory we developed. The optimal convergence rates and the adaptive estimators we consider extend the ones studied by Pensky and Sapatinas (2009, 2010) for independent and identically distributed Gaussian errors to the case of long-range dependent Gaussian errors.

math.ST

Strong approximations for long memory sequences based partial sums, counting and their Vervaat processes

We study the asymptotic behaviour of partial sums of long range dependent random variables and that of their counting process, together with an appropriately normalized integral process of the sum of these two processes, the so-called Vervaat process. The first two of these processes are approximated by an appropriately constructed fractional Brownian motion, while the Vervaat process in turn is approximated by the square of the same fractional Brownian motion.

math.PR

Limit theorems for long memory stochastic volatility models with infinite variance: Partial Sums and Sample Covariances

Long Memory Stochastic volatility (LMSV) models capture two standardized features of financial data: the log-returns are uncorrelated, but their squares, or absolute values are (highly) dependent and they may have heavy tails. EGARCH and related models were introduced to model leverage, i.e. negative dependence between previous returns and future volatility. Limit theorems for partial sums, sample variance and sample covariances are basic tools to investigate the presence of long memory and heavy tails and their consequences. In this paper we extend the existing literature on the asymptotic behaviour of the partial sums and the sample covariances of long memory stochastic volatility models in the case of infinite variance. We also consider models with leverage, for which our results are entirely new in the infinite variance case. Depending on the nterplay between the tail behaviour and the intensity of dependence, wo types of convergence rates and limiting distributions can arise. In articular, we show that the asymptotic behaviour of partial sums is the same for both LMSV and models with leverage, whereas there is a crucial difference when sample covariances are considered.

math.ST

Empirical process of residuals for regression models with long memory errors

We consider the residual empirical process in random design regression with long memory errors. We establish its limiting behaviour, showing that its rates of convergence are different from the rates of convergence for to the empirical process based on (unobserved) errors. Also, we study a residual empirical process with estimated parameters. Its asymptotic distribution can be used to construct Kolmogorov-Smirnov, Cramér-Smirnov-von Mises, or other goodness-of-fit tests. Theoretical results are justified by simulation studies.

math.ST

Some results on random design regression with long memory errors and predictors

This paper studies nonparametric regression with long memory (LRD) errors and predictors. First, we formulate general conditions which guarantee the standard rate of convergence for a nonparametric kernel estimator. Second, we calculate the Mean Integrated Squared Error (MISE). In particular, we show that LRD of errors may influence MISE. On the other hand, an estimator for a shape function is typically not influenced by LRD in errors. Finally, we investigate properties of a data-driven bandwidth choice. We show that Averaged Squared Error (ASE) is a good approximation of MISE, however, this is not the case for a cross-validation criterion.

math.ST

Tail behaviour of the area under a random process, with applications to queueing systems, insurance and percolations

The areas under workload process and under queuing process in a single server queue over the busy period have many applications not only in queuing theory but also in risk theory or percolation theory. We focus here on the tail behaviour of distribution of these two integrals. We present various open problems and conjectures, which are supported by partial results for some special cases.

math.PR

The tail empirical process for some long memory sequences

This paper describes limiting behaviour of tail empirical process associated with long memory stochastic volatility models. We show that such process has dichotomous behaviour, according to an interplay between a Hurst parameter and a tail index. In particular, the limit may be non-Gaussian and/or degenerate, indicating an influence of long memory. On the other hand, tail empirical process with random levels never suffers from long memory. This is very desirable from a practical point of view, since such the process may be used to construct Hill estimator of the tail index. To prove our results we need to establish several new results for regularly varying distribution functions, which may be of independent interest.

math.ST

Kink estimation in stochastic regression with dependent errors and predictors

In this article we study the estimation of the location of jump points in the first derivative (referred to as kinks) of a regression function μin two random design models with different long-range dependent (LRD) structures. The method is based on the zero-crossing technique and makes use of high-order kernels. The rate of convergence of the estimator is contingent on the level of dependence and the smoothness of the regression function μ. In one of the models, the convergence rate is the same as the minimax rate for kink estimation in the fixed design scenario with i.i.d. errors which suggests that the method is optimal in the minimax sense.

math.ST

Reduction principles for quantile and Bahadur-Kiefer processes of long-range dependent linear sequences

In this paper we consider quantile and Bahadur-Kiefer processes for long range dependent linear sequences. These processes, unlike in previous studies, are considered on the whole interval $(0,1)$. As it is well-known, quantile processes can have very erratic behavior on the tails. We overcome this problem by considering these processes with appropriate weight functions. In this way we conclude strong approximations that yield some remarkable phenomena that are not shared with i.i.d. sequences, including weak convergence of the Bahadur-Kiefer processes, a different pointwise behavior of the general and uniform Bahadur-Kiefer processes, and a somewhat "strange" behavior of the general quantile process.

math.ST