SearcharxivSearch

arXiv subjects

Stilian Stoev

Publications and source records attributed to Stilian Stoev.

At least 19 recordsLinked to original sources

On the optimal prediction of extreme events

The prediction of the extremely large values of a response variable $Y$ in terms of a vector of covariates $X=(X_i)_{i=1}^d$ is a fundamental problem arising in many scientific and engineering domains. The scarcity of data in the extremes makes the optimal solution of this problem of particular importance. The optimal predictors of such events can be explicitly characterized in just a few cases and it is of fundamental practical and theoretical interest to develop optimal estimators over large classes of models and predictors. In this work, the focus is on the case where $(Y,X)$ have a multivariate regularly varying distribution and one seeks an optimal predictor expressed as a positive homogeneous function $h(X)$ of the covariates. The asymptotic prediction precision in this setting coincides with the tail-dependence coefficient $\lambda(Y,h(X))$ and it can be expressed as an integral functional of the associated angular measure of $(Y,X)$. Thus, finding asymptotically optimal homogeneous predictors amounts to solving a variational problem. We obtain a general solution to this problem, which is expressed in terms of a non-extreme conditional quantile of a tilted distribution derived from the angular measure. This leads to a general inference methodology for the optimal predictors in the peaks-over-threshold framework form extreme value theory. We establish the universal consistency for these estimators over large classes of angular measures. A general-purpose implementation of the resulting inference procedure is shown to work remarkably well against optimal oracle estimators, as well as in the challenging problem of extreme solar flare prediction.

math.ST

On the universal calibration of heavy-tailed combination tests

It is often of interest to test a global null hypothesis using multiple, possibly dependent $p$-values by combining their strengths while controlling the type-I error. Recently, several heavy-tailed combination tests, such as the harmonic mean test and the Cauchy combination test, have been proposed: they transform $p$-values into heavy-tailed random variables before combining them into a single test statistic. The resulting tests, which are calibrated under some form of independence assumption among the $p$-values, have been shown to be rather robust to dependence asymptotically as the $\alpha$ level gets small. Yet, it has remained an open problem to understand this general phenomenon and characterize how such tests behave under dependence. Using the framework of multivariate regular variation from extreme value theory, we show that for a class of combination tests that are homogeneous, the asymptotic level of the test can be expressed using the angular measure under multivariate regular variation. This measure characterizes the dependence of the transformed heavy-tailed variables in their upper tails, or equivalently, the dependence of the $p$-values near zero. We use this result to study several tests. The harmonic mean test, which coincides with the Pareto linear combination test, is shown to be universally calibrated regardless of the tail dependence; further, this test is shown to be the only one that achieves universal calibration among all homogeneous heavy-tailed combination tests. In contrast, the Cauchy combination test is shown to be universally honest but often conservative; the Dunn-\v{S}id\'ak correction, also known as Tippett's method, while being honest, is calibrated if and only if the underlying $p$-values are independent near zero. These theoretical findings are corroborated with simulations and an application to independence testing with survey data.

math.ST

A functional regression model for heterogeneous BioGeoChemical Argo data in the Southern Ocean

Leveraging available measurements of our environment can help us understand complex processes. One example is Argo Biogeochemical data, which aims to collect measurements of oxygen, nitrate, pH, and other variables at varying depths in the ocean. We focus on the oxygen data in the Southern Ocean, which has implications for ocean biology and the Earth's carbon cycle. Systematic monitoring of such data has only recently begun to be established, and the data is sparse. In contrast, Argo measurements of temperature and salinity are much more abundant. In this work, we introduce and estimate a functional regression model describing dependence in oxygen, temperature, and salinity data at all depths covered by the Argo data simultaneously. Our model elucidates important aspects of the joint distribution of temperature, salinity, and oxygen. Due to fronts that establish distinct spatial zones in the Southern Ocean, we augment this functional regression model with a mixture component. By modelling spatial dependence in the mixture component and in the data itself, we provide predictions onto a grid and improve location estimates of fronts. Our approach is scalable to the size of the Argo data, and we demonstrate its success in cross-validation and a comprehensive interpretation of the model.

stat.ME

On the optimal prediction of extreme events in heavy-tailed time series with applications to solar flare forecasting

The prediction of extreme events in time series is a fundamental problem arising in many financial, scientific, engineering, and other applications. We begin by establishing a general Neyman-Pearson-type characterization of optimal extreme event predictors in terms of density ratios. This yields new insights and several closed-form optimal extreme event predictors for additive models. These results naturally extend to time series, where we study optimal extreme event prediction for both light- and heavy-tailed autoregressive and moving average models. Using a uniform law of large numbers for ergodic time series, we establish the asymptotic optimality of an empirical version of the optimal predictor for autoregressive models. Using multivariate regular variation, we obtain an expression for the optimal extremal precision in heavy-tailed infinite moving averages, which provides theoretical bounds on the ability to predict extremes in this general class of models. We address the important problem of predicting solar flares by applying our theory and methodology to a state-of-the-art time series consisting of solar soft X-ray flux measurements. Our results demonstrate the success and limitations in solar flare forecasting of long-memory autoregressive models and long-range-dependent, heavy-tailed FARIMA models.

math.ST

Multivariate Matérn Models -- A Spectral Approach

The classical Matérn model has been a staple in spatial statistics. Novel data-rich applications in environmental and physical sciences, however, call for new, flexible vector-valued spatial and space-time models. Therefore, the extension of the classical Matérn model has been a problem of active theoretical and methodological interest. In this paper, we offer a new perspective to extending the Matérn covariance model to the vector-valued setting. We adopt a spectral, stochastic integral approach, which allows us to address challenging issues on the validity of the covariance structure and at the same time to obtain new, flexible, and interpretable models. In particular, our multivariate extensions of the Matérn model allow for asymmetric covariance structures. Moreover, the spectral approach provides an essentially complete flexibility in modeling the local structure of the process. We establish closed-form representations of the cross-covariances when available, compare them with existing models, simulate Gaussian instances of these new processes, and demonstrate estimation of the model's parameters through maximum likelihood. An application of the new class of multivariate Matérn models to environmental data indicate their success in capturing inherent covariance-asymmetry phenomena.

stat.ME

Detection of Sparse Anomalies in High-Dimensional Network Telescope Signals

Network operators and system administrators are increasingly overwhelmed with incessant cyber-security threats ranging from malicious network reconnaissance to attacks such as distributed denial of service and data breaches. A large number of these attacks could be prevented if the network operators were better equipped with threat intelligence information that would allow them to block or throttle nefarious scanning activities. Network telescopes or "darknets" offer a unique window into observing Internet-wide scanners and other malicious entities, and they could offer early warning signals to operators that would be critical for infrastructure protection and/or attack mitigation. A network telescope consists of unused or "dark" IP spaces that serve no users, and solely passively observes any Internet traffic destined to the "telescope sensor" in an attempt to record ubiquitous network scanners, malware that forage for vulnerable devices, and other dubious activities. Hence, monitoring network telescopes for timely detection of coordinated and heavy scanning activities is an important, albeit challenging, task. The challenges mainly arise due to the non-stationarity and the dynamic nature of Internet traffic and, more importantly, the fact that one needs to monitor high-dimensional signals (e.g., all TCP/UDP ports) to search for "sparse" anomalies. We propose statistical methods to address both challenges in an efficient and "online" manner; our work is validated both with synthetic data as well as real-world data from a large network telescope.

cs.CR

Spectral Density Estimation of Function-Valued Spatial Processes

The spectral density function describes the second-order properties of a stationary stochastic process on $\mathbb{R}^d$. This paper considers the nonparametric estimation of the spectral density of a continuous-time stochastic process taking values in a separable Hilbert space. Our estimator is based on kernel smoothing and can be applied to a wide variety of spatial sampling schemes including those in which data are observed at irregular spatial locations. Thus, it finds immediate applications in Spatial Statistics, where irregularly sampled data naturally arise. The rates for the bias and variance of the estimator are obtained under general conditions in a mixed-domain asymptotic setting. When the data are observed on a regular grid, the optimal rate of the estimator matches the minimax rate for the class of covariance functions that decay according to a power law. The asymptotic normality of the spectral density estimator is also established under general conditions for Gaussian Hilbert-space valued processes. Finally, with a view towards practical applications the asymptotic results are specialized to the case of discretely-sampled functional data in a reproducing kernel Hilbert space.

math.ST

Tail-dependence, exceedance sets, and metric embeddings

There are many ways of measuring and modeling tail-dependence in random vectors: from the general framework of multivariate regular variation and the flexible class of max-stable vectors down to simple and concise summary measures like the matrix of bivariate tail-dependence coefficients. This paper starts by providing a review of existing results from a unifying perspective, which highlights connections between extreme value theory and the theory of cuts and metrics. Our approach leads to some new findings in both areas with some applications to current topics in risk management. We begin by using the framework of multivariate regular variation to show that extremal coefficients, or equivalently, the higher-order tail-dependence coefficients of a random vector can simply be understood in terms of random exceedance sets, which allows us to extend the notion of Bernoulli compatibility. In the special but important case of bi-variate tail-dependence, we establish a correspondence between tail-dependence matrices and $L^1$- and $\ell_1$-embeddable finite metric spaces via the spectral distance, which is a metric on the space of jointly $1$-Fréchet random variables. Namely, the coefficients of the cut-decomposition of the spectral distance and of the Tawn-Molchanov max-stable model realizing the corresponding bi-variate extremal dependence coincide. We show that line metrics are rigid and if the spectral distance corresponds to a line metric, the higher order tail-dependence is determined by the bi-variate tail-dependence matrix. Finally, the correspondence between $\ell_1$-embeddable metric spaces and tail-dependence matrices allows us to revisit the realizability problem, i.e. checking whether a given matrix is a valid tail-dependence matrix. We confirm a conjecture of Shyamalkumar & Tao (2020) that this problem is NP-complete.

math.PR

Tangent fields, intrinsic stationarity, and self-similarity (with a supplement on Matheron Theory)

This paper studies the local structure of continuous random fields on $\mathbb R^d$ taking values in a complete separable linear metric space ${\mathbb V}$. Extending seminal work of Falconer, we show that the generalized $(1+k)$-th order increment tangent fields are self-similar and almost everywhere intrinsically stationary in the sense of Matheron. These results motivate the further study of the structure of ${\mathbb V}$-valued intrinsic random functions of order $k$ (IRF$_k$,\ $k=0,1,\cdots$). To this end, we focus on the special case where ${\mathbb V}$ is a Hilbert space. Building on the work of Sasvari and Berschneider, we establish the spectral characterization of all second order ${\mathbb V}$-valued IRF$_k$'s, extending the classical Matheron theory. Using these results, we further characterize the class of Gaussian, operator self-similar ${\mathbb V}$-valued IRF$_k$'s, generalizing results of Dobrushin and Didier, Meerschaert and Pipiras, among others. These processes are the Hilbert-space-valued versions of the general $k$-th order operator fractional Brownian fields and are characterized by their self-similarity operator exponent as well as a finite trace class operator valued spectral measure. We conclude with several examples motivating future applications to probability and statistics. In a technical Supplement of independent interest, we provide a unified treatment of the Matheron spectral theory for second-order stationary and intrinsically stationary processes taking values in a separable Hilbert space. We give the proofs of the Bochner-Neeb and Bochner-Schwartz theorems.

math.PR

A functional-data approach to the Argo data

The Argo data is a modern oceanography dataset that provides unprecedented global coverage of temperature and salinity measurements in the upper 2,000 meters of depth of the ocean. We study the Argo data from the perspective of functional data analysis (FDA). We develop spatio-temporal functional kriging methodology for mean and covariance estimation to predict temperature and salinity at a fixed location as a smooth function of depth. By combining tools from FDA and spatial statistics, including smoothing splines, local regression, and multivariate spatial modeling and prediction, our approach provides advantages over current methodology that consider pointwise estimation at fixed depths. Our approach naturally leverages the irregularly-sampled data in space, time, and depth to fit a space-time functional model for temperature and salinity. The developed framework provides new tools to address fundamental scientific problems involving the entire upper water column of the oceans such as the estimation of ocean heat content, stratification, and thermohaline oscillation. For example, we show that our functional approach yields more accurate ocean heat content estimates than ones based on discrete integral approximations in pressure. Further, using the derivative function estimates, we obtain a new product of a global map of the mixed layer depth, a key component in the study of heat absorption and nutrient circulation in the oceans. The derivative estimates also reveal evidence for density inversions in areas distinguished by mixing of particularly different water masses.

stat.AP

On the rate of concentration of maxima in Gaussian arrays

Recently in Gao and Stoev (2018) it was established that the concentration of maxima phenomenon is the key to solving the exact sparse support recovery problem in high dimensions. This phenomenon, known also as relative stability, has been little studied in the context of dependence. Here, we obtain bounds on the rate of concentration of maxima in Gaussian triangular arrays. These results are used to establish sufficient conditions for the uniform relative stability of functions of Gaussian arrays, leading to new models that exhibit phase transitions in the exact support recovery problem. Finally, the optimal rate of concentration for Gaussian arrays is studied under more general assumptions than the ones implied by the classic condition of Berman (1964).

math.ST

Distributionally Robust Inference for Extreme Value-at-Risk

Under general multivariate regular variation conditions, the extreme Value-at-Risk of a portfolio can be expressed as an integral of a known kernel with respect to a generally unknown spectral measure supported on the unit simplex. The estimation of the spectral measure is challenging in practice and virtually impossible in high dimensions. This motivates the problem studied in this work, which is to find universal lower and upper bounds of the extreme Value-at-Risk under practically estimable constraints. That is, we study the infimum and supremum of the extreme Value-at-Risk functional, over the infinite dimensional space of all possible spectral measures that meet a finite set of constraints. We focus on extremal coefficient constraints, which are popular and easy to interpret in practice. Our contributions are twofold. Firstly, we show that optimization problems over an infinite dimensional space of spectral measures are in fact dual problems to linear semi-infinite programs (LSIPs) -- linear optimization problems in an Euclidean space with an uncountable set of linear constraints. This allows us to prove that the optimal solutions are in fact attained by discrete spectral measures supported on finitely many atoms. Second, in the case of balanced portfolia, we establish further structural results for the lower bounds as well as closed form solutions for both the lower- and upper-bounds of extreme Value-at-Risk in the special case of a single extremal coefficient constraint. The solutions unveil important connections to the Tawn-Molchanov max-stable models. The results are illustrated with a real data example.

math.ST

Fundamental Limits of Exact Support Recovery in High Dimensions

We study the support recovery problem for a high-dimensional signal observed with additive noise. With suitable parametrization of the signal sparsity and magnitude of its non-zero components, we characterize a phase-transition phenomenon akin to the signal detection problem studied by Ingster in 1998. Specifically, if the signal magnitude is above the so-called strong classification boundary, we show that several classes of well-known procedures achieve asymptotically perfect support recovery as the dimension goes to infinity. This is so, for a very broad class of error distributions with light, rapidly varying tails which may have arbitrary dependence. Conversely, if the signal is below the boundary, then for a very broad class of error dependence structures, no thresholding estimators (including ones with data-dependent thresholds) can achieve perfect support recovery. The proofs of these results exploit a certain concentration of maxima phenomenon known as relative stability. We provide a complete characterization of the relative stability phenomenon for Gaussian triangular arrays in terms their correlation structure. The proof uses classic Sudakov-Fernique and Slepian lemma arguments along with a curious application of Ramsey's coloring theorem. We note that our study of the strong classification boundary is in a finer, point-wise, rather than minimax, sense. We also establish the Bayes optimality and sub-optimality of thresholding procedures. Consequently, we obtain a minimax-type characterization of the strong classification boundary for errors with log-concave densities.

math.ST

Principal components analysis of regularly varying functions

The paper is concerned with asymptotic properties of the principal components analysis of functional data. The currently available results assume the existence of the fourth moment. We develop analogous results in a setting which does not require this assumption. Instead, we assume that the observed functions are regularly varying. We derive the asymptotic distribution of the sample covariance operator and of the sample functional principal components. We obtain a number of results on the convergence of moments and almost sure convergence. We apply the new theory to establish the consistency of the regression operator in a functional linear model.

math.ST

Exchangeable random partitions from max-infinitely-divisible distributions

The hitting partitions are random partitions that arise from the investigation of so-called hitting scenarios of max-infinitely-divisible (max-i.d.)~distributions. We study a class of max-i.d.~laws with exchangeable hitting partitions obtained by size-biased sampling from the jumps of a Lévy subordinator. We obtain explicit formulae for the distributions of these partitions in the case of the multivariate $α$-logistic and another family of exchangeable max-i.d.\ distributions. Specifically, the hitting partitions for these two cases are shown to coincide with the well-known Poisson--Dirichlet partitions ${\rm PD}(α,0),\ α\in (0,1)$ and ${\rm PD}(0,θ),\ θ>0$.

math.PR

Data-adaptive trimming of the Hill estimator and detection of outliers in the extremes of heavy-tailed data

We introduce a trimmed version of the Hill estimator for the index of a heavy-tailed distribution, which is robust to perturbations in the extreme order statistics. In the ideal Pareto setting, the estimator is essentially finite-sample efficient among all unbiased estimators with a given strict upper break-down point. For general heavy-tailed models, we establish the asymptotic normality of the estimator under second order regular variation conditions and also show it is minimax rate-optimal in the Hall class of distributions. We also develop an automatic, data-driven method for the choice of the trimming parameter which yields a new type of robust estimator that can adapt to the unknown level of contamination in the extremes. This adaptive robustness property makes our estimator particularly appealing and superior to other robust estimators in the setting where the extremes of the data are contaminated. As an important application of the data-driven selection of the trimming parameters, we obtain a methodology for the principled identification of extreme outliers in heavy tailed data. Indeed, the method has been shown to correctly identify the number of outliers in the previously explored Condroz data set.

stat.ME

Trimming the Hill estimator: robustness, optimality and adaptivity

We introduce a trimmed version of the Hill estimator for the index of a heavy-tailed distribution, which is robust to perturbations in the extreme order statistics. In the ideal Pareto setting, the estimator is essentially finite-sample efficient among all unbiased estimators with a given strict upper break-down point. For general heavy-tailed models, we establish the asymptotic normality of the estimator under second order conditions and discuss its minimax optimal rate in the Hall class. We introduce the so-called trimmed Hill plot, which can be used to select the number of top order statistics to trim. We also develop an automatic, data-driven procedure for the choice of trimming. This results in a new type of robust estimator that can {\em adapt} to the unknown level of contamination in the extremes. As a by-product we also obtain a methodology for identifying extreme outliers in heavy tailed data. The competitive performance of the trimmed Hill and adaptive trimmed Hill estimators is illustrated with simulations.

stat.ME

Inference for Monotone Trends Under Dependence

We focus on the problem estimating a monotone trend function under additive and dependent noise. New point-wise confidence interval estimators under both short- and long-range dependent errors are introduced and studied. These intervals are obtained via the method of inversion of certain discrepancy statistics arising in hypothesis testing problems. The advantage of this approach is that it avoids the estimation of nuisance parameters such as the derivative of the unknown function, which existing methods are forced to deal with. While the methodology is motivated by earlier work in the independent context, the dependence of the errors, especially longrange dependence leads to new challenges, such as the study of convex minorants of drifted fractional Brownian motion that may be of independent interest. We also unravel a new family of universal limit distributions (and tabulate selected quantiles) that can henceforth be used for inference in monotone function problems involving dependence.

math.ST