SearcharxivSearch

arXiv subjects

Mark Podolskij

Publications and source records attributed to Mark Podolskij.

At least 19 recordsLinked to original sources

Weighted Nuclear Elastic Net Estimation of (Near-) Low-Rank Drift Matrices in Ornstein-Uhlenbeck Processes

We study estimation of the drift matrix in a continuously observed high-dimensional Ornstein-Uhlenbeck process when the drift is exactly or approximately low rank. In this setting, exact low rank induces non-stable directions and hence a non-ergodic regime, resulting in a poorly conditioned empirical covariance matrix. To address this difficulty, we introduce a Weighted Nuclear Elastic Net Estimator that combines ridge regularization with a nuclear-norm penalty expressed in the empirical likelihood geometry. Under a general diagonalizable spectral framework, we establish oracle inequalities relative to arbitrary low-rank comparison matrices. For near low-rank drifts, the approximation error is naturally measured through the singular-value decay of the drift after weighting by the regularized empirical covariance. The stochastic term is controlled by self-normalized martingale arguments under appropriate choice of the tuning parameter. For a symmetric positive-semidefinite exact low-rank model, we verify the empirical-curvature condition required to translate the weighted bound into a Frobenius-norm bound. With an appropriate choice of tuning parameters, the resulting estimator satisfies, up to a logarithmic factor, the standard rank-$r$ matrix-estimation scaling $r d/T$: specifically, its squared Frobenius error is of order $r d\log(T)/T$ with high probability, under an explicit dimension-horizon condition.

math.ST

Pre-averaging estimators of the ex-post covariance matrix in noisy diffusion models with non-synchronous data

We show how pre-averaging can be applied to the problem of measuring the ex-post covariance of financial asset returns under microstructure noise and non-synchronous trading. A pre-averaged realised covariance is proposed, and we present an asymptotic theory for this new estimator, which can be configured to possess an optimal convergence rate or to ensure positive semi-definite covariance matrix estimates. We also derive a noise-robust Hayashi-Yoshida estimator that can be implemented on the original data without prior alignment of prices. We uncover the finite sample properties of our estimators with simulations and illustrate their practical use on high-frequency equity data.

econ.EM

On covariation estimation for multivariate continuous It\^o semimartingales with noise in non-synchronous observation schemes

This paper presents a Hayashi-Yoshida type estimator for the covariation matrix of continuous It\^o semimartingales observed with noise. The coordinates of the multivariate process are assumed to be observed at highly frequent non-synchronous points. The estimator of the covariation matrix is designed via a certain combination of the local averages and the Hayashi-Yoshida estimator. Our method does not require any synchronization of the observation scheme (as e.g. previous tick method or refreshing time method) and it is robust to some dependence structure of the noise process. We show the associated central limit theorem for the proposed estimator and provide a feasible asymptotic result. Our proofs are based on a blocking technique and a stable convergence theorem for semimartingales. Finally, we show simulation results for the proposed estimator to illustrate its finite sample properties.

econ.EM

Asymptotic theory of range-based multipower variation

In this paper, we present a realized range-based multipower variation theory, which can be used to estimate return variation and draw jump-robust inference about the diffusive volatility component, when a high-frequency record of asset prices is available. The standard range-statistic -- routinely used in financial economics to estimate the variance of securities prices -- is shown to be biased when the price process contains jumps. We outline how the new theory can be applied to remove this bias by constructing a hybrid range-based estimator. Our asymptotic theory also reveals that when high-frequency data are sparsely sampled, as is often done in practice due to the presence of microstructure noise, the range-based multipower variations can produce significant efficiency gains over comparable subsampled return-based estimators. The analysis is supported by a simulation study and we illustrate the practical use of our framework on some recent TAQ equity data.

econ.EM

Fact or friction: Jumps at ultra high frequency

This paper shows that jumps in financial asset prices are often erroneously identified and are, in fact, rare events accounting for a very small proportion of the total price variation. We apply new econometric techniques to a comprehensive set of ultra high-frequency equity and foreign exchange tick data recorded at millisecond precision, allowing us to examine the price evolution at the individual order level. We show that in both theory and practice, traditional measures of jump variation based on lower-frequency data tend to spuriously assign a burst of volatility to the jump component. As a result, the true price variation coming from jumps is overstated. Our estimates based on tick data suggest that the jump variation is an order of magnitude smaller than typical estimates found in the existing literature.

econ.EM

Realized range-based estimation of integrated variance

We provide a set of probabilistic laws for estimating the quadratic variation of continuous semimartingales with realized range-based variance -- a statistic that replaces every squared return of realized variance with a normalized squared range. If the entire sample path of the process is available, and under a set of weak conditions, our statistic is consistent and has a mixed Gaussian limit, whose precision is five times greater than that of realized variance. In practice, of course, inference is drawn from discrete data and true ranges are unobserved, leading to downward bias. We solve this problem to get a consistent, mixed normal estimator, irrespective of non-trading effects. This estimator has varying degrees of efficiency over realized variance, depending on how many observations that are used to construct the high-low. The methodology is applied to TAQ data and compared with realized variance. Our findings suggest that the empirical path of quadratic variation is also estimated better with the realized range-based variance.

econ.EM

Inference from high-frequency data: A subsampling approach

In this paper, we show how to estimate the asymptotic (conditional) covariance matrix, which appears in central limit theorems in high-frequency estimation of asset return volatility. We provide a recipe for the estimation of this matrix by subsampling; an approach that computes rescaled copies of the original statistic based on local stretches of high-frequency data, and then it studies the sampling variation of these. We show that our estimator is consistent both in frictionless markets and models with additive microstructure noise. We derive a rate of convergence for it and are also able to determine an optimal rate for its tuning parameters (e.g., the number of subsamples). Subsampling does not require an extra set of estimators to do inference, which renders it trivial to implement. As a variance-covariance matrix estimator, it has the attractive feature that it is positive semi-definite by construction. Moreover, the subsampler is to some extent automatic, as it does not exploit explicit knowledge about the structure of the asymptotic covariance. It therefore tends to adapt to the problem at hand and be robust against misspecification of the noise process. As such, this paper facilitates assessment of the sampling errors inherent in high-frequency estimation of volatility. We highlight the finite sample properties of the subsampler in a Monte Carlo study, while some initial empirical work demonstrates its use to draw feasible inference about volatility in financial markets.

econ.EM

Is the diurnal pattern sufficient to explain intraday variation in volatility? A nonparametric assessment

In this paper, we propose a nonparametric way to test the hypothesis that time-variation in intraday volatility is caused solely by a deterministic and recurrent diurnal pattern. We assume that noisy high-frequency data from a discretely sampled jump-diffusion process are available. The test is then based on asset returns, which are deflated by the seasonal component and therefore homoskedastic under the null. To construct our test statistic, we extend the concept of pre-averaged bipower variation to a general It\^o semimartingale setting via a truncation device. We prove a central limit theorem for this statistic and construct a positive semi-definite estimator of the asymptotic covariance matrix. The $t$-statistic (after pre-averaging and jump-truncation) diverges in the presence of stochastic volatility and has a standard normal distribution otherwise. We show that replacing the true diurnal factor with a model-free jump- and noise-robust estimator does not affect the asymptotic theory. A Monte Carlo simulation also shows this substitution has no discernable impact in finite samples. The test is, however, distorted by small infinite-activity price jumps. To improve inference, we propose a new bootstrap approach, which leads to almost correctly sized tests of the null hypothesis. We apply the developed framework to a large cross-section of equity high-frequency data and find that the diurnal pattern accounts for a rather significant fraction of intraday variation in volatility, but important sources of heteroskedasticity remain present in the data.

econ.EM

Realised quantile-based estimation of the integrated variance

In this paper, we propose a new jump robust quantile-based realised variance measure of ex-post return variation that can be computed using potentially noisy data. The estimator is consistent for the integrated variance and we present feasible central limit theorems which show that it converges at the best attainable rate and has excellent efficiency. Asymptotically, the quantile-based realised variance is immune to finite activity jumps and outliers in the price series, while in modified form the estimator is applicable with market microstructure noise and therefore operational on high-frequency data. Simulations show that it has superior robustness properties in finite sample, while an empirical application illustrates its use on equity data.

econ.EM

Local asymptotic normality for discretely observed McKean-Vlasov diffusions

We study the local asymptotic normality (LAN) property for the likelihood function associated with discretely observed $d$-dimensional McKean-Vlasov stochastic differential equations over a fixed time interval. The model involves a joint parameter in both the drift and diffusion coefficients, introducing challenges due to its dependence on the process distribution. We derive a stochastic expansion of the log-likelihood ratio using Malliavin calculus techniques and establish the LAN property under appropriate conditions. The main technical challenge arises from the implicit nature of the transition densities, which we address through integration by parts and Gaussian-type bounds. This work extends existing LAN results for interacting particle systems to the mean-field regime, contributing to statistical inference in non-linear stochastic models

math.ST

On goodness-of-fit testing for volatility in McKean-Vlasov models

This paper develops a statistical framework for goodness-of-fit testing of volatility functions in i.i.d. McKean-Vlasov stochastic differential equations, which model large diffusion systems with distribution-dependent dynamics. Although integrated volatility estimation for classical SDEs is well established, formal model validation and goodness-of-fit testing for McKean-Vlasov systems remain largely unexplored, particularly in settings combining large particle limits with high-frequency observations. We propose a goodness-of-fit test based on discrete observations of a particle system and study its asymptotic properties in a joint regime where both the number of particles and the sampling frequency tend to infinity. We establish asymptotic normality for the relevant volatility estimators and derive a central limit theorem for the resulting test statistic. These results provide a rigorous basis for assessing volatility specifications in high-dimensional mean-field diffusion models.

stat.ME

Consistent support recovery for high-dimensional diffusions

Statistical inference for stochastic processes has advanced significantly due to applications in diverse fields, but challenges remain in high-dimensional settings where parameters are allowed to grow with the sample size. This paper analyzes a d-dimensional ergodic diffusion process under sparsity constraints, focusing on the adaptive Lasso estimator, which improves variable selection and bias over the standard Lasso. We derive conditions under which the adaptive Lasso achieves support recovery property and asymptotic normality for the drift parameter, with a focus on linear models. Explicit parameter relationships guide tuning for optimal performance, and a marginal estimator is proposed for p>>d scenarios under partial orthogonality assumption. Numerical studies confirm the adaptive Lasso's superiority over standard Lasso and MLE in accuracy and support recovery, providing robust solutions for high-dimensional stochastic processes.

math.ST

Sampling effects on Lasso estimation of drift functions in high-dimensional diffusion processes

In this paper, we address high-dimensional parametric estimation of the drift function in diffusion models, specifically focusing on a $d$-dimensional ergodic diffusion process observed at discrete time points. We consider both a general linear form for the drift function and the particular case of the Ornstein-Uhlenbeck (OU) process. Assuming sparsity of the parameter vector, we examine the statistical behavior of the Lasso estimator for the unknown parameter. Our primary contribution is the proof of an oracle inequality for the Lasso estimator, which holds on the intersection of three specific sets defined for our analysis. We carefully control the probability of these sets, tackling the central challenge of our study. This approach allows us to derive error bounds for the $l_1$ and $l_2$ norms, assessing the performance of the proposed Lasso estimator. Our results demonstrate that, under certain conditions, the discretization error becomes negligible, enabling us to achieve the same optimal rate of convergence as if the continuous trajectory of the process were observed. We validate our theoretical findings through numerical experiments, which show that the Lasso estimator significantly outperforms the maximum likelihood estimator (MLE) in terms of support recovery.

math.ST

On nonparametric estimation of the interaction function in particle system models

This paper delves into a nonparametric estimation approach for the interaction function within diffusion-type particle system models. We introduce two estimation methods based upon an empirical risk minimization. Our study encompasses an analysis of the stochastic and approximation errors associated with both procedures, along with an examination of certain minimax lower bounds. In particular, we show that there is a natural metric under which the corresponding minimax estimation error of the interaction function converges to zero with parametric rate. This result is rather suprising given complexity of the underlying estimation problem and rather large classes of interaction functions for which the above parametric rate holds.

math.ST

Optimal estimation of local time and occupation time measure for an α-stable Levy process

We present a novel theoretical result on estimation of local time and occupation time measure of an α-stable Lévy process with α in (1, 2). Our approach is based upon computing the conditional expectation of the desired quantities given high frequency data, which is an L^2-optimal statistic by construction. We prove the corresponding stable central limit theorems and discuss a statistical application. In particular, this work extends the results of [Ivanovs and i Podolskij (2021)], which investigated the case of the Brownian motion.

math.PR

Polynomial rates via deconvolution for nonparametric estimation in McKean-Vlasov SDEs

This paper investigates the estimation of the interaction function for a class of McKean-Vlasov stochastic differential equations. The estimation is based on observations of the associated particle system at time $T$, considering the scenario where both the time horizon $T$ and the number of particles $N$ tend to infinity. Our proposed method recovers polynomial rates of convergence for the resulting estimator. This is achieved under the assumption of exponentially decaying tails for the interaction function. Additionally, we conduct a thorough analysis of the transform of the associated invariant density as a complex function, providing essential insights for our main results.

math.ST

Limit theorems for general functionals of Brownian local times

In this paper, we present the asymptotic theory for integrated functions of increments of Brownian local times in space. Specifically, we determine their first-order limit, along with the asymptotic distribution of the fluctuations. Our key result establishes that a standardized version of our statistic converges stably in law towards a mixed normal distribution. Our contribution builds upon a series of prior works by S. Campese, X. Chen, Y. Hu, W.V. Li, M.B. Markus, D. Nualart and J. Rosen \cite{C17,CLMR10,HN09,HN10,MR08,R11,R11b}, which delved into special cases of the considered problem, such as quadratic, cubic and polynomial cases. We establish the limit theorem for general functions that satisfy mild smoothness and growth conditions. This extends the scope beyond the polynomial cases studied in previous works, providing a more comprehensive understanding of the asymptotic properties of the considered functionals.

math.PR

Parameter estimation of discretely observed interacting particle systems

In this paper, we consider the problem of joint parameter estimation for drift and diffusion coefficients of a stochastic McKean-Vlasov equation and for the associated system of interacting particles. The analysis is provided in a general framework, as both coefficients depend on the solution of the process and on the law of the solution itself. Starting from discrete observations of the interacting particle system over a fixed interval $[0, T]$, we propose a contrast function based on a pseudo likelihood approach. We show that the associated estimator is consistent when the discretization step ($Δ_n$) and the number of particles ($N$) satisfy $Δ_n \rightarrow 0$ and $N \rightarrow \infty$, and asymptotically normal when additionally the condition $Δ_n N \rightarrow 0$ holds.

math.ST