SearcharxivSearch

arXiv subjects

Nestor Parolya

Publications and source records attributed to Nestor Parolya.

At least 19 recordsLinked to original sources

Consistent Estimation of the High-Dimensional Efficient Frontier

In this paper, we analyze the asymptotic behavior of the main characteristics of the mean-variance efficient frontier employing random matrix theory. Our particular interest covers the case when the dimension $p$ and the sample size $n$ tend to infinity simultaneously and their ratio $p/n$ tends to a positive constant $c\in(0,1)$. We neither impose any distributional nor structural assumptions on the asset returns. For the developed theoretical framework, some regularity conditions, like the existence of the $4$th moments, are needed. It is shown that two out of three quantities of interest are biased and overestimated by their sample counterparts under the high-dimensional asymptotic regime. This becomes evident based on the asymptotic deterministic equivalents of the sample plug-in estimators. Using them we construct consistent estimators of the three characteristics of the efficient frontier. It it shown that the additive and/or the multiplicative biases of the sample estimates are solely functions of the concentration ratio $c$. Furthermore, the asymptotic normality of the considered estimators of the parameters of the efficient frontier is proved. Verifying the theoretical results based on an extensive simulation study we show that the proposed estimator for the efficient frontier is a valuable alternative to the sample estimator for high dimensional data. Finally, we present an empirical application, where we estimate the efficient frontier based on the stocks included in S\&P 500 index.

q-fin.ST

Reviving pseudo-inverses: Asymptotic properties of large dimensional Moore-Penrose and Ridge-type inverses with applications

In this paper, we derive high-dimensional asymptotic properties of the Moore-Penrose inverse and, as a byproduct, of various ridge-type inverses of the sample covariance matrix. In particular, the analytical expressions of the asymptotic behavior of the weighted sample trace moments of generalized inverse matrices are deduced in terms of the partial exponential Bell polynomials which can be easily computed in practice. The existent results for pseudo-inverses are extended in several directions: (i) First, the population covariance matrix is not assumed to be a multiple of the identity matrix; (ii) Second, the assumption of normality is not used in the derivation; (iii) Third, the asymptotic results are derived under the high-dimensional asymptotic regime. Our findings provide universal methodology for construction of fully data-driven improved shrinkage estimators of the precision matrix, optimal portfolio weights and beyond. It is found that the Moore-Penrose inverse acts asymptotically as a certain regularizer of the true covariance matrix and it seems that its proper transformation (shrinkage) performs similarly to or even outperforms the existing benchmarks in many applications, while keeping the computational time as minimal as possible.

math.ST

Linear quasi-shrinkage estimator for high-dimensional optimization with linear constraints

In large-scale data-driven optimization problems, parameters are often only known approximately due to noisy and small-sized samples. We consider optimization problems with linear constraints where the true parameter matrix is not precisely known, and the number of constraints and variables are comparable and large. Our goal is to construct a linear estimator of the true parameter matrix by minimizing the Frobenius distance between the estimator and the true parameter matrix. Our method offers three key advantages: 1) the coefficients of the linear estimator are consistently estimated from the observations and require no further calibration; 2) it only requires the sample size to be greater than one and it delivers stable performance across varied sample sizes; and 3) the constraints of the formulated optimization problem using the estimator remain linear, ensuring computational efficiency when the number of constraints and variables are large. Simulation shows that our linear estimator consistently produces stable outcomes in terms of the objective value, the ratio of violated constraints and the magnitude of constraint violation across various scenarios, compared to the nominal and robust methods. Additionally, it demonstrates resilience against high levels of noise, making it a robust choice under uncertainty.

math.OC

Sampling Distributions of Optimal Portfolio Weights and Characteristics in Low and Large Dimensions

Optimal portfolio selection problems are determined by the (unknown) parameters of the data generating process. If an investor wants to realise the position suggested by the optimal portfolios, he/she needs to estimate the unknown parameters and to account for the parameter uncertainty in the decision process. Most often, the parameters of interest are the population mean vector and the population covariance matrix of the asset return distribution. In this paper, we characterise the exact sampling distribution of the estimated optimal portfolio weights and their characteristics. This is done by deriving their sampling distribution by its stochastic representation. This approach possesses several advantages, {e.g.} (i) it determines the sampling distribution of the estimated optimal portfolio weights by expressions, which could be used to draw samples from this distribution efficiently; (ii) the application of the derived stochastic representation provides an easy way to obtain the asymptotic approximation of the sampling distribution. The later property is used to show that the high-dimensional asymptotic distribution of optimal portfolio weights is a multivariate normal and to determine its parameters. Moreover, a consistent estimator of optimal portfolio weights and their characteristics is derived under the high-dimensional settings. Via an extensive simulation study, we investigate the finite-sample performance of the derived asymptotic approximation and study its robustness to the violation of the model assumptions used in the derivation of the theoretical results.

q-fin.PM

Logarithmic law of large random correlation matrices

Consider a random vector $\mathbf{y}=\mathbfΣ^{1/2}\mathbf{x}$, where the $p$ elements of the vector $\mathbf{x}$ are i.i.d. real-valued random variables with zero mean and finite fourth moment, and $\mathbfΣ^{1/2}$ is a deterministic $p\times p$ matrix such that the spectral norm of the population correlation matrix $\mathbf{R}$ of $\mathbf{y}$ is uniformly bounded. In this paper, we find that the log determinant of the sample correlation matrix $\hat{\mathbf{R}}$ based on a sample of size $n$ from the distribution of $\mathbf{y}$ satisfies a CLT (central limit theorem) for $p/n\to γ\in (0, 1]$ and $p\leq n$. Explicit formulas for the asymptotic mean and variance are provided. In case the mean of $\mathbf{y}$ is unknown, we show that after recentering by the empirical mean the obtained CLT holds with a shift in the asymptotic mean. This result is of independent interest in both large dimensional random matrix theory and high-dimensional statistical literature of large sample correlation matrices for non-normal data. At last, the obtained findings are applied for testing of uncorrelatedness of $p$ random variables. Surprisingly, in the null case $\mathbf{R}=\mathbf{I}$, the test statistic becomes completely pivotal and the extensive simulations show that the obtained CLT also holds if the moments of order four do not exist at all, which conjectures a promising and robust test statistic for heavy-tailed high-dimensional data.

math.ST

Log determinant of large correlation matrices under infinite fourth moment

In this paper, we show the central limit theorem for the logarithmic determinant of the sample correlation matrix $\mathbf{R}$ constructed from the $(p\times n)$-dimensional data matrix $\mathbf{X}$ containing independent and identically distributed random entries with mean zero, variance one and infinite fourth moments. Precisely, we show that for $p/n\to γ\in (0,1)$ as $n,p\to \infty$ the logarithmic law \begin{equation*} \frac{\log \det \mathbf{R} -(p-n+\frac{1}{2})\log(1-p/n)+p-p/n}{\sqrt{-2\log(1-p/n)- 2 p/n}} \overset{d}{\rightarrow} N(0,1)\, \end{equation*} is still valid if the entries of the data matrix $\mathbf{X}$ follow a symmetric distribution with a regularly varying tail of index $α\in (3,4)$. The latter assumptions seem to be crucial, which is justified by the simulations: if the entries of $\mathbf{X}$ have the infinite absolute third moment and/or their distribution is not symmetric, the logarithmic law is not valid anymore. The derived results highlight that the logarithmic determinant of the sample correlation matrix is a very stable and flexible statistic for heavy-tailed big data and open a novel way of analysis of high-dimensional random matrices with self-normalized entries.

math.PR

Two is better than one: Regularized shrinkage of large minimum variance portfolio

In this paper we construct a shrinkage estimator of the global minimum variance (GMV) portfolio by a combination of two techniques: Tikhonov regularization and direct shrinkage of portfolio weights. More specifically, we employ a double shrinkage approach, where the covariance matrix and portfolio weights are shrunk simultaneously. The ridge parameter controls the stability of the covariance matrix, while the portfolio shrinkage intensity shrinks the regularized portfolio weights to a predefined target. Both parameters simultaneously minimize with probability one the out-of-sample variance as the number of assets $p$ and the sample size $n$ tend to infinity, while their ratio $p/n$ tends to a constant $c>0$. This method can also be seen as the optimal combination of the well-established linear shrinkage approach of Ledoit and Wolf (2004, JMVA) and the shrinkage of the portfolio weights by Bodnar et al. (2018, EJOR). No specific distribution is assumed for the asset returns except of the assumption of finite $4+\varepsilon$ moments. The performance of the double shrinkage estimator is investigated via extensive simulation and empirical studies. The suggested method significantly outperforms its predecessor (without regularization) and the nonlinear shrinkage approach in terms of the out-of-sample variance, Sharpe ratio and other empirical measures in the majority of scenarios. Moreover, it obeys the most stable portfolio weights with uniformly smallest turnover.

q-fin.ST

Is the empirical out-of-sample variance an informative risk measure for the high-dimensional portfolios?

The main contribution of this paper is the derivation of the asymptotic behaviour of the out-of-sample variance, the out-of-sample relative loss, and of their empirical counterparts in the high-dimensional setting, i.e., when both ratios $p/n$ and $p/m$ tend to some positive constants as $m\to\infty$ and $n\to\infty$, where $p$ is the portfolio dimension, while $n$ and $m$ are the sample sizes from the in-sample and out-of-sample periods, respectively. The results are obtained for the traditional estimator of the global minimum variance (GMV) portfolio, for the two shrinkage estimators introduced by \cite{frahm2010} and \cite{bodnar2018estimation}, and for the equally-weighted portfolio, which is used as a target portfolio in the specification of the two considered shrinkage estimators. We show that the behaviour of the empirical out-of-sample variance may be misleading is many practical situations. On the other hand, this will never happen with the empirical out-of-sample relative loss, which seems to provide a natural normalization of the out-of-sample variance in the high-dimensional setup. As a result, an important question arises if this risk measure can safely be used in practice for portfolios constructed from a large asset universe.

q-fin.ST

Dynamic Shrinkage Estimation of the High-Dimensional Minimum-Variance Portfolio

In this paper, new results in random matrix theory are derived which allow us to construct a shrinkage estimator of the global minimum variance (GMV) portfolio when the shrinkage target is a random object. More specifically, the shrinkage target is determined as the holding portfolio estimated from previous data. The theoretical findings are applied to develop theory for dynamic estimation of the GMV portfolio, where the new estimator of its weights is shrunk to the holding portfolio at each time of reconstruction. Both cases with and without overlapping samples are considered in the paper. The non-overlapping samples corresponds to the case when different data of the asset returns are used to construct the traditional estimator of the GMV portfolio weights and to determine the target portfolio, while the overlapping case allows intersections between the samples. The theoretical results are derived under weak assumptions imposed on the data-generating process. No specific distribution is assumed for the asset returns except from the assumption of finite $4+\varepsilon$, $\varepsilon>0$, moments. Also, the population covariance matrix with unbounded spectrum can be considered. The performance of new trading strategies is investigated via an extensive simulation. Finally, the theoretical findings are implemented in an empirical illustration based on the returns on stocks included in the S\&P 500 index.

q-fin.ST

Optimal shrinkage-based portfolio selection in high dimensions

In this paper we estimate the mean-variance portfolio in the high-dimensional case using the recent results from the theory of random matrices. We construct a linear shrinkage estimator which is distribution-free and is optimal in the sense of maximizing with probability $1$ the asymptotic out-of-sample expected utility, i.e., mean-variance objective function for different values of risk aversion coefficient which in particular leads to the maximization of the out-of-sample expected utility and to the minimization of the out-of-sample variance. One of the main features of our estimator is the inclusion of the estimation risk related to the sample mean vector into the high-dimensional portfolio optimization. The asymptotic properties of the new estimator are investigated when the number of assets $p$ and the sample size $n$ tend simultaneously to infinity such that $p/n \rightarrow c\in (0,+\infty)$. The results are obtained under weak assumptions imposed on the distribution of the asset returns, namely the existence of the $4+\varepsilon$ moments is only required. Thereafter we perform numerical and empirical studies where the small- and large-sample behavior of the derived estimator is investigated. The suggested estimator shows significant improvements over the existent approaches including the nonlinear shrinkage estimator and the three-fund portfolio rule, especially when the portfolio dimension is larger than the sample size. Moreover, it is robust to deviations from normality.

q-fin.ST

Statistical inference for the EU portfolio in high dimensions

In this paper, using the shrinkage-based approach for portfolio weights and modern results from random matrix theory we construct an effective procedure for testing the efficiency of the expected utility (EU) portfolio and discuss the asymptotic behavior of the proposed test statistic under the high-dimensional asymptotic regime, namely when the number of assets $p$ increases at the same rate as the sample size $n$ such that their ratio $p/n$ approaches a positive constant $c\in(0,1)$ as $n\to\infty$. We provide an extensive simulation study where the power function and receiver operating characteristic curves of the test are analyzed. In the empirical study, the methodology is applied to the returns of S\&P 500 constituents.

q-fin.PM

Spectral analysis of large reflexive generalized inverse and Moore-Penrose inverse matrices

A reflexive generalized inverse and the Moore-Penrose inverse are often confused in statistical literature but in fact they have completely different behaviour in case the population covariance matrix is not a multiple of identity. In this paper, we study the spectral properties of a reflexive generalized inverse and of the Moore-Penrose inverse of the sample covariance matrix. The obtained results are used to assess the difference in the asymptotic behaviour of their eigenvalues.

math.ST

Tests for the weights of the global minimum variance portfolio in a high-dimensional setting

In this study, we construct two tests for the weights of the global minimum variance portfolio (GMVP) in a high-dimensional setting, namely, when the number of assets $p$ depends on the sample size $n$ such that $\frac{p}{n}\to c \in (0,1)$ as $n$ tends to infinity. In the case of a singular covariance matrix with rank equal to $q$ we assume that $q/n\to \tilde{c}\in(0, 1)$ as $n\to\infty$. The considered tests are based on the sample estimator and on the shrinkage estimator of the GMVP weights. We derive the asymptotic distributions of the test statistics under the null and alternative hypotheses. Moreover, we provide a simulation study where the power functions and the receiver operating characteristic curves of the proposed tests are compared with other existing approaches. We observe that the test based on the shrinkage estimator performs well even for values of $c$ close to one.

q-fin.ST

Mean-Variance Efficiency of Optimal Power and Logarithmic Utility Portfolios

We derive new results related to the portfolio choice problem for power and logarithmic utilities. Assuming that the portfolio returns follow an approximate log-normal distribution, the closed-form expressions of the optimal portfolio weights are obtained for both utility functions. Moreover, we prove that both optimal portfolios belong to the set of mean-variance feasible portfolios and establish necessary and sufficient conditions such that they are mean-variance efficient. Furthermore, an application to the stock market is presented and the behavior of the optimal portfolio is discussed for different values of the relative risk aversion coefficient. It turns out that the assumption of log-normality does not seem to be a strong restriction.

q-fin.PM

Testing for Independence of Large Dimensional Vectors

In this paper new tests for the independence of two high-dimensional vectors are investigated. We consider the case where the dimension of the vectors increases with the sample size and propose multivariate analysis of variance-type statistics for the hypothesis of a block diagonal covariance matrix. The asymptotic properties of the new test statistics are investigated under the null hypothesis and the alternative hypothesis using random matrix theory. For this purpose we study the weak convergence of linear spectral statistics of central and (conditionally) non-central Fisher matrices. In particular, a central limit theorem for linear spectral statistics of large dimensional (conditionally) non-central Fisher matrices is derived which is then used to analyse the power of the tests under the alternative. The theoretical results are illustrated by means of a simulation study where we also compare the new tests with several alternative, in particular with the commonly used corrected likelihood ratio test. It is demonstrated that the latter test does not keep its nominal level, if the dimension of one sub-vector is relatively small compared to the dimension of the other sub-vector. On the other hand the tests proposed in this paper provide a reasonable approximation of the nominal level in such situations. Moreover, we observe that one of the proposed tests is most powerful under a variety of correlation scenarios.

math.ST

Optimal Shrinkage Estimator for High-Dimensional Mean Vector

In this paper we derive the optimal linear shrinkage estimator for the high-dimensional mean vector using random matrix theory. The results are obtained under the assumption that both the dimension $p$ and the sample size $n$ tend to infinity in such a way that $p/n \to c\in(0,\infty)$. Under weak conditions imposed on the underlying data generating mechanism, we find the asymptotic equivalents to the optimal shrinkage intensities and estimate them consistently. The proposed nonparametric estimator for the high-dimensional mean vector has a simple structure and is proven to minimize asymptotically, with probability $1$, the quadratic loss when $c\in(0,1)$. When $c\in(1, \infty)$ we modify the estimator by using a feasible estimator for the precision covariance matrix. To this end, an exhaustive simulation study and an application to real data are provided where the proposed estimator is compared with known benchmarks from the literature. It turns out that the existing estimators of the mean vector, including the new proposal, converge to the sample mean vector when the true mean vector has an unbounded Euclidean norm.

math.ST

Bayesian mean-variance analysis: Optimal portfolio selection under parameter uncertainty

The paper solves the problem of optimal portfolio choice when the parameters of the asset returns distribution, like the mean vector and the covariance matrix are unknown and have to be estimated by using historical data of the asset returns. The new approach employs the Bayesian posterior predictive distribution which is the distribution of the future realization of the asset returns given the observable sample. The parameters of the posterior predictive distributions are functions of the observed data values and, consequently, the solution of the optimization problem is expressed in terms of data only and does not depend on unknown quantities. In contrast, the optimization problem of the traditional approach is based on unknown quantities which are estimated in the second step leading to a suboptimal solution. We also derive a very useful stochastic representation of the posterior predictive distribution whose application leads not only to the solution of the considered optimization problem, but provides the posterior predictive distribution of the optimal portfolio return used to construct a prediction interval. A Bayesian efficient frontier, a set of optimal portfolios obtained by employing the posterior predictive distribution, is constructed as well. Theoretically and using real data we show that the Bayesian efficient frontier outperforms the sample efficient frontier, a common estimator of the set of optimal portfolios known to be overoptimistic.

q-fin.ST

Central limit theorems for functionals of large sample covariance matrix and mean vector in matrix-variate location mixture of normal distributions

In this paper we consider the asymptotic distributions of functionals of the sample covariance matrix and the sample mean vector obtained under the assumption that the matrix of observations has a matrix-variate location mixture of normal distributions. The central limit theorem is derived for the product of the sample covariance matrix and the sample mean vector. Moreover, we consider the product of the inverse sample covariance matrix and the mean vector for which the central limit theorem is established as well. All results are obtained under the large-dimensional asymptotic regime where the dimension $p$ and the sample size $n$ approach to infinity such that $p/n\to c\in[0 , +\infty)$ when the sample covariance matrix does not need to be invertible and $p/n\to c\in [0, 1)$ otherwise.

math.ST