Searcharxiv⌕ Search

arXiv subjects

J. Martin van Zyl

Publications and source records attributed to J. Martin van Zyl.

13 recordsLinked to original sources

An Empirical Study of the Behaviour of the Sample Kurtosis in Samples from Symmetric Stable Distributions

Kurtosis is seen as a measure of the discrepancy between the observed data and a Gaussian distribution and is defined when the 4th moment is finite. In this work an empirical study is conducted to investigate the behaviour of the sample estimate of kurtosis with respect to sample size and the tail index when applied to heavy-tailed data where the 4th moment does not exist. The study will focus on samples from the symmetric stable distributions. It was found that the expected value of excess kurtosis divided by the sample size is finite for any value of the tail index and the sample estimate of kurtosis increases as a linear function of sample size and tail index. It is very sensitive to changes in the tail-index.

q-fin.ST↗

The sample fraction in peaks-over-threshold problems where the second-order expansion is valid with specific reference to the generalized Pareto distribution

In samples from a heavy-tailed distribution a second-order approximation is often use to approximate the tail function. Based on the parameters of the approximation, an optimal sample fraction can be estimated which is then used to estimate the index. Given that the observations are above a threshold and has an approximate generalized Pareto distribution, an expression is derived for the percentile above which the second-order approximation is valid

math.ST↗

The performance of univariate goodness-of-fit tests for normality based on the empirical characteristic function in large samples

An empirical power comparison is made between two tests based on the empirical characteristic function and some of the best performing tests for normality. A simple normality test based on the empirical characteristic function calculated in a single point is shown to outperform the more complicated Epps-Pulley test and the frequentist tests included in the study in large samples.

stat.CO↗

A goodness-of-fit test based on the empirical characteristic function and a comparison of tests for normality

The normal distribution has the unique property that the cumulant generating function has only two terms, namely those involving the mean and the variance. This property is used to construct a simple by using the log of the modulus of the empirical characteristic function to what would be expected under normality. The test statistic is easy to calculate. Using a simulation study the proposed test is shown to have excellent power, especially in large samples.

stat.ME↗

The efficiency of the likelihood ratio to choose between a t-distribution and a normal distribution

A decision must often be made between heavy-tailed and Gaussian errors for a regression or a time series model, and the t-distribution is frequently used when it is assumed that the errors are heavy-tailed distributed. The performance of the likelihood ratio to choose between the two distributions is investigated using entropy properties and a simulation study. The proportion of times or probability that the likelihood of the correct assumption will be bigger than the likelihood of the incorrect assumption is estimated.

stat.CO↗

Estimating the Tail Index by using Model Averaging

The ideas of model averaging are used to find weights in peak-over-threshold problems using a possible range of thresholds. A range of the largest observations are chosen and considered as possible thresholds, each time performing estimation. Weights based on an information criterion for each threshold are calculated. A weighted estimate of the threshold and shape parameter can be calculated.

stat.OT↗

An empirical study to order citation statistics between subject fields

An empirical study is conducted to compare citations per publication, statistics and observed Hirsch indexes between subject fields using summary statistics of countries. No distributional assumptions are made and ratios are calculated. These ratios can be used to make approximate comparisons between researchers of different subject fields with respect to the Hirsch index.

cs.DL↗

Regression with an infinite number of observations applied to estimating the parameters of the stable distribution using the empirical characteristic function

A function of the empirical characteristic function,exists for the stable distribution, which leads to a linear regression and can be used to estimate the parameters. Two approaches are often used, one to find optimal values of t, but these points are dependent on the unknown parameters. And using a fixed number of values for t. In this work the results when all points in an interval is used, thus where least squares using an infinite number of observations,is approximated. It was found that this procedure performs good in small samples.

stat.CO↗

Applying least absolute deviation regression to regression-type estimation of the index of a stable distribution using the characteristic function

Least absolute deviation regression is applied using a fixed number of points for all values of the index to estimate the index and scale parameter of the stable distribution using regression methods based on the empirical characteristic function. The recognized fixed number of points estimation procedure uses ten points in the interval zero to one, and least squares estimation. It is shown that using the more robust least absolute regression based on iteratively re-weighted least squares outperforms the least squares procedure with respect to bias and also mean square error in smaller samples.

stat.CO↗

An empirical study to check the accuracy of approximating averages of ratios using ratios of averages

For a number of researchers a number of publications for each author is simulated using the zeta distribution and then for each publication a number of citations per publication simulated. Bootstrap confidence intervals indicate that the difference between the average of ratios and the ratio of averages are not significant, and there are no significant differences in the distributions in realistic problems when using the two-sample Kolmogorov-Smirnov test to compare distributions. It was found that the log-logistic distribution which is a general form for the ratio of two correlated Pareto random variables, give a good fit to the estimated ratios.

stat.AP↗

Estimation of the shape parameter of a generalized Pareto distribution based on a transformation to Pareto distributed variables

Random variables of the generalized Pareto distribution, can be transformed to that of the Pareto distribution. Explicit expressions exist for the maximum likelihood estimators of the parameters of the Pareto distribution. The performance of the estimation of the shape parameter of generalized Pareto distributed using transformed observations, based on the probability weighted method is tested. It was found to improve the performance of the probability weighted estimator and performs good with respect to bias and MSE.

q-fin.CP↗

Testing approximate normality of an estimator using the estimated MSE and bias with an application to the shape parameter of the generalized Pareto distribution

Often it is not easy to choose between estimators, based on the estimated MSE and bias using simulation studies. Normality in small samples and a variance of the estimator, which is correct and easy to calculate using a single sample, give the added advantage that hypotheses concerning the parameter can be tested in new samples. A procedure to check normality is proposed where previously published MSE and bias are used to perform a test for normality. A confidence interval for the index of the S&P500 index is found by applying the results to estimators of the generalized Pareto distribution.

stat.AP↗

A weighted least squares procedure to approximate least absolute deviation estimation in time series with specific reference to infinite variance unit root problems

A weighted regression procedure is proposed for regression type problems where the innovations are heavy-tailed. This method approximates the least absolute regression method in large samples, and the main advantage will be if the sample is large and for problems with many independent variables. In such problems bootstrap methods must often be utilized to test hypotheses and especially in such a case this procedure has an advantage over least absolute regression. The procedure will be illustrated on first-order autoregressive problems, including the random walk. A bootstrap procedure is used to test the unit root hypothesis and good results were found.

stat.CO↗