SearcharxivSearch

arXiv subjects

Chris Tofallis

Publications and source records attributed to Chris Tofallis.

5 recordsLinked to original sources

Fitting an Equation to Data Impartially

We consider the problem of fitting a relationship (e.g. a potential scientific law) to data involving multiple variables. Ordinary (least squares) regression is not suitable for this because the estimated relationship will differ according to which variable is chosen as being dependent, and the dependent variable is unrealistically assumed to be the only variable which has any measurement error (noise). We present a very general method for estimating a linear functional relationship between multiple noisy variables, which are treated impartially, i.e. no distinction between dependent and independent variables. The data are not assumed to follow any distribution, but all variables are treated as being equally reliable. Our approach extends the geometric mean functional relationship to multiple dimensions. This is especially useful with variables measured in different units, as it is naturally scale-invariant, whereas orthogonal regression is not. This is because our approach is not based on minimizing distances, but on the symmetric concept of correlation. The estimated coefficients are easily obtained from the covariances or correlations, and correspond to geometric means of associated least squares coefficients. The ease of calculation will hopefully allow widespread application of impartial fitting to estimate relationships in a neutral way.

stat.ME

Objective Weights for Scoring: The Automatic Democratic Method

When comparing performance (of products, services, entities, etc.), multiple attributes are involved. This paper deals with a way of weighting these attributes when one is seeking an overall score. It presents an objective approach to generating the weights in a scoring formula which avoids personal judgement. The first step is to find the maximum possible score for each assessed entity. These upper bound scores are found using Data Envelopment Analysis. In the second step the weights in the scoring formula are found by regressing the unique DEA scores on the attribute data. Reasons for using least squares and avoiding other distance measures are given. The method is tested on data where the true scores and weights are known. The method enables the construction of an objective scoring formula which has been generated from the data arising from all assessed entities and is, in that sense, democratic.

stat.ME

A better measure of relative prediction accuracy for model selection and model estimation

Surveys show that the mean absolute percentage error (MAPE) is the most widely used measure of forecast accuracy in businesses and organizations. It is however, biased: When used to select among competing prediction methods it systematically selects those whose predictions are too low. This is not widely discussed and so is not generally known among practitioners. We explain why this happens. We investigate an alternative relative accuracy measure which avoids this bias: the log of the accuracy ratio: log (prediction / actual). Relative accuracy is particularly relevant if the scatter in the data grows as the value of the variable grows (heteroscedasticity). We demonstrate using simulations that for heteroscedastic data (modelled by a multiplicative error factor) the proposed metric is far superior to MAPE for model selection. Another use for accuracy measures is in fitting parameters to prediction models. Minimum MAPE models do not predict a simple statistic and so theoretical analysis is limited. We prove that when the proposed metric is used instead, the resulting least squares regression model predicts the geometric mean. This important property allows its theoretical properties to be understood.

stat.ME

Investment Volatility: A Critique of Standard Beta Estimation and a Simple Way Forward

Beta is a widely used quantity in investment analysis. We review the common interpretations that are applied to beta in finance and show that the standard method of estimation - least squares regression - is inconsistent with these interpretations. We present the case for an alternative beta estimator which is more appropriate, as well as being easier to understand and to calculate. Unlike regression, the line fit we propose treats both variables in the same way. Remarkably, it provides a slope that is precisely the ratio of the volatility of the investment's rate of return to the volatility of the market index rate of return (or the equivalent excess rates of returns). Hence, this line fitting method gives an alternative beta, which corresponds exactly to the relative volatility of an investment - which is one of the usual interpretations attached to beta.

q-fin.PM

Model Building with Multiple Dependent Variables and Constraints

The most widely used method for finding relationships between several quantities is multiple regression. This however is restricted to a single dependent variable. We present a more general method which allows models to be constructed with multiple variables on both sides of an equation and which can be computed easily using a spreadsheet program. The underlying principle (originating from canonical correlation analysis) is that of maximising the correlation between the two sides of the model equation. This paper presents a fitting procedure which makes it possible to force the estimated model to satisfy constraint conditions which it is required to possess, these may arise from theory, prior knowledge or be intuitively obvious. We also show that the least squares approach to the problem is inadequate as it produces models which are not scale invariant.

math.ST