SearcharxivSearch

arXiv subjects

Christophe Croux

Publications and source records attributed to Christophe Croux.

18 recordsLinked to original sources

Robust XGBoosting for Regression

XGBoost is a very popular and powerful method for prediction. It iteratively fits simple decision trees to the residuals of the previous step. An efficient and scalable implementation is available. The standard loss function for XGBoost is the quadratic loss, but a Huber loss can also be used. In this paper, we study the robustness of XGBoost and show that its performance can be affected by vertical outliers and leverage points. To address this, we explore alternative loss functions, based on M-, S-, and {\tau} -estimators from robust regression. Our results indicate that a two-step procedure, referred to as MM-XGBoost, provides the best trade-off between robustness and prediction accuracy.

cs.LG

Detecting Anti-dumping Circumvention: A Network Approach

Despite the increasing integration of the global economic system, anti-dumping measures are a common tool used by governments to protect their national economy. In this paper, we propose a methodology to detect cases of anti-dumping circumvention through re-routing trade via a third country. Based on the observed full network of trade flows, we propose a measure to proxy the evasion of an anti-dumping duty for a subset of trade flows directed to the European Union, and look for possible cases of circumvention of an active anti-dumping duty. Using panel regression, we are able correctly classify 86% of the trade flows, on which an investigation of anti-dumping circumvention has been opened by the European authorities.

econ.GN

Robust multivariate methods in Chemometrics

This chapter presents an introduction to robust statistics with applications of a chemometric nature. Following a description of the basic ideas and concepts behind robust statistics, including how robust estimators can be conceived, the chapter builds up to the construction (and use) of robust alternatives for some methods for multivariate analysis frequently used in chemometrics, such as principal component analysis and partial least squares. The chapter then provides an insight into how these robust methods can be used or extended to classification. To conclude, the issue of validation of the results is being addressed: it is shown how uncertainty statements associated with robust estimates, can be obtained.

stat.ME

Sliced Average Variance Estimation for Multivariate Time Series

Supervised dimension reduction for time series is challenging as there may be temporal dependence between the response $y$ and the predictors $\boldsymbol x$. Recently a time series version of sliced inverse regression, TSIR, was suggested, which applies approximate joint diagonalization of several supervised lagged covariance matrices to consider the temporal nature of the data. In this paper we develop this concept further and propose a time series version of sliced average variance estimation, TSAVE. As both TSIR and TSAVE have their own advantages and disadvantages, we consider furthermore a hybrid version of TSIR and TSAVE. Based on examples and simulations we demonstrate and evaluate the differences between the three methods and show also that they are superior to apply their iid counterparts to when also using lagged values of the explaining variables as predictors.

stat.ME

Volatility Spillovers and Heavy Tails: A Large t-Vector AutoRegressive Approach

Volatility is a key measure of risk in financial analysis. The high volatility of one financial asset today could affect the volatility of another asset tomorrow. These lagged effects among volatilities - which we call volatility spillovers - are studied using the Vector AutoRegressive (VAR) model. We account for the possible fat-tailed distribution of the VAR model errors using a VAR model with errors following a multivariate Student t-distribution with unknown degrees of freedom. Moreover, we study volatility spillovers among a large number of assets. To this end, we use penalized estimation of the VAR model with t-distributed errors. We study volatility spillovers among energy, biofuel and agricultural commodities and reveal bidirectional volatility spillovers between energy and biofuel, and between energy and agricultural commodities.

q-fin.ST

Lasso-based forecast combinations for forecasting realized variances

Volatility forecasts are key inputs in financial analysis. While lasso based forecasts have shown to perform well in many applications, their use to obtain volatility forecasts has not yet received much attention in the literature. Lasso estimators produce parsimonious forecast models. Our forecast combination approach hedges against the risk of selecting a wrong degree of model parsimony. Apart from the standard lasso, we consider several lasso extensions that account for the dynamic nature of the forecast model. We apply forecast combined lasso estimators in a comprehensive forecasting exercise using realized variance time series of ten major international stock market indices. We find the lasso extended "ordered lasso" to give the most accurate realized variance forecasts. Multivariate forecast models, accounting for volatility spillovers between different stock markets, outperform univariate forecast models for longer forecast horizons.

stat.AP

Multi-class Vector AutoRegressive Models for Multi-store Sales Data

Retailers use the Vector AutoRegressive (VAR) model as a standard tool to estimate the effects of prices, promotions and sales in one product category on the sales of another product category. Besides, these price, promotion and sales data are available for not just one store, but a whole chain of stores. We propose to study cross-category effects using a multi-class VAR model: we jointly estimate cross-category effects for several distinct but related VAR models, one for each store. Our methodology encourages effects to be similar across stores, while still allowing for small differences between stores to account for store heterogeneity. Moreover, our estimator is sparse: unimportant effects are estimated as exactly zero, which facilitates the interpretation of the results. A simulation study shows that the proposed multi-class estimator improves estimation accuracy by borrowing strength across classes. Finally, we provide three visual tools showing (i) the clustering of stores on identical cross-category effects, (ii) the networks of product categories and (iii) the similarity matrices of shared cross-category effects across stores.

stat.AP

Commodity Dynamics: A Sparse Multi-class Approach

The correct understanding of commodity price dynamics can bring relevant improvements in terms of policy formulation both for developing and developed countries. Agricultural, metal and energy commodity prices might depend on each other: although we expect few important effects among the total number of possible ones, some price effects among different commodities might still be substantial. Moreover, the increasing integration of the world economy suggests that these effects should be comparable for different markets. This paper introduces a sparse estimator of the Multi-class Vector AutoRegressive model to detect common price effects between a large number of commodities, for different markets or investment portfolios. In a first application, we consider agricultural, metal and energy commodities for three different markets. We show a large prevalence of effects involving metal commodities in the Chinese and Indian markets, and the existence of asymmetric price effects. In a second application, we analyze commodity prices for five different investment portfolios, and highlight the existence of important effects from energy to agricultural commodities. The relevance of biofuels is hereby confirmed. Overall, we find stronger similarities in commodity price effects among portfolios than among markets.

econ.GN

An algorithm for the multivariate group lasso with covariance estimation

We study a group lasso estimator for the multivariate linear regression model that accounts for correlated error terms. A block coordinate descent algorithm is used to compute this estimator. We perform a simulation study with categorical data and multivariate time series data, typical settings with a natural grouping among the predictor variables. Our simulation studies show the good performance of the proposed group lasso estimator compared to alternative estimators. We illustrate the method on a time series data set of gene expressions.

stat.CO

The predictive power of the business and bank sentiment of firms: A high-dimensional Granger Causality approach

We study the predictive power of industry-specific economic sentiment indicators for future macro-economic developments. In addition to the sentiment of firms towards their own business situation, we study their sentiment with respect to the banking sector - their main credit providers. The use of industry-specific sentiment indicators results in a high-dimensional forecasting problem. To identify the most predictive industries, we present a bootstrap Granger Causality test based on the Adaptive Lasso. This test is more powerful than the standard Wald test in such high-dimensional settings. Forecast accuracy is improved by using only the most predictive industries rather than all industries.

stat.AP

Identifying Demand Effects in a Large Network of Product Categories

Planning marketing mix strategies requires retailers to understand within- as well as cross-category demand effects. Most retailers carry products in a large variety of categories, leading to a high number of such demand effects to be estimated. At the same time, we do not expect cross-category effects between all categories. This paper outlines a methodology to estimate a parsimonious product category network without prior constraints on its structure. To do so, sparse estimation of the Vector AutoRegressive Market Response Model is presented. We find that cross-category effects go beyond substitutes and complements, and that categories have asymmetric roles in the product category network. Destination categories are most influential for other product categories, while convenience and occasional categories are most responsive. Routine categories are moderately influential and moderately responsive.

stat.AP

Robust high-dimensional precision matrix estimation

The dependency structure of multivariate data can be analyzed using the covariance matrix $Σ$. In many fields the precision matrix $Σ^{-1}$ is even more informative. As the sample covariance estimator is singular in high-dimensions, it cannot be used to obtain a precision matrix estimator. A popular high-dimensional estimator is the graphical lasso, but it lacks robustness. We consider the high-dimensional independent contamination model. Here, even a small percentage of contaminated cells in the data matrix may lead to a high percentage of contaminated rows. Downweighting entire observations, which is done by traditional robust procedures, would then results in a loss of information. In this paper, we formally prove that replacing the sample covariance matrix in the graphical lasso with an elementwise robust covariance matrix leads to an elementwise robust, sparse precision matrix estimator computable in high-dimensions. Examples of such elementwise robust covariance estimators are given. The final precision matrix estimator is positive definite, has a high breakdown point under elementwise contamination and can be computed fast.

stat.ME

The shooting S-estimator for robust regression

To perform multiple regression, the least squares estimator is commonly used. However, this estimator is not robust to outliers. Therefore, robust methods such as S-estimation have been proposed. These estimators flag any observation with a large residual as an outlier and downweight it in the further procedure. However, a large residual may be caused by an outlier in only one single predictor variable, and downweighting the complete observation results in a loss of information. Therefore, we propose the shooting S-estimator, a regression estimator that is especially designed for situations where a large number of observations suffer from contamination in a small number of predictor variables. The shooting S-estimator combines the ideas of the coordinate descent algorithm with simple S-regression, which makes it robust against componentwise contamination, at the cost of failing the regression equivariance property.

stat.ME

Sparse canonical correlation analysis from a predictive point of view

Canonical correlation analysis (CCA) describes the associations between two sets of variables by maximizing the correlation between linear combinations of the variables in each data set. However, in high-dimensional settings where the number of variables exceeds the sample size or when the variables are highly correlated, traditional CCA is no longer appropriate. This paper proposes a method for sparse CCA. Sparse estimation produces linear combinations of only a subset of variables from each data set, thereby increasing the interpretability of the canonical variates. We consider the CCA problem from a predictive point of view and recast it into a regression framework. By combining an alternating regression approach together with a lasso penalty, we induce sparsity in the canonical vectors. We compare the performance with other sparse CCA techniques in different simulation settings and illustrate its usefulness on a genomic data set.

stat.ME

Robust Sparse Canonical Correlation Analysis

Canonical correlation analysis (CCA) is a multivariate statistical method which describes the associations between two sets of variables. The objective is to find linear combinations of the variables in each data set having maximal correlation. This paper discusses a method for Robust Sparse CCA. Sparse estimation produces canonical vectors with some of their elements estimated as exactly zero. As such, their interpretability is improved. We also robustify the method such that it can cope with outliers in the data. To estimate the canonical vectors, we convert the CCA problem into an alternating regression framework, and use the sparse Least Trimmed Squares estimator. We illustrate the good performance of the Robust Sparse CCA method in several simulation studies and two real data examples.

stat.ME

Sparse cointegration

Cointegration analysis is used to estimate the long-run equilibrium relations between several time series. The coefficients of these long-run equilibrium relations are the cointegrating vectors. In this paper, we provide a sparse estimator of the cointegrating vectors. The estimation technique is sparse in the sense that some elements of the cointegrating vectors will be estimated as zero. For this purpose, we combine a penalized estimation procedure for vector autoregressive models with sparse reduced rank regression. The sparse cointegration procedure achieves a higher estimation accuracy than the traditional Johansen cointegration approach in settings where the true cointegrating vectors have a sparse structure, and/or when the sample size is low compared to the number of time series. We also discuss a criterion to determine the cointegration rank and we illustrate its good performance in several simulation settings. In a first empirical application we investigate whether the expectations hypothesis of the term structure of interest rates, implying sparse cointegrating vectors, holds in practice. In a second empirical application we show that forecast performance in high-dimensional systems can be improved by sparsely estimating the cointegration relations.

stat.ME

The Influence Function of Penalized Regression Estimators

To perform regression analysis in high dimensions, lasso or ridge estimation are a common choice. However, it has been shown that these methods are not robust to outliers. Therefore, alternatives as penalized M-estimation or the sparse least trimmed squares (LTS) estimator have been proposed. The robustness of these regression methods can be measured with the influence function. It quantifies the effect of infinitesimal perturbations in the data. Furthermore it can be used to compute the asymptotic variance and the mean squared error. In this paper we compute the influence function, the asymptotic variance and the mean squared error for penalized M-estimators and the sparse LTS estimator. The asymptotic biasedness of the estimators make the calculations nonstandard. We show that only M-estimators with a loss function with a bounded derivative are robust against regression outliers. In particular, the lasso has an unbounded influence function.

math.ST

Sparse least trimmed squares regression for analyzing high-dimensional large data sets

Sparse model estimation is a topic of high importance in modern data analysis due to the increasing availability of data sets with a large number of variables. Another common problem in applied statistics is the presence of outliers in the data. This paper combines robust regression and sparse model estimation. A robust and sparse estimator is introduced by adding an $L_1$ penalty on the coefficient estimates to the well-known least trimmed squares (LTS) estimator. The breakdown point of this sparse LTS estimator is derived, and a fast algorithm for its computation is proposed. In addition, the sparse LTS is applied to protein and gene expression data of the NCI-60 cancer cell panel. Both a simulation study and the real data application show that the sparse LTS has better prediction performance than its competitors in the presence of leverage points.

stat.AP