SearcharxivSearch

arXiv subjects

George Kapetanios

Publications and source records attributed to George Kapetanios.

14 recordsLinked to original sources

Grow and Pollute but Invest and Clean: Dynamic Associations between Parent Firm Characteristics and Facility Toxic Releases

This paper examines how relationships between parent-firm characteristics and facility-level toxic releases evolve over time. Using 238,304 observations for 7,447 U.S. manufacturing facilities from 1992 to 2023, we link on-site releases from the Toxics Release Inventory to financial, managerial, and macroeconomic data. A time-varying mean-group estimator accommodates changes in average coefficients over time and heterogeneity across facilities. The main result is a persistent contrast between operating scale and investment-related adjustment: sales is positively associated with release growth, whereas investment intensity is negatively associated. This pattern remains in joint specifications, parsimonious models, and the main robustness exercises. Other financial and managerial associations are identified as well. These findings suggest that environmental policy evaluation may benefit from distinguishing release changes associated with production expansion from those associated with investment-related adjustment and from allowing for variation across periods and production settings. For corporate managers, the same distinction can inform how emissions considerations are incorporated into capital budgeting and production planning.

econ.EM

Nonlinear Boosting with Multiple Testing in High-Dimensional Generalised Linear Models with Binary Responses

This paper proposes a nonlinear boosting with multiple testing (BMT) approach to variable selection in high-dimensional generalised linear models with binary responses. At each stage of the BMT procedure, the model is updated by adding only the most significant covariate, conditional on those already selected in previous stages, while taking into account the multiple testing nature of the problem. It is shown that, under the stated conditions, the BMT procedure selects all covariates whose true coefficients are nonzero, and no other covariates, with probability tending to one. Furthermore, the procedure enjoys an oracle property, in the sense that the post-BMT maximum likelihood estimator of the parameters of the model is asymptotically equivalent to an oracle estimator that knows the correct sparse model in advance. Monte Carlo experiments demonstrate that BMT outperforms competing methods, delivering high covariate-selection accuracy and low parameter estimation error. An empirical example illustrates that BMT delivers a predictive model for the probability that U.S. inflation exceeds a given threshold over a 12-month horizon which has very good out-of-sample performance.

econ.EM

Estimation and Inference for Latent Dual Networks Using High-Dimensional IV Screening

We develop a novel methodology for estimation and inference in high-dimensional panel network models with latent dual structures. The framework allows outcomes to be affected simultaneously by positive and negative interaction channels, accommodating settings in which some interactions reinforce outcomes while others generate competition and displacement effects. The proposed method identifies and estimates the network directly from the structural model using observed data without the need to pre-specify the network. Network recovery is achieved through a sequential instrumental-variable screening procedure. We establish exact support recovery and oracle-equivalent post-selection inference. An application to U.S. corporate leverage data reveals the coexistence of reinforcing and displacement interactions in firms' financial decisions.

econ.EM

Network Effects in Corporate Emissions: Evidence from a Data-Dependent Spatial Panel Model

We study spillover effects in corporate toxic emissions using a heterogeneous panel network of U.S. industrial facilities from 2000-2023. Rather than imposing a network structure a priori, we uncover an unobserved web of influence directly from the data using recent advances in high-dimensional network econometrics. Indirect effects transmitted through the estimated network account for about 28% of the total impact of key firm balance-sheet characteristics. By contrast, distance-based networks generate no statistically discernible spillovers, while a priori firm- or industry-based networks substantially overstate within-group spillins relative to the data-driven network. These findings show that who is linked to whom, and with what strength, matters critically for assessing systemic environmental risk and for designing targeted regulation. Methodologically, the paper provides a flexible framework for quantifying facility-level emissions spillovers and their consequences in financial and policy settings.

econ.GN

Model Selection in High-Dimensional Linear Regression using Boosting with Multiple Testing

High-dimensional regression specification and analysis is a complex and active area of research in statistics, machine learning, and econometrics. This paper proposes a new approach, Boosting with Multiple Testing (BMT), which combines forward stepwise variable selection with the multiple testing framework of Chudik et al (2018). At each stage, the model is updated by adding only the most significant regressor conditional on those already included, while a family-wise multiple testing filter is applied to the remaining candidates. In this way, the method retains the strong screening properties of Chudik et al (2018) while operating in a less greedy manner with respect to proxy and noise variables. Using sharp probability inequalities for heterogeneous strongly mixing processes from Dendramis et al (2022), we show that BMT enjoys oracle type properties relative to an approximating model that includes all true signals and excludes pure noise variables: this model is selected with probability tending to one, and the resulting estimator achieves standard parametric rates for prediction error and coefficient estimation. Additional results establish conditions under which BMT recovers the exact true model and avoids selection of proxy signals. Monte Carlo experiments indicate that BMT performs very well relative to OCMT and Lasso type procedures, delivering higher model selection accuracy and smaller RMSE for the estimated coefficients, especially under strong multicollinearity of the regressors. Two empirical illustrations based on a large set of macro-financial indicators as covariates, show that BMT yields sparse, interpretable specifications with favourable out-of-sample performance.

econ.EM

Detecting Network Instability via Multiscale Detrended Cross-Correlations and MST Topology

We introduce a multiscale measure of network instability based on the joint use of Detrended Cross-Correlation Analysis (DCCA) and Minimum Spanning Tree (MST) filtering. The proposed metric, the Elastic Detrended Cross-Correlation Ratio (Elastic DCCR), is defined as a finite-difference measure of the logarithmic sensitivity of the average MST length to the observation scale. It captures how the structure of cross-correlation networks deforms across different investment horizons. When applied to a network of global equity indices, the Elastic DCCR rises sharply during episodes of financial stress, reflecting increased short-term coordination among investors and a contraction of correlation distances. The measure reveals scale-dependent reconfigurations in network topology that are not visible in single-scale analyses, and highlights clear differences between stressed and stable market regimes. The approach does not assume covariance stationarity and relies only on scale-dependent detrended correlations; as a result, it is broadly applicable to other complex systems in which interaction strength varies with scale.

physics.soc-ph

Unlocking the Regression Space

This paper introduces and analyzes a framework that accommodates general heterogeneity in regression modeling. It demonstrates that regression models with fixed or time-varying parameters can be estimated using the OLS and time-varying OLS methods, respectively, across a broad class of regressors and noise processes not covered by existing theory. The proposed setting facilitates the development of asymptotic theory and the estimation of robust standard errors. The robust confidence interval estimators accommodate substantial heterogeneity in both regressors and noise. The resulting robust standard error estimates coincide with White's (1980) heteroskedasticity-consistent estimator but are applicable to a broader range of conditions, including models with missing data. They are computationally simple and perform well in Monte Carlo simulations. Their robustness, generality, and ease of implementation make them highly suitable for empirical applications. Finally, the paper provides a brief empirical illustration.

econ.EM

Heterogeneous Exposures to Systematic and Idiosyncratic Risk across Crypto Assets: A Divide-and-Conquer Approach

This paper analyzes realized return behavior across a broad set of crypto assets by estimating heterogeneous exposures to idiosyncratic and systematic risk. A key challenge arises from the latent nature of broader economy-wide risk sources: macro-financial proxies are unavailable at high-frequencies, while the abundance of low-frequency candidates offers limited guidance on empirical relevance. To address this, we develop a two-stage ``divide-and-conquer'' approach. The first stage estimates exposures to high-frequency idiosyncratic and market risk only, using asset-level IV regressions. The second stage identifies latent economy-wide factors by extracting the leading principal component from the model residuals and mapping it to lower-frequency macro-financial uncertainty and sentiment-based indicators via high-dimensional variable selection. Structured patterns of heterogeneity in exposures are uncovered using Mean Group estimators across asset categories. The method is applied to a broad sample of crypto assets, covering more than 80% of total market capitalization. We document short-term mean reversion and significant average exposures to idiosyncratic volatility and illiquidity. Green and DeFi assets are, on average, more exposed to market-level and economy-wide risk than their non-Green and non-DeFi counterparts. By contrast, stablecoins are less exposed to idiosyncratic, market-level, and economy-wide risk factors relative to non-stablecoins. At a conceptual level, our study develops a coherent framework for isolating distinct layers of risk in crypto markets. Empirically, it sheds light on how return sensitivities vary across digital asset categories -- insights that are important for both portfolio design and regulatory oversight.

econ.EM

Investor behavior and multiscale cross-correlations: Unveiling regime shifts in global financial markets

We propose an algorithm to capture emergent patterns in the cross-correlations of financial markets, highlighting regime changes on a global scale. In our approach, financial markets are viewed as complex adaptive systems, and multiscale properties and cross-correlations are considered, particularly during stress conditions such as the COVID-19 pandemic, the invasion of Ukraine by Russia in 2022, and Brexit. We investigate whether significant disruptions reflect an imbalance in investment horizons among investors, and we propose a measure based on this imbalance to depict the impact on global financial markets. The detrended cross-correlation cost (DCCC), which is derived from detrended cross-correlation analysis, uses cross-correlations at different timescales to capture variations in investment horizons amid financial uncertainties. Our algorithm, which combines DCCC analysis and the minimum-spanning-tree filtering approach, tracks system interconnectedness and investor imbalances. We tested the DCCC indicator using daily price series of G7, Russian, and Chinese markets over the past decade and found that it increases sharply during ``crash'' periods compared to ``business as usual'' periods. Our empirical results confirm that short-term investment horizons dominate during financial instabilities; this validates our hypothesis and indicates that the DCCC can serve as a leading indicator of shifts in financial-market regimes.

econ.GN

Heterogeneous Grouping Structures in Panel Data

In this paper we examine the existence of heterogeneity within a group, in panels with latent grouping structure. The assumption of within group homogeneity is prevalent in this literature, implying that the formation of groups alleviates cross-sectional heterogeneity, regardless of the prior knowledge of groups. While the latter hypothesis makes inference powerful, it can be often restrictive. We allow for models with richer heterogeneity that can be found both in the cross-section and within a group, without imposing the simple assumption that all groups must be heterogeneous. We further contribute to the method proposed by \cite{su2016identifying}, by showing that the model parameters can be consistently estimated and the groups, while unknown, can be identifiable in the presence of different types of heterogeneity. Within the same framework we consider the validity of assuming both cross-sectional and within group homogeneity, using testing procedures. Simulations demonstrate good finite-sample performance of the approach in both classification and estimation, while empirical applications across several datasets provide evidence of multiple clusters, as well as reject the hypothesis of within group homogeneity.

econ.EM

On Robust Inference in Time Series Regression

Least squares regression with heteroskedasticity consistent standard errors ("OLS-HC regression") has proved very useful in cross section environments. However, several major difficulties, which are generally overlooked, must be confronted when transferring the HC technology to time series environments via heteroskedasticity and autocorrelation consistent standard errors ("OLS-HAC regression"). First, in plausible time-series environments, OLS parameter estimates can be inconsistent, so that OLS-HAC inference fails even asymptotically. Second, most economic time series have autocorrelation, which renders OLS parameter estimates inefficient. Third, autocorrelation similarly renders conditional predictions based on OLS parameter estimates inefficient. Finally, the structure of popular HAC covariance matrix estimators is ill-suited for capturing the autoregressive autocorrelation typically present in economic time series, which produces large size distortions and reduced power in HAC-based hypothesis testing, in all but the largest samples. We show that all four problems are largely avoided by the use of a simple and easily-implemented dynamic regression procedure, which we call DURBIN. We demonstrate the advantages of DURBIN with detailed simulations covering a range of practical issues.

econ.EM

High Dimensional Generalised Penalised Least Squares

In this paper we develop inference for high dimensional linear models, with serially correlated errors. We examine Lasso under the assumption of strong mixing in the covariates and error process, allowing for fatter tails in their distribution. While the Lasso estimator performs poorly under such circumstances, we estimate via GLS Lasso the parameters of interest and extend the asymptotic properties of the Lasso under more general conditions. Our theoretical results indicate that the non-asymptotic bounds for stationary dependent processes are sharper, while the rate of Lasso under general conditions appears slower as $T,p\to \infty$. Further we employ the debiased Lasso to perform inference uniformly on the parameters of interest. Monte Carlo results support the proposed estimator, as it has significant efficiency gains over traditional methods.

econ.EM

Deep Neural Network Estimation in Panel Data Models

In this paper we study neural networks and their approximating power in panel data models. We provide asymptotic guarantees on deep feed-forward neural network estimation of the conditional mean, building on the work of Farrell et al. (2021), and explore latent patterns in the cross-section. We use the proposed estimators to forecast the progression of new COVID-19 cases across the G7 countries during the pandemic. We find significant forecasting gains over both linear panel and nonlinear time series models. Containment or lockdown policies, as instigated at the national-level by governments, are found to have out-of-sample predictive power for new COVID-19 cases. We illustrate how the use of partial derivatives can help open the "black-box" of neural networks and facilitate semi-structural analysis: school and workplace closures are found to have been effective policies at restricting the progression of the pandemic across the G7 countries. But our methods illustrate significant heterogeneity and time-variation in the effectiveness of specific containment policies.

econ.EM

A New Test for Market Efficiency and Uncovered Interest Parity

We suggest a new single-equation test for Uncovered Interest Parity (UIP) based on a dynamic regression approach. The method provides consistent and asymptotically efficient parameter estimates, and is not dependent on assumptions of strict exogeneity. This new approach is asymptotically more efficient than the common approach of using OLS with HAC robust standard errors in the static forward premium regression. The coefficient estimates when spot return changes are regressed on the forward premium are all positive and remarkably stable across currencies. These estimates are considerably larger than those of previous studies, which frequently find negative coefficients. The method also has the advantage of showing dynamic effects of risk premia, or other events that may lead to rejection of UIP or the efficient markets hypothesis.

econ.EM