SearcharxivSearch

arXiv subjects

Vasilis Sarafidis

Publications and source records attributed to Vasilis Sarafidis.

8 recordsLinked to original sources

Nonlinear Boosting with Multiple Testing in High-Dimensional Generalised Linear Models with Binary Responses

This paper proposes a nonlinear boosting with multiple testing (BMT) approach to variable selection in high-dimensional generalised linear models with binary responses. At each stage of the BMT procedure, the model is updated by adding only the most significant covariate, conditional on those already selected in previous stages, while taking into account the multiple testing nature of the problem. It is shown that, under the stated conditions, the BMT procedure selects all covariates whose true coefficients are nonzero, and no other covariates, with probability tending to one. Furthermore, the procedure enjoys an oracle property, in the sense that the post-BMT maximum likelihood estimator of the parameters of the model is asymptotically equivalent to an oracle estimator that knows the correct sparse model in advance. Monte Carlo experiments demonstrate that BMT outperforms competing methods, delivering high covariate-selection accuracy and low parameter estimation error. An empirical example illustrates that BMT delivers a predictive model for the probability that U.S. inflation exceeds a given threshold over a 12-month horizon which has very good out-of-sample performance.

econ.EM

Estimation and Inference for Latent Dual Networks Using High-Dimensional IV Screening

We develop a novel methodology for estimation and inference in high-dimensional panel network models with latent dual structures. The framework allows outcomes to be affected simultaneously by positive and negative interaction channels, accommodating settings in which some interactions reinforce outcomes while others generate competition and displacement effects. The proposed method identifies and estimates the network directly from the structural model using observed data without the need to pre-specify the network. Network recovery is achieved through a sequential instrumental-variable screening procedure. We establish exact support recovery and oracle-equivalent post-selection inference. An application to U.S. corporate leverage data reveals the coexistence of reinforcing and displacement interactions in firms' financial decisions.

econ.EM

Network Effects in Corporate Emissions: Evidence from a Data-Dependent Spatial Panel Model

We study spillover effects in corporate toxic emissions using a heterogeneous panel network of U.S. industrial facilities from 2000-2023. Rather than imposing a network structure a priori, we uncover an unobserved web of influence directly from the data using recent advances in high-dimensional network econometrics. Indirect effects transmitted through the estimated network account for about 28% of the total impact of key firm balance-sheet characteristics. By contrast, distance-based networks generate no statistically discernible spillovers, while a priori firm- or industry-based networks substantially overstate within-group spillins relative to the data-driven network. These findings show that who is linked to whom, and with what strength, matters critically for assessing systemic environmental risk and for designing targeted regulation. Methodologically, the paper provides a flexible framework for quantifying facility-level emissions spillovers and their consequences in financial and policy settings.

econ.GN

Model Selection in High-Dimensional Linear Regression using Boosting with Multiple Testing

High-dimensional regression specification and analysis is a complex and active area of research in statistics, machine learning, and econometrics. This paper proposes a new approach, Boosting with Multiple Testing (BMT), which combines forward stepwise variable selection with the multiple testing framework of Chudik et al (2018). At each stage, the model is updated by adding only the most significant regressor conditional on those already included, while a family-wise multiple testing filter is applied to the remaining candidates. In this way, the method retains the strong screening properties of Chudik et al (2018) while operating in a less greedy manner with respect to proxy and noise variables. Using sharp probability inequalities for heterogeneous strongly mixing processes from Dendramis et al (2022), we show that BMT enjoys oracle type properties relative to an approximating model that includes all true signals and excludes pure noise variables: this model is selected with probability tending to one, and the resulting estimator achieves standard parametric rates for prediction error and coefficient estimation. Additional results establish conditions under which BMT recovers the exact true model and avoids selection of proxy signals. Monte Carlo experiments indicate that BMT performs very well relative to OCMT and Lasso type procedures, delivering higher model selection accuracy and smaller RMSE for the estimated coefficients, especially under strong multicollinearity of the regressors. Two empirical illustrations based on a large set of macro-financial indicators as covariates, show that BMT yields sparse, interpretable specifications with favourable out-of-sample performance.

econ.EM

Chasing Opportunity: Spillovers and Drivers of U.S. State Population Growth

We study the drivers and spatial diffusion of U.S. state population growth using a dynamic spatial model for 49 states, 1965-2017. Methodologically, we recover the spatial network structure from the data, rather than imposing it a priori via contiguity or distance, and combine this with an IV estimator that permits heterogeneous slopes and interactive fixed effects. This unified design delivers consistent estimation and inference in a flexible spatial panel model with endogenous regressors, a data-inferred network structure, and pervasive cross-state dependence. To our knowledge, it is the first estimation framework in spatial econometrics to combine all three elements within a single setting. Empirically, population growth exhibits broad yet heterogeneous conditional convergence: about three-quarters of states converge, while a small high-growth group mildly diverges. Effects of the core drivers, amenities, labour income, migration frictions, are stable across various network specifications. On the other hand, the productivity effect emerges only when the network is estimated from the data. Spatial spillovers are sizable, with indirect effects roughly one-third of total impacts, and diffusion extending beyond contiguous neighbours.

econ.EM

Heterogeneous Exposures to Systematic and Idiosyncratic Risk across Crypto Assets: A Divide-and-Conquer Approach

This paper analyzes realized return behavior across a broad set of crypto assets by estimating heterogeneous exposures to idiosyncratic and systematic risk. A key challenge arises from the latent nature of broader economy-wide risk sources: macro-financial proxies are unavailable at high-frequencies, while the abundance of low-frequency candidates offers limited guidance on empirical relevance. To address this, we develop a two-stage ``divide-and-conquer'' approach. The first stage estimates exposures to high-frequency idiosyncratic and market risk only, using asset-level IV regressions. The second stage identifies latent economy-wide factors by extracting the leading principal component from the model residuals and mapping it to lower-frequency macro-financial uncertainty and sentiment-based indicators via high-dimensional variable selection. Structured patterns of heterogeneity in exposures are uncovered using Mean Group estimators across asset categories. The method is applied to a broad sample of crypto assets, covering more than 80% of total market capitalization. We document short-term mean reversion and significant average exposures to idiosyncratic volatility and illiquidity. Green and DeFi assets are, on average, more exposed to market-level and economy-wide risk than their non-Green and non-DeFi counterparts. By contrast, stablecoins are less exposed to idiosyncratic, market-level, and economy-wide risk factors relative to non-stablecoins. At a conceptual level, our study develops a coherent framework for isolating distinct layers of risk in crypto markets. Empirically, it sheds light on how return sensitivities vary across digital asset categories -- insights that are important for both portfolio design and regulatory oversight.

econ.EM

Residual Income Valuation and Stock Returns: Evidence from a Value-to-Price Investment Strategy

We hypothesize that portfolio sorts based on the V/P ratio generate excess returns and consist of companies that are undervalued for prolonged periods. Results, for the US market show that high V/P portfolios outperform low V/P portfolios across horizons extending from one to three years. The V/P ratio is positively correlated to future stock returns after controlling for firm characteristics, which are well known risk proxies. Findings also indicate that profitability and investment add explanatory power to the Fama and French three factor model and for stocks with V/P ratio close to 1. However, these factors cannot explain all variation in excess returns especially for years two and three and for stocks with high V/P ratio. Finally, portfolios with the highest V/P stocks select companies that are significantly mispriced relative to their equity (investment) and profitability growth persistence in the future.

econ.EM

IV Estimation of Heterogeneous Spatial Dynamic Panel Models with Interactive Effects

This paper develops a Mean Group Instrumental Variables (MGIV) estimator for spatial dynamic panel data models with interactive effects, under large N and T asymptotics. Unlike existing approaches that typically impose slope-parameter homogeneity, MGIV accommodates cross-sectional heterogeneity in slope coefficients. The proposed estimator is linear, making it computationally efficient and robust. Furthermore, it avoids the incidental parameters problem, enabling asymptotically valid inferences without requiring bias correction. The Monte Carlo experiments indicate strong finite-sample performance of the MGIV estimator across various sample sizes and parameter configurations. The practical utility of the estimator is illustrated through an application to regional economic growth in Europe. By explicitly incorporating heterogeneity, our approach provides fresh insights into the determinants of regional growth, underscoring the critical roles of spatial and temporal dependencies.

econ.EM