SearcharxivSearch

arXiv subjects

Xiaojun Song

Publications and source records attributed to Xiaojun Song.

At least 19 recordsLinked to original sources

Specification Testing for Dyadic Regression Models

This paper develops omnibus specification tests for linear conditional-mean models with undirected dyadic data. We establish a uniform projection theorem that reduces the dyadic process to its latent first-order node projections under shared-node dependence. We then show that a raw first-order node-multiplier bootstrap is valid when this node component is nondegenerate but double-counts dyad-specific variation when dyads are independent. An exact covariance decomposition motivates a corrected Gaussian bootstrap that is valid in both regimes. The resulting Kolmogorov-Smirnov and Cram\'er-von Mises tests are consistent against fixed alternatives and have nontrivial power against rate-appropriate local alternatives. Simulations show that the corrected Kolmogorov-Smirnov test provides the most stable size control while retaining substantial local power. An application to the Lazega law-firm network rejects additive linear and quadratic specifications but finds no remaining misspecification after including an economically relevant interaction.

econ.EM

Kernel Minimum Distance Estimation and Testing with Conditional Moment Restrictions: A Unified Framework

We propose a unified Kernel Minimum Distance (KMD) framework for estimating and testing models defined by conditional moment restrictions. By embedding conditional moments into a Reproducing Kernel Hilbert Space (RKHS), we construct a closed-form $V$-statistic objective function that quantifies the distance from the restrictions. We establish the $\sqrt{n}$-consistency and asymptotic normality of the associated minimum distance estimator. Within this framework, the minimized objective function naturally yields a consistent omnibus specification test. Unlike projection-based methods that require auxiliary nonparametric estimation for Neyman orthogonalization, our test inherently captures the estimation effect via a projected kernel structure. We derive asymptotic properties of the test statistics under the null hypothesis, the alternative hypothesis, and a sequence of local alternatives converging to the null at the parametric rate $n^{-1/2}$. The validity of a computationally simple multiplier bootstrap is established to facilitate inference. Simulation results demonstrate robust finite-sample performance, and the framework is illustrated by analyzing Engel curves using UK Family Expenditure Survey data.

econ.EM

Orthogonal Integrated Conditional Moment Tests for Treatment Effect Heterogeneity

We propose a nonparametric integrated conditional moment (ICM) test for treatment effect heterogeneity across subpopulations defined by a given covariate subvector. Under unconfoundedness, the null is recast as a conditional moment restriction based on a Neyman-orthogonal score, which reduces the first-order sensitivity of the empirical process to nuisance parameter estimation. The test statistics are constructed as continuous functionals of a marked empirical process. We establish a uniform feasible-to-oracle approximation and derive the asymptotic properties of these test statistics under the null and fixed alternatives. We further show that the test has nontrivial power against local alternatives converging to the null at the $n^{-1/2}$ rate, and develop an easy-to-implement multiplier bootstrap for feasible inference. We also develop extensions to tests of parametric CATE specifications and to settings with endogenous treatment and a binary instrument. Finally, we apply the proposed testing approach to study whether the effect of maternal smoking during pregnancy on infant birth weight varies with maternal age.

econ.EM

Kernel Two-Sample Testing via Directional Components Analysis

Standard kernel two-sample tests, such as those based on the Maximum Mean Discrepancy (MMD), aggregate squared differences across all directions in a Reproducing Kernel Hilbert Space (RKHS). However, in finite samples, trailing directional components are noisy, which degrades test power. We propose a novel kernel-based test that resolves this by truncating the spectral decomposition of the MMD, retaining only the well-estimated leading eigen-directions. By aggregating these robust components, our method achieves superior power and robustness, particularly in high-dimensional and unbalanced settings. Furthermore, we introduce a computationally efficient parametric bootstrap procedure for approximating critical values, which is theoretically justified and significantly faster than permutation-based alternatives. Extensive simulations and empirical studies demonstrate that our method maintains strict Type I error control while delivering higher power than existing MMD-based tests.

stat.ME

Paired Sample Tests for High-dimensional Uncorrelatedness via Random Integration

This paper proposes a novel nonparametric test to assess the uncorrelatedness between two high-dimensional random vectors. We develop our test by generalizing the random integration proposed by Jiang et al. (2023, 2024), and the resulting test statistic estimates a weighted squared $\mathscr{L}_2$ norm of the covariance matrix. Asymptotic properties of the test statistic are derived by letting both the sample size $n$ and the dimension $p$ diverge to infinity. Under the null hypothesis of uncorrelatedness, our proposed test statistic is asymptotically normal with zero mean and unit variance, without requiring any specification of the relative magnitude regarding $n$ and $p$. Monte Carlo simulations demonstrate the good finite-sample performance of our proposed methods. Compared with many existing tests, our test statistic is more powerful at detecting ``weak but pervasive'' dependence while maintaining a comparable empirical size. The advantages of the proposed methods are further illustrated by an empirical analysis that assesses the correlation between DNA methylation and gene expression.

stat.ME

Robust Inference for Dyadic Data with Dependent Ordered Nodes

Dyadic regression models are commonly analyzed under the conventional dyadic dependence framework, where two observations may be dependent only if the corresponding dyads share a node. This paper studies inference when nodes are ordered and nearby nodes are exposed to common latent shocks, so that dyads with no shared endpoint may still be dependent. Although each additional covariance term may be weak, the number of nearby-node dyad pairs grows with the sample size, making their aggregate contribution asymptotically non-negligible. We develop an inferential framework for dyadic arrays with ordered-node dependence and propose two variance estimators: a dependent-node dyadic cluster-robust variance estimator that retains covariance terms between dyads with nearby endpoints, and a row-column moving-block jackknife method that deletes adjacent blocks of nodes together with all dyads touching those nodes. We establish the asymptotic validity of both procedures under weak dependence along the ordered node index. Monte Carlo evidence shows improvements in size control, with the jackknife procedure displaying comparatively stable finite-sample performance. An application to international trade gravity regressions shows that accounting for ordered-node dependence substantially weakens the statistical evidence for free trade agreement effects.

econ.EM

Testing Heteroskedasticity Under Measurement Error

In this paper, we propose a novel approach to detect heteroskedasticity in regression models with regressors contaminated by measurement error. Specifically, inspired by the integrated conditional moment (ICM) approach, we construct test statistics based on a deconvolved residual-marked empirical process and establish their asymptotic properties in both ordinary smooth and supersmooth cases, assuming the measurement error distribution is known. The issue of an unknown measurement error distribution is addressed by employing estimators of the measurement error characteristic function based on repeated measurements. Furthermore, depending on whether the measurement error distribution is known or not, to obtain critical values from the case-dependent limiting null distributions, we propose two computationally attractive multiplier bootstrap methods where the "parameter estimation effect" is successfully addressed. Finally, simulation results and empirical studies about corn yields and household budget shares confirm the favorable properties of the proposed tests.

econ.EM

Data-driven Smooth Tests for Normality in ANOVA When the Number of Groups is Large

The normality assumption for random errors is fundamental in the analysis of variance (ANOVA) models. However, it is rarely subjected to formal testing in practice, and theoretically justified procedures are largely unavailable, especially when the number of groups diverges. In this paper, we develop Neyman's smooth tests for assessing normality in a broad class of ANOVA models, allowing the number of groups to diverge. The proposed test statistics are constructed via the Gaussian probability integral transformation of ANOVA residuals. We show that using residuals induces non-negligible parameter estimation effects, whose structure depends on the underlying ANOVA model and plays a crucial role in shaping the form of the test statistics and their asymptotic behavior. Under the null hypothesis of normality, the resulting statistics follow an asymptotic Chi-square distribution, with degrees of freedom determined by the order of the smooth test (i.e., the number of components included in the smooth test). We further propose a modified Schwarz's selection rule to automatically determine the order, thereby yielding fully data-driven smooth tests that require no additional tuning parameters. Simulation studies and a real-data example indicate that the proposed tests perform well in practice and are readily applicable.

econ.EM

A Projection Approach to Nonparametric Significance and Conditional Independence Testing

This paper develops a novel nonparametric significance test based on a tailored nonparametric-type projected weighting function that exhibits appealing theoretical and numerical properties. We derive the asymptotic properties of the proposed test and show that it can detect local alternatives at the parametric rate. Using the nonparametric orthogonal projection, we construct a computationally convenient multiplier bootstrap to obtain critical values from the case-dependent asymptotic null distribution. Compared with the existing literature, our approach overcomes the need for a stronger compact support assumption on the density of covariates arising from random denominators. We also extend the tailor-made projection procedure to test the conditional independence assumption. The simulation experiments further illustrate the advantages of our proposed method in testing significance and conditional independence in finite samples.

econ.EM

Specification tests for regression models with measurement errors

In this paper, we propose new specification tests for regression models with measurement errors in the explanatory variables. Inspired by the integrated conditional moment (ICM) approach, we use a deconvoluted residual-marked empirical process and construct ICM-type test statistics based on it. The issue of measurement errors is addressed by applying a deconvolution kernel estimator in constructing the residuals. We demonstrate that employing an orthogonal projection onto the tangent space of nuisance parameters not only eliminates the parameter estimation effect but also facilitates the simulation of critical values via a computationally simple multiplier bootstrap procedure. It is the first time a multiplier bootstrap has been proposed in the literature of specification testing with measurement errors. We also develop specification tests and the multiplier bootstrap procedure when the measurement error distribution is unknown. The finite-sample performance of the proposed tests for both known and unknown measurement error distributions is evaluated through Monte Carlo simulations, which demonstrate their efficacy.

econ.EM

Asymmetric Huber Periodogram

This paper introduces a novel spectral M-estimator, called the asymmetric Huber periodogram (AHP), as a generalization of the ordinary periodogram (PG), the quantile periodogram (QP), and the Huber periodogram (HP). The AHP is constructed via trigonometric asymmetric Huber regression (AHR), in which a specially designed loss function replaces the squared $\ell_2$ loss used to define the PG. Relative to the QP, the AHP can be more computationally efficient, and relative to the HP and PG, it provides a more comprehensive characterization by examining the data across the range of the asymmetry parameter. We establish the theoretical properties of the AHP and investigate its relationship with the corresponding asymmetric Huber spectrum (AHS). Building on the asymptotic theory, we develop confidence intervals (CIs) for the AHS and propose a Fisher-type test. Simulation studies and two applications further demonstrate the AHP's effectiveness in detecting hidden periodicities, robustness to outliers, and utility for time series clustering.

stat.ME

Deep learning based doubly robust test for Granger causality

Granger causality is popular for analyzing time series data in many applications from natural science to social science including genomics, neuroscience, economics, and finance. Consequently, the Granger causality test has become one of the main concerns of the econometrician for decades. Taking advantage of the theoretical breakthroughs in deep learning in recent years, we propose a doubly robust Granger causality test (DRGCT). Our method offers several key advantages. The first and most direct benefit is for the users, DRGCT allows them to handle large lag orders while alleviating the curse of dimensionality that traditional nonlinear Granger causality tests usually face. Second, introducing a doubly robust test statistic for time series based on neural networks that achieves a parametric convergence rate not only suggests a new paradigm for nonparametric inference in econometrics, but also broadens the application scope of deep learning. Third, a multiplier bootstrap method, combined with the doubly robust approach, provides an efficient way to obtain critical values, effectively reducing computational time and avoiding redundant calculations. We prove that the test asymptotically controls the type I error, while achieving power approaches one, and validate the effectiveness of our test through numerical simulations. In real data analysis, we apply DRGCT to revisit the price-volume relationship problem in the stock markets of America, China, and Japan.

stat.ME

Almost Dominance: Inference and Application

This paper proposes a general framework for inference on three types of almost dominances: almost Lorenz dominance, almost inverse stochastic dominance, and almost stochastic dominance. We first generalize almost Lorenz dominance to almost upward and downward Lorenz dominances. We then provide a bootstrap inference procedure for the Lorenz dominance coefficients, which measure the degrees of almost Lorenz dominance. Furthermore, we propose almost upward and downward inverse stochastic dominances and provide inference on the inverse stochastic dominance coefficients. We also show that our results can easily be extended to almost stochastic dominance. Simulation studies demonstrate the finite sample properties of the proposed estimators and the bootstrap confidence intervals. This framework can be applied to economic analysis, particularly in the areas of social welfare, inequality, and decision making under uncertainty. As an empirical example, we apply the methods to the inequality growth in the United Kingdom and find evidence for almost upward inverse stochastic dominance.

econ.EM

A Powerful Chi-Square Specification Test with Support Vectors

Specification tests, such as Integrated Conditional Moment (ICM) and Kernel Conditional Moment (KCM) tests, are crucial for model validation but often lack power in finite samples. This paper proposes a novel framework to enhance specification test performance using Support Vector Machines (SVMs) for direction learning. We introduce two alternative SVM-based approaches: one maximizes the discrepancy between nonparametric and parametric classes, while the other maximizes the separation between residuals and the origin. Both approaches lead to a $t$-type test statistic that converges to a standard chi-square distribution under the null hypothesis. Our method is computationally efficient and capable of detecting any arbitrary alternative. Simulation studies demonstrate its superior performance compared to existing methods, particularly in large-dimensional settings.

econ.EM

Mixture Conditional Regression with Ultrahigh Dimensional Text Data for Estimating Extralegal Factor Effects

Testing judicial impartiality is a problem of fundamental importance in empirical legal studies, for which standard regression methods have been popularly used to estimate the extralegal factor effects. However, those methods cannot handle control variables with ultrahigh dimensionality, such as found in judgment documents recorded in text format. To solve this problem, we develop a novel mixture conditional regression (MCR) approach, assuming that the whole sample can be classified into a number of latent classes. Within each latent class, a standard linear regression model can be used to model the relationship between the response and a key feature vector, which is assumed to be of a fixed dimension. Meanwhile, ultrahigh dimensional control variables are then used to determine the latent class membership, where a Naïve Bayes type model is used to describe the relationship. Hence, the dimension of control variables is allowed to be arbitrarily high. A novel expectation-maximization algorithm is developed for model estimation. Therefore, we are able to estimate the interested key parameters as efficiently as if the true class membership were known in advance. Simulation studies are presented to demonstrate the proposed MCR method. A real dataset of Chinese burglary offenses is analyzed for illustration purpose.

stat.ME

Specification tests for generalized propensity scores using double projections

This paper proposes a new class of nonparametric tests for the correct specification of models based on conditional moment restrictions, paying particular attention to generalized propensity score models. The test procedure is based on two different projection arguments, leading to test statistics that are suitable to setups with many covariates, and are (asymptotically) invariant to the estimation method used to estimate the nuisance parameters. We show that our proposed tests are able to detect a broad class of local alternatives converging to the null at the usual parametric rate and illustrate its attractive power properties via simulations. We also extend our proposal to test parametric or semiparametric single-index-type models.

econ.EM

Testing linearity in semi-functional partially linear regression models

This paper proposes a Kolmogorov-Smirnov type statistic and a Cramér-von Mises type statistic to test linearity in semi-functional partially linear regression models. Our test statistics are based on a residual marked empirical process indexed by a randomly projected functional covariate,which is able to circumvent the "curse of dimensionality" brought by the functional covariate. The asymptotic properties of the proposed test statistics under the null, the fixed alternative, and a sequence of local alternatives converging to the null at the $n^{1/2}$ rate are established. A straightforward wild bootstrap procedure is suggested to estimate the critical values that are required to carry out the tests in practical applications. Results from an extensive simulation study show that our tests perform reasonably well in finite samples.Finally, we apply our tests to the Tecator and AEMET datasets to check whether the assumption of linearity is supported by these datasets.

math.ST

Unified Inference on Moment Restrictions with Nuisance Parameters

This paper proposes a simple unified inference approach on moment restrictions in the presence of nuisance parameters. The proposed test is constructed based on a new characterization that avoids the estimation of nuisance parameters and can be broadly applied across diverse settings. Under suitable conditions, the test is shown to be asymptotically size controlled and consistent for both independent and dependent samples. Monte Carlo simulations show that the test performs well in finite samples. Numerical results from the application to conditional moment restriction models with weak instruments demonstrate that the proposed method may improve upon existing approaches in the literature.

stat.ME