Searcharxiv⌕ Search

arXiv subjects

Yoshiyuki Ninomiya

Publications and source records attributed to Yoshiyuki Ninomiya.

13 recordsLinked to original sources

Covariate balancing estimation and model selection for difference-in-differences approach

Remarkable progress has been made in difference-in-differences (DID) approaches to causal inference that estimate the average effect of a treatment on the treated (ATT). Of these, the semiparametric DID (SDID) approach incorporates a propensity score analysis into the DID setup. Supposing that the ATT is a function of covariates, we estimate it by weighting the inverse of the propensity score. In this study, as one way to make the estimation robust to the propensity score modeling, we incorporate covariate balancing. Then, by attentively constructing the moment conditions used in the covariate balancing, we show that the proposed estimator is doubly robust. In addition to the estimation, we also address model selection. In practice, covariate selection is an essential task in statistical analysis, but even in the basic setting of the SDID approach, there are no reasonable information criteria. Here, we derive a model selection criterion as an asymptotically bias-corrected estimator of risk based on the loss function used in the SDID estimation. We show that a penalty term can be derived that is considerably different from almost twice the number of parameters that often appears in AIC-type information criteria. Numerical experiments show that the proposed method estimates the ATT more robustly compared with the method using propensity scores given by maximum likelihood estimation, and that the proposed criterion clearly reduces the risk targeted in the SDID approach in comparison with the intuitive generalization of the existing information criterion. In addition, real data analysis confirms that there is a large difference between the results of the proposed method and those of the existing method.

stat.ME↗

Developing an information criterion for spatial data analysis through Bayesian generalized fused lasso

In the field of spatial data analysis, spatially varying coefficients (SVC) models, which allow regression coefficients to vary by region and flexibly capture spatial heterogeneity, have continued to be developed in various directions. Moreover, the Bayesian generalized fused lasso is often used as a method that efficiently provides estimation under the natural assumption that regression coefficients of adjacent regions tend to take the same value. In most Bayesian methods, the selection of prior distribution is an essential issue, and in the setting of SVC model with the Bayesian generalized fused lasso, determining the complexity of the class of prior distributions is also a challenging aspect, further amplifying the difficulty of the problem. For example, the widely applicable information criterion (WAIC), which has become standard in Bayesian model selection, does not target determining the complexity. Therefore, in this study, we adapted a criterion called the prior intensified information criterion (PIIC) to this setting. Specifically, under an asymptotic setting that retains the influence of the prior distribution, that is, under an asymptotic setting that deliberately does not provide selection consistency, we derived the asymptotic properties of our generalized fused lasso estimator. Then, based on these properties, we constructed an information criterion as an asymptotically bias-corrected estimator of predictive risk. In numerical experiments, we confirmed that PIIC outperforms WAIC in the sense of reducing the predictive risk, and in a real data analysis, we observed that the two criteria give rise to substantially different results.

stat.ME↗

Akaike information criterion for segmented regression models

In segmented regression, when the regression function is continuous at the change-points that are the boundaries of the segments, it is also called joinpoint regression, and the analysis package developed by \cite{KimFFM00} has become a standard tool for analyzing trends in longitudinal data in the field of epidemiology. In addition, it is sometimes natural to expect the regression function to be discontinuous at the change-points, and in the field of epidemiology, this model is used in \cite{JiaZS22}, which is considered important due to the analysis of COVID-19 data. On the other hand, model selection is also indispensable in segmented regression, including the estimation of the number of change-points; however, it can be said that only BIC-type information criteria have been developed. In this paper, we derive an information criterion based on the original definition of AIC, aiming to minimize the divergence between the true structure and the estimated structure. Then, using the statistical asymptotic theory specific to the segmented regression, we confirm that the penalty for the change-point parameter is 6 in the discontinuous case. On the other hand, in the continuous case, we show that the penalty for the change-point parameter remains 2 despite the rapid change in the derivative coefficients. Through numerical experiments, we observe that our AIC tends to reduce the divergence compared to BIC. In addition, through analyzing the same real data as in \cite{JiaZS22}, we find that the selection between continuous and discontinuous using our AIC yields new insights and that our AIC and BIC may yield different results.

stat.ME↗

Doubly Robust Criterion for Causal Inference

The semiparametric estimation approach, which includes inverse-probability-weighted and doubly robust estimation using propensity scores, is a standard tool in causal inference, and it is rapidly being extended in various directions. On the other hand, although model selection is indispensable in statistical analysis, an information criterion for selecting an appropriate regression structure has just started to be developed. In this paper, based on the original definition of Akaike information criterion (AIC; \citealt{Aka73}), we derive an AIC-type criterion for propensity score analysis. Here, we define a risk function based on the Kullback-Leibler divergence as the cornerstone of the information criterion and treat a general causal inference model that is not necessarily a linear one. The causal effects to be estimated are those in the general population, such as the average treatment effect on the treated or the average treatment effect on the untreated. In light of the fact that this field attaches importance to doubly robust estimation, which allows either the model of the assignment variable or the model of the outcome variable to be wrong, we make the information criterion itself doubly robust so that either one can be wrong and it will still be an asymptotically unbiased estimator of the risk function. In simulation studies, we compare the derived criterion with an existing criterion obtained from a formal argument and confirm that the former outperforms the latter. Specifically, we check that the divergence between the estimated structure from the derived criterion and the true structure is clearly small in all simulation settings and that the probability of selecting the true or nearly true model is clearly higher. Real data analyses confirm that the results of variable selection using the two criteria differ significantly.

stat.ME↗

Prior Intensified Information Criterion

The widely applicable information criterion (WAIC) has been used as a model selection criterion for Bayesian statistics in recent years. It is an asymptotically unbiased estimator of the Kullback-Leibler divergence between a Bayesian predictive distribution and the true distribution. Not only is the WAIC theoretically more sound than other information criteria, its usefulness in practice has also been reported. On the other hand, the WAIC is intended for settings in which the prior distribution does not have an asymptotic influence, and as we set the class of the prior distribution to be more complex, it never fails to select the most complex one. To alleviate these concerns, this paper proposed the prior intensified information criterion (PIIC). In addition, it customizes this criterion to incorporate sparse estimation and causal inference. Numerical experiments show that the PIIC clearly outperforms the WAIC in terms of prediction performance when the above concerns are manifested. A real data analysis confirms that the results of variable selection and Bayesian estimators of the WAIC and PIIC differ significantly.

stat.ME↗

Information criteria for sparse methods in causal inference

For propensity score analysis and sparse estimation, we develop an information criterion for determining the regularization parameters needed in variable selection. First, for Gaussian distribution-based causal inference models, we extend Stein's unbiased risk estimation theory, which leads to a generalized Cp criterion that has almost no weakness in conventional sparse estimation, and derive an inverse-probability-weighted sparse estimation version of the criterion without resorting to asymptotics. Next, for general causal inference models that are not necessarily Gaussian distribution-based, we extend the asymptotic theory on LASSO for propensity score analysis, with the intention of implementing doubly robust sparse estimation. From the asymptotic theory, an AIC-type information criterion for inverse-probability-weighted sparse estimation is given, and then a criterion with double robustness in itself is derived for doubly robust sparse estimation. Numerical experiments compare the proposed criterion with the existing criterion derived from a formal argument and verify that the proposed criterion is superior in almost all cases, that the difference is not negligible in many cases, and that the results of variable selection differ significantly. Real data analysis confirms that the difference between variable selection and estimation by these criteria is actually large. Finally, generalizations to general sparse estimation using group LASSO, elastic net, and non-convex regularization are made in order to indicate that the proposed criterion is highly extensible.

stat.ME↗

Information criteria for detecting change-points in the Cox proportional hazards model

The Cox proportional hazards model, commonly used in clinical trials, assumes proportional hazards. However, it does not hold when, for example, there is a delayed onset of the treatment effect. In such a situation, an acute change in the hazard ratio function is expected to exist. This paper considers the Cox model with change-points and derives AIC-type information criteria for detecting those change-points. The change-point model does not allow for conventional statistical asymptotics due to its irregularity, thus a formal AIC that penalizes twice the number of parameters would not be analytically derived, and using it would clearly give overfitting analysis results. Therefore, we will construct specific asymptotics using the partial likelihood estimation method in the Cox model with change-points. Based on the original derivation method for AIC, we propose information criteria that are mathematically guaranteed. If the partial likelihood is used in the estimation, information criteria with penalties much larger than twice the number of parameters could be obtained in an explicit form. Numerical experiments confirm that the proposed criterion is clearly superior in terms of the original purpose of AIC, which is to provide an estimate that is close to the true structure. We also apply the proposed criterion to actual clinical trial data to indicate that it will easily lead to different results from the formal AIC.

stat.ME↗

Selective Inference in Propensity Score Analysis

Selective inference (post-selection inference) is a methodology that has attracted much attention in recent years in the fields of statistics and machine learning. Naive inference based on data that are also used for model selection tends to show an overestimation, and so the selective inference conditions the event that the model was selected. In this paper, we develop selective inference in propensity score analysis with a semiparametric approach, which has become a standard tool in causal inference. Specifically, for the most basic causal inference model in which the causal effect can be written as a linear sum of confounding variables, we conduct Lasso-type variable selection by adding an $\ell_1$ penalty term to the loss function that gives a semiparametric estimator. Confidence intervals are then given for the coefficients of the selected confounding variables, conditional on the event of variable selection, with asymptotic guarantees. An important property of this method is that it does not require modeling of nonparametric regression functions for the outcome variables, as is usually the case with semiparametric propensity score analysis.

stat.ME↗

Smoothly varying ridge regularization

A basis expansion with regularization methods is much appealing to the flexible or robust nonlinear regression models for data with complex structures. When the underlying function has inhomogeneous smoothness, it is well known that conventional reguralization methods do not perform well. In this case, an adaptive procedure such as a free-knot spline or a local likelihood method is often introduced as an effective method. However, both methods need intensive computational loads. In this study, we consider a new efficient basis expansion by proposing a smoothly varying regularization method which is constructed by some special penalties. We call them adaptive-type penalties. In our modeling, adaptive-type penalties play key rolls and it has been successful in giving good estimation for inhomogeneous smoothness functions. A crucial issue in the modeling process is the choice of a suitable model among candidates. To select the suitable model, we derive an approximated generalized information criterion (GIC). The proposed method is investigated through Monte Carlo simulations and real data analysis. Numerical results suggest that our method performs well in various situations.

stat.ME↗

Use of spurious correlation for multiplicity adjustment

We consider one of the most basic multiple testing problems that compares expectations of multivariate data among several groups. As a test statistic, a conventional (approximate) $t$-statistic is considered, and we determine its rejection region using a common rejection limit. When there are unknown correlations among test statistics, the multiplicity adjusted $p$-values are dependent on the unknown correlations. They are usually replaced with their estimates that are always consistent under any hypothesis. In this paper, we propose the use of estimates, which are not necessarily consistent and are referred to as spurious correlations, in order to improve statistical power. Through simulation studies, we verify that the proposed method asymptotically controls the family-wise error rate and clearly provides higher statistical power than existing methods. In addition, the proposed and existing methods are applied to a real multiple testing problem that compares quantitative traits among groups of mice and the results are compared.

stat.ME↗

$C_p$ criterion for semiparametric approach in causal inference

For marginal structural models, which recently play an important role in causal inference, we consider a model selection problem in the framework of a semiparametric approach using inverse-probability-weighted estimation or doubly robust estimation. In this framework, the modeling target is a potential outcome which may be a missing value, and so we cannot apply the AIC nor its extended version to this problem. In other words, there is no analytical information criterion obtained according to its classical derivation for this problem. Hence, we define a mean squared error appropriate for treating the potential outcome, and then we derive its asymptotic unbiased estimator as a $C_{p}$ criterion from an asymptotics for the semiparametric approach and using an ignorable treatment assignment condition. In simulation study, it is shown that the proposed criterion exceeds a conventionally derived existing criterion in the squared error and model selection frequency. Specifically, in all simulation settings, the proposed criterion provides clearly smaller squared errors and higher frequencies selecting the true or nearly true model. Moreover, in real data analysis, we check that there is a clear difference between the selections by the two criteria.

stat.ME↗

On the Consistency of the Bias Correction Term of the AIC for the Non-Concave Penalized Likelihood Method

Penalized likelihood methods with an $\ell_γ$-type penalty, such as the Bridge, the SCAD, and the MCP, allow us to estimate a parameter and to do variable selection, simultaneously, if $γ\in (0,1]$. In this method, it is important to choose a tuning parameter which controls the penalty level, since we can select the model as we want when we choose it arbitrarily. Nowadays, several information criteria have been developed to choose the tuning parameter without such an arbitrariness. However the bias correction term of such information criteria depend on the true parameter value in general, then we usually plug-in a consistent estimator of it to compute the information criteria from the data. In this paper, we derive a consistent estimator of the bias correction term of the AIC for the non-concave penalized likelihood method and propose a simple AIC-type information criterion for such models.

stat.ME↗

AIC for Non-concave Penalized Likelihood Method

Non-concave penalized maximum likelihood methods, such as the Bridge, the SCAD, and the MCP, are widely used because they not only do parameter estimation and variable selection simultaneously but also have a high efficiency as compared to the Lasso. They include a tuning parameter which controls a penalty level, and several information criteria have been developed for selecting it. While these criteria assure the model selection consistency and so have a high value, it is a severe problem that there are no appropriate rules to choose the one from a class of information criteria satisfying such a preferred asymptotic property. In this paper, we derive an information criterion based on the original definition of the AIC by considering the minimization of the prediction error rather than the model selection consistency. Concretely speaking, we derive a function of the score statistic which is asymptotically equivalent to the non-concave penalized maximum likelihood estimator, and then we provide an asymptotically unbiased estimator of the Kullback-Leibler divergence between the true distribution and the estimated distribution based on the function. Furthermore, through simulation studies, we check that the performance of the proposed information criterion gives almost the same as or better than that of the cross-validation.

stat.ME↗