SearcharxivSearch

arXiv subjects

Kosuke Morikawa

Publications and source records attributed to Kosuke Morikawa.

17 recordsLinked to original sources

TS-Neyman: Posterior Sampling for Adaptive Stratified Estimation

Many model evaluation tasks reduce to estimating an average loss, error rate, or subgroup metric on a stratified pool when each label, human rating, or simulator call is costly. The precision-optimal Neyman allocation depends on within-stratum variances, which must be learned from the same observations used for estimation. We formulate this as a sequential allocation problem and use the exact one-step marginal variance reduction as the priority index. Replacing the unknown variances by independent inverse-chi-squared posterior draws yields TS-Neyman, a Thompson-sampling rule that preserves the oracle marginal-gain structure while randomizing over variance uncertainty. For any fixed finite number of strata, we prove almost-sure convergence of the TS-Neyman allocation proportions to the Neyman target, asymptotic optimality of the variance proxy, and a central limit theorem for the resulting adaptive stratified estimator. In two five-stratum budget-scaling benchmarks, one bounded-loss benchmark and one binary model-error benchmark in the spirit of Dai et al. 2023, TS-Neyman's relative efficiency stays within 5 percent of the oracle on the bounded-loss population and within about 15 percent on the binary benchmark. In an additional CivilComments real-data replay with confidence-based strata, it stays within about 8 percent of the oracle and improves on equal allocation by roughly 7 to 14 percent in MSE across budgets, while plug-in greedy and two-stage plug-in can degrade by over an order of magnitude under sparse pilots. Common-pilot warm-start and prior-sensitivity studies show that this behavior is stable under working-model and working-prior misspecification.

stat.ME

Semiparametric Efficient Data Integration Using the Dual-Frame Sampling Framework

Integrating probability and non-probability samples is increasingly important, yet unknown non-probability inclusion mechanisms complicate identification and efficient estimation. We develop semiparametric theory for dual-frame data integration and propose two complementary estimators. The first models the conditional inclusion probability parametrically and, under correct specification, attains the semiparametric efficiency bound. The probability sample serves as an outcome-observed validation sample, identifying the conditional inclusion parameters without instrumental variables even under informative selection. We derive the efficient influence function while leaving the conditional distribution of the design inclusion probability unrestricted. The resulting estimating equations satisfy exact augmentation invariance, which supports both same-sample sieve estimation and pooled cross-fitting under weak nuisance-rate conditions. The second estimator, motivated by a two-stage sampling approximation, does not require a non-probability inclusion model; although not fully efficient, it is efficient within a restricted augmentation class, is robust to misspecification of the non-probability inclusion model, and its oracle version weakly dominates optimally augmented probability-sample-only inference. Simulations and a public-data repeated-sampling study show efficiency gains under correct specification and stable performance under misspecification and weak identification. Across the main simulations, the efficient procedures substantially reduce RMSE relative to probability-sample-only estimation, while the restricted estimator remains stable under inclusion-model misspecification and weak identification.

stat.ME

Data-Adaptive Integration with External Summary Data for Outcome Mean Estimation

Combining an internal individual-level study with readily available external summary statistics promises major efficiency gains at minimal additional cost, yet heterogeneity between sources can bias estimates for the internal target population. We develop a generalized entropy-balancing integration strategy that calibrates the internal individual-level sample to externally reported moments while retaining the internal population as the target, explicitly permitting a biased external sample. The weighted-regression version of our estimator is doubly robust: it remains consistent when either the outcome-regression model or the entropy-balancing model is correctly specified. When multiple balancing specifications are plausible, we introduce a data-adaptive entropy-family selection rule. For the final borrowing decision, we propose a bootstrap-based criterion comparing stabilized mean squared error (MSE) estimates for the selected entropy-balancing estimator and the internal sample mean. This criterion is selection consistent under fixed alternatives and reverts to the internal estimator when a nonvanishing bias is detected. Separately, under a linear homoscedastic benchmark, the asymptotic efficiency criteria admit geometric interpretations through the Mahalanobis distance and Pearson chi-squared divergence. The entropy-balancing estimators and numerical-experiment routines are implemented in the R package daisy. Simulations show stable MSE reductions for the weighted-regression estimator across calibrated distributional shifts and the predicted reversion toward the internal estimator under fixed simultaneous misspecification as the sample size increases. An application to nationwide public-access defibrillation records in Japan illustrates the resulting MSE-based borrowing decision.

stat.ME

Robust Estimation and Model Selection for the Controlled Directed Effect with Unmeasured Mediator-Outcome Confounders

Controlled Direct Effect (CDE) is one of the causal estimands used to evaluate both exposure and mediation effects on an outcome. When there are unmeasured confounders existing between the mediator and the outcome, the ordinary identification assumption does not work. In this manuscript, we consider an identification condition to identify CDE in the presence of unmeasured confounders. The key assumptions are: 1) the random allocation of the exposure, and 2) the existence of instrumental variables directly related to the mediator. Under these conditions, we propose a novel doubly robust estimation method, which work well if either the propensity score model or the baseline outcome model is correctly specified. Additionally, we propose a Generalized Information Criterion (GIC)-based model selection criterion for CDE that ensures model selection consistency. Our proposed procedure and related methods are applied to both simulation and real datasets to confirm the performance of these methods. Our proposed method can select the correct model with high probability and accurately estimate CDE.

stat.ME

Efficient Multiple-Robust Estimation for Nonresponse Data Under Informative Sampling

Nonresponse after probability sampling is a universal challenge in survey sampling, often necessitating adjustments to mitigate sampling and selection bias simultaneously. This study explored the removal of bias and effective utilization of available information, not just in nonresponse but also in the scenario of data integration, where summary statistics from other data sources are accessible. We reformulate these settings within a two-step monotone missing data framework, where the first step of missingness arises from sampling and the second originates from nonresponse. Subsequently, we derive the semiparametric efficiency bound for the target parameter. We also propose adaptive estimators utilizing methods of moments and empirical likelihood approaches to attain the lower bound. The proposed estimator exhibits both efficiency and double robustness. However, attaining efficiency with an adaptive estimator requires the correct specification of certain working models. To reinforce robustness against the misspecification of working models, we extend the property of double robustness to multiple robustness by proposing a two-step empirical likelihood method that effectively leverages empirical weights. A numerical study is undertaken to investigate the finite-sample performance of the proposed methods. We further applied our methods to a dataset from the National Health and Nutrition Examination Survey data by efficiently incorporating summary statistics from the National Health Interview Survey data.

stat.ME

Verifiable identification condition for nonignorable nonresponse data with categorical instrumental variables

We consider a model identification problem in which an outcome variable contains nonignorable missing values. Statistical inference requires a guarantee of the model identifiability to obtain estimators enjoying theoretically reasonable properties such as consistency and asymptotic normality. Recently, instrumental or shadow variables, combined with the completeness condition in the outcome model, have been highlighted to make a model identifiable. However, the completeness condition may not hold even for simple models when the instrument is categorical. We propose a sufficient condition for model identifiability, which is applicable to cases where establishing the completeness condition is difficult. Using observed data, we demonstrate that the proposed conditions are easy to check for many practical models and outline their usefulness in numerical experiments and real data analysis.

stat.ME

An empirical likelihood approach to reduce selection bias in voluntary samples

We address the weighting problem in voluntary samples under a nonignorable sample selection model. Under the assumption that the sample selection model is correctly specified, we can compute a consistent estimator of the model parameter and construct the propensity score estimator of the population mean. We use the empirical likelihood method to construct the final weights for voluntary samples by incorporating the bias calibration constraints and the benchmarking constraints. Linearization variance estimation of the proposed method is developed. A limited simulation study is also performed to check the performance of the proposed methods.

stat.ME

Semiparametric imputation using latent sparse conditional Gaussian mixtures for multivariate mixed outcomes

This paper proposes a flexible Bayesian approach to multiple imputation using conditional Gaussian mixtures. We introduce novel shrinkage priors for covariate-dependent mixing proportions in the mixture models to automatically select the suitable number of components used in the imputation step. We develop an efficient sampling algorithm for posterior computation and multiple imputation via Markov Chain Monte Carlo methods. The proposed method can be easily extended to the situation where the data contains not only continuous variables but also discrete variables such as binary and count values. We also propose approximate Bayesian inference for parameters defined by loss functions based on posterior predictive distributing of missing observations, by extending bootstrap-based Bayesian inference for complete data. The proposed method is demonstrated through numerical studies using simulated and real data.

stat.ME

Semiparametric adaptive estimation under informative sampling

In survey sampling, survey data do not necessarily represent the target population, and the samples are often biased. However, information on the survey weights aids in the elimination of selection bias. The Horvitz-Thompson estimator is a well-known unbiased, consistent, and asymptotically normal estimator; however, it is not efficient. Thus, this study derives the semiparametric efficiency bound for various target parameters by considering the survey weight as a random variable and consequently proposes a semiparametric optimal estimator with certain working models on the survey weights. The proposed estimator is consistent, asymptotically normal, and efficient in a class of the regular and asymptotically linear estimators. Further, a limited simulation study is conducted to investigate the finite sample performance of the proposed method. The proposed method is applied to the 1999 Canadian Workplace and Employee Survey data.

stat.ME

Identification enhanced generalised linear model estimation with nonignorable missing outcomes

Missing data often result in undesirable bias and loss of efficiency. These issues become substantial when the response mechanism is nonignorable, meaning that the response model depends on unobserved variables. To manage nonignorable nonresponse, it is necessary to estimate the joint distribution of unobserved variables and response indicators. However, model misspecification and identification issues can prevent robust estimates, even with careful estimation of the target joint distribution. In this study, we modeled the distribution of the observed parts and derived sufficient conditions for model identifiability, assuming a logistic regression model as the response mechanism and generalized linear models as the main outcome model of interest. More importantly, the derived sufficient conditions do not require any instrumental variables, which are often assumed to guarantee model identifiability but cannot be practically determined beforehand. To analyze missing data in applications, we propose practical guidelines and sensitivity analysis to determine the response mechanism. Furthermore, we present the performance of the proposed estimators in numerical studies and apply the proposed method to two sets of real data: exit polls from the 19th South Korean election and public data collected from the Korean Survey of Household Finances and Living Conditions.

stat.ME

Adjusting for publication bias in meta-analysis via inverse probability weighting using clinical trial registries

Publication bias is a major concern in conducting systematic reviews and meta-analyses. Various sensitivity analysis or bias-correction methods have been developed based on selection models and they have some advantages over the widely used bias-correction method of the trim-and-fill method. However, likelihood methods based on selection models may have difficulty in obtaining precise estimates and reasonable confidence intervals or require a complicated sensitivity analysis process. In this paper, we develop a simple publication bias adjustment method utilizing information on conducted but still unpublished trials from clinical trial registries. We introduce an estimating equation for parameter estimation in the selection function by regarding the publication bias issue as a missing data problem under missing not at random. With the estimated selection function, we introduce the inverse probability weighting (IPW) method to estimate the overall mean across studies. Furthermore, the IPW versions of heterogeneity measures such as the between-study variance and the I2 measure are proposed. We propose methods to construct asymptotic confidence intervals and suggest intervals based on parametric bootstrapping as an alternative. Through numerical experiments, we observed that the estimators successfully eliminate biases and the confidence intervals had empirical coverage probabilities close to the nominal level. On the other hand, the asymptotic confidence interval is much wider in some scenarios than the bootstrap confidence interval. Therefore, the latter is recommended for practical use.

stat.ME

Forecasting temporal variation of aftershocks immediately after a main shock using Gaussian process regression

Uncovering the distribution of magnitudes and arrival times of aftershocks is a key to comprehend the characteristics of the sequence of earthquakes, which enables us to predict seismic activities and hazard assessments. However, identifying the number of aftershocks immediately after the main shock is practically difficult due to contaminations of arriving seismic waves. To overcome the difficulty, we construct a likelihood based on the detected data incorporating a detection function to which the Gaussian process regression (GPR) is applied. The GPR is capable of estimating not only the parameters of the distribution of aftershocks together with the detection function but also credible intervals for both of the parameters and the detection function. A property that distributions of both the Gaussian process and aftershocks are exponential functions leads to an efficient Bayesian computational algorithm to estimate the hyperparameters. After the validations through numerical tests, the proposed method is retrospectively applied to the catalog data related to the 2004 Chuetsu earthquake towards early forecasting of the aftershocks. The result shows that the proposed method stably estimates the parameters of the distribution simultaneously their credible intervals even within three hours after the main shock.

physics.geo-ph

Bayesian Semiparametric Modeling of Response Mechanism for Nonignorable Missing Data

Statistical inference with nonresponse is quite challenging, especially when the response mechanism is nonignorable. In this case, the validity of statistical inference depends on untestable correct specification of the response model. To avoid the misspecification, we propose semiparametric Bayesian estimation in which an outcome model is parametric, but the response model is semiparametric in that we do not assume any parametric form for the nonresponse variable. We adopt penalized spline methods to estimate the unknown function. We also consider a fully nonparametric approach to modeling the response mechanism by using radial basis function methods. Using Polya-gamma data augmentation, we developed an efficient posterior computation algorithm via Gibbs sampling in which most full conditional distributions can be obtained in familiar forms. The performance of the proposed method is demonstrated in simulation studies and an application to longitudinal data.

stat.ME

A Profile Likelihood Approach to Semiparametric Estimation with Nonignorable Nonresponse

Statistical inference with nonresponse is quite challenging, especially when the response mechanism is nonignorable. The existing methods often require correct model specifications for both outcome and response models. However, due to nonresponse, both models cannot be verified from data directly and model misspecification can lead to a seriously biased inference. To overcome this limitation, we develop a robust and efficient semiparametric method based on the profile likelihood. The proposed method uses the robust semiparametric response model, in which fully unspecified function of study variable is assumed. An efficient computation algorithm using fractional imputation is developed. A quasi-likelihood approach for testing ignorability is also developed. The consistency and asymptotic normality of the proposed method are established. The finite-sample performance is examined in the extensive simulation studies and an application to the Korean Labor and Income Panel Study dataset is also presented.

stat.ME

Semiparametric Optimal Estimation With Nonignorable Nonresponse Data

When the response mechanism is believed to be not missing at random (NMAR), a valid analysis requires stronger assumptions on the response mechanism than standard statistical methods would otherwise require. Semiparametric estimators have been developed under the model assumptions on the response mechanism. In this paper, a new statistical test is proposed to guarantee model identifiability without using any instrumental variable. Furthermore, we develop optimal semiparametric estimation for parameters such as the population mean. Specifically, we propose two semiparametric optimal estimators that do not require any model assumptions other than the response mechanism. Asymptotic properties of the proposed estimators are discussed. An extensive simulation study is presented to compare with some existing methods. We present an application of our method using Korean Labor and Income Panel Survey data.

stat.ME

Statistical Inference with Different Missing-data Mechanisms

When data are missing due to at most one cause from some time to next time, we can make sampling distribution inferences about the parameter of the data by modeling the missing-data mechanism correctly. Proverbially, in case its mechanism is missing at random (MAR), it can be ignored, but in case not missing at random (NMAR), it can not be. There are no methods, however, to analyze when missing of the data can occur because of several causes despite of there being many such data in practice. Hence the aim of this paper is to propose how to inference on such data. Concretely, we extend the missing-data indicator from usual binary random vectors to discrete random vectors, define missing-data mechanism for every causes and research ignorability of a mixture of missing-data mechanisms such as "MAR & MAR" and "MAR & NMAR". In particular, when the combination of mechanisms is "MAR & NMAR", generally the component of MAR can not be ignored, but in special case, it can be.

stat.ME

Identification Problem for The Analysis of Binary Data with Non-ignorable Missing

When a missing-data mechanism is NMAR or non-ignorable, missingness is itself vital information and it must be taken into the likelihood, which, however, needs to introduce additional parameters to be estimated. The incompleteness of the data and introduction of more parameters can cause the identification problem. When a response variable is binary, it becomes a more serious problem because of less information of bi- nary data, however, there are no methods to briefly verify whether a mode is identified or not. Therefore, we provide a new necessary and sufficient condition to easily check model identifiability when analyzing binary data with non-ignorable missing by condi- tional models. This condition can give us what condition is needed for a model to have identifiability as well as make easily check the identifiability of a model.

stat.ME