SearcharxivSearch

arXiv subjects

Yiou Li

Publications and source records attributed to Yiou Li.

6 recordsLinked to original sources

On Efficient Design of Pilot Experiment for Generalized Linear Models

The experimental design for a generalized linear model (GLM) is important but challenging since the design criterion often depends on model specification including the link function, the linear predictor, and the unknown regression coefficients. Prior to constructing locally or globally optimal designs, a pilot experiment is usually conducted to provide some insights on the model specifications. In pilot experiments, little information on the model specification of GLM is available. Surprisingly, there is very limited research on the design of pilot experiments for GLMs. In this work, we obtain some theoretical understanding of the design efficiency in pilot experiments for GLMs. Guided by the theory, we propose to adopt a low-discrepancy design with respect to some target distribution for pilot experiments. The performance of the proposed design is assessed through several numerical examples.

stat.ME

A Maximin $\Phi_{p}$-Efficient Design for Multivariate GLM

Experimental designs for a generalized linear model (GLM) often depend on the specification of the model, including the link function, the predictors, and unknown parameters, such as the regression coefficients. To deal with uncertainties of these model specifications, it is important to construct optimal designs with high efficiency under such uncertainties. Existing methods such as Bayesian experimental designs often use prior distributions of model specifications to incorporate model uncertainties into the design criterion. Alternatively, one can obtain the design by optimizing the worst-case design efficiency with respect to uncertainties of model specifications. In this work, we propose a new Maximin $\Phi_p$-Efficient (or Mm-$\Phi_p$ for short) design which aims at maximizing the minimum $\Phi_p$-efficiency under model uncertainties. Based on the theoretical properties of the proposed criterion, we develop an efficient algorithm with sound convergence properties to construct the Mm-$\Phi_p$ design. The performance of the proposed Mm-$\Phi_p$ design is assessed through several numerical examples.

stat.ME

Covariate Balancing Based on Kernel Density Estimates for Controlled Experiments

Controlled experiments are widely used in many applications to investigate the causal relationship between input factors and experimental outcomes. A completely randomized design is usually used to randomly assign treatment levels to experimental units. When covariates of the experimental units are available, the experimental design should achieve covariate balancing among the treatment groups, such that the statistical inference of the treatment effects is not confounded with any possible effects of covariates. However, covariate imbalance often exists, because the experiment is carried out based on a single realization of the complete randomization. It is more likely to occur and worsen when the size of the experimental units is small or moderate. In this paper, we introduce a new covariate balancing criterion, which measures the differences between kernel density estimates of the covariates of treatment groups. To achieve covariate balance before the treatments are randomly assigned, we partition the experimental units by minimizing the criterion, then randomly assign the treatment levels to the partitioned groups. Through numerical examples, we show that the proposed partition approach can improve the accuracy of the difference-in-mean estimator and outperforms the complete randomization and rerandomization approaches.

stat.ME

Is a Transformed Low Discrepancy Design Also Low Discrepancy?

Experimental designs intended to match arbitrary target distributions are typically constructed via a variable transformation of a uniform experimental design. The inverse distribution function is one such transformation. The discrepancy is a measure of how well the empirical distribution of any design matches its target distribution. This chapter addresses the question of whether a variable transformation of a low discrepancy uniform design yields a low discrepancy design for the desired target distribution. The answer depends on the two kernel functions used to define the respective discrepancies. If these kernels satisfy certain conditions, then the answer is yes. However, these conditions may be undesirable for practical reasons. In such a case, the transformation of a low discrepancy uniform design may yield a design with a large discrepancy. We illustrate how this may occur. We also suggest some remedies. One remedy is to ensure that the original uniform design has optimal one-dimensional projection, but this remedy works best if the design is dense, or in other words, the ratio of sample size divided by the dimension of the random variable is relatively large. Another remedy is to use the transformed design as the input to a coordinate-exchange algorithm that optimizes the desired discrepancy, and this works for both dense or sparse designs. The effectiveness of these two remedies is illustrated via simulation.

stat.CO

An Efficient Algorithm for Elastic I-optimal Design of Generalized Linear Models

The generalized linear models (GLMs) are widely used in statistical analysis and the related design issues are undoubtedly challenging. The state-of-the-art works mostly apply to design criteria on the estimates of regression coefficients. The prediction accuracy is usually critical in modern decision making and artificial intelligence applications. It is of importance to study optimal designs from the prediction aspects for generalized linear models. In this work, we consider the Elastic I-optimality as a prediction-oriented design criterion for generalized linear models, and develop efficient algorithms for such $\text{EI}$-optimal designs. By investigating theoretical properties for the optimal weights of any set of design points and extending the general equivalence theorem to the $\text{EI}$-optimality for GLMs, the proposed efficient algorithm adequately combines the Fedorov-Wynn algorithm and multiplicative algorithm. It achieves great computational efficiency with guaranteed convergence. Numerical examples are conducted to evaluate the feasibility and computational efficiency of the proposed algorithm.

stat.ME

Kernel Discrepancy-Based Rerandomization for Controlled Experiments

This paper introduces a kernel discrepancy-based framework for rerandomization to enhance the precision of causal inference in controlled experiments. We demonstrate that the kernel discrepancy is the key part of the variance upper bound for the difference-in-means estimator, thereby establishing a theoretical rationale for its use. It quantifies the difference between empirical covariate distributions of treatment groups. We can choose a suitable kernel function and the corresponding discrepancy to accommodate simple or complex relationships between the outcome and the covariates. The proposed framework efficiently applies to any number of treatment groups, overcoming a significant limitation of existing methods. Furthermore, we develop a computationally efficient composite strategy for factorial experiments by recursively applying two- or multi-group rerandomizations. Numerical studies demonstrate that our approach significantly reduces estimator variance, with the linear kernel being optimal for linear relationships and the $\mathcal{L}_2$-discrepancy offering robust performance under model uncertainty.

stat.ME