SearcharxivSearch

arXiv subjects

Ming-Chung Chang

Publications and source records attributed to Ming-Chung Chang.

5 recordsLinked to original sources

Nearly Optimal Subdata Selection

When, in terms of the number of data points, the size of a dataset exceeds available computing resources, or when labeling is expensive, an attractive solution consists of selecting only some of the data points (subdata) for further consideration. A central question for selecting subdata of size $n$ from $N$ available data points is which $n$ points to select. While an answer to this question depends on the objective, one approach for a parametric model and a focus on parameter estimation is to select subdata that retains maximal information. Identifying such subdata is a classical NP-hard problem due to its inherent discreteness. Based on optimal approximate design theory, we develop a new methodology for information-based subdata selection, resulting in subdata that approaches the optimal solution. To achieve this, we develop a novel algorithm that applies to a general model, accommodates arbitrary choices of $N$ and $n$, and supports multiple optimality criteria, and we prove its convergence. Moreover, the new methodology facilitates an assessment of the efficiency of subdata selected by any method by obtaining tight lower and upper bounds for the efficiency. We show that the subdata obtained through the new methodology is highly efficient and outperforms all existing methods.

stat.ME

Generative and Nonparametric Approaches for Conditional Distribution Estimation: Methods, Perspectives, and Comparative Evaluations

The inference of conditional distributions is a fundamental problem in statistics, essential for prediction, uncertainty quantification, and probabilistic modeling. A wide range of methodologies have been developed for this task. This article reviews and compares several representative approaches spanning classical nonparametric methods and modern generative models. We begin with the single-index method of Hall and Yao (2005), which estimates the conditional distribution through a dimension-reducing index and nonparametric smoothing of the resulting one-dimensional cumulative conditional distribution function. We then examine the basis-expansion approaches, including FlexCode (Izbicki and Lee, 2017) and DeepCDE (Dalmasso et al., 2020), which convert conditional density estimation into a set of nonparametric regression problems. In addition, we discuss two recent generative simulation-based methods that leverage modern deep generative architectures: the generative conditional distribution sampler (Zhou et al., 2023) and the conditional denoising diffusion probabilistic model (Fu et al., 2024; Yang et al., 2025). A systematic numerical comparison of these approaches is provided using a unified evaluation framework that ensures fairness and reproducibility. The performance metrics used for the estimated conditional distribution include the mean-squared errors of conditional mean and standard deviation, as well as the Wasserstein distance. We also discuss their flexibility and computational costs, highlighting the distinct advantages and limitations of each approach.

stat.ML

Accumulated Aggregated D-Optimal Designs for Estimating Main Effects in Black-Box Models

Estimating how individual input variables affect the output of a black-box model is a central task in explainable machine learning. However, existing methods suffer from two key limitations: sensitivity to out-of-distribution (OOD) evaluations, which arises when query points are placed far from the data manifold, and instability under feature correlation, which can lead to unreliable effect estimates in practice. We introduce a unified view of main effect estimation as a design problem, which reveals that all existing methods differ only in their choice of evaluation locations. Building on this formulation, we propose A2D2E, an Estimator based on Accumulated Aggregated D-Optimal Designs, which replaces evaluations with a D-optimal hypercube design to minimize the variance of main effect estimation. A2D2E is model-agnostic, requires no differentiability of the predictor, and admits a closed-form estimator with complexity comparable to existing approaches. We establish that A2D2E is consistent to the same population target as ALE, and extend this result to the realistic setting where only a surrogate model is available. Through extensive simulations across multiple predictive models and dependence settings, we demonstrate that A2D2E outperforms ALE-based methods, with the largest gains under high feature correlation.

stat.ML

Design Selection for Two-Level Multi-Stratum Factorial Experiments Based on Swarm Intelligence Optimization

For unstructured experimental units, the minimum aberration due to Fries and Hunter (1980) is a popular criterion for choosing regular fractional factorial designs. Following which, many related studies have focused on multi-stratum factorial designs, in which multiple error terms arise from the complicated structures of experimental units. Chang and Cheng (2018) proposed a Bayesian-inspired aberration criterion for selecting multi-stratum factorial designs, which can be considered as a generalized version of that in Fries and Hunter (1980). However, they did not propose algorithms for searching for minimum aberration designs. The particle swarm optimization (PSO) algorithm is a popular optimization method that has been widely used in various applications. In this paper, we propose a new version of the PSO to select regular as well as nonregular multi-stratum designs. To select regular ones, we treat defining words as particles in the PSO and link the PSO with design key matrices. For nonregular multi-stratum designs, we treat treatment combinations as particles in the PSO. Several numerical illustrations are provided.

stat.ME

An aberration criterion for conditional models

Conditional models with one pair of conditional and conditioned factors in Mukerjee et al. (2017) are extended to two pairs in this paper. The extension includes the parametrization, effect hierarchy, sufficient conditions for universal optimality, aberration, complementary set theory and the strategy for finding minimum aberration designs. A catalog of 16-run minimum aberration designs under conditional models is provided. For five to twelve factors, all 16-run minimum aberration designs under conditional models are also minimum aberration under traditional models.

stat.ME