Searcharxiv⌕ Search

arXiv subjects

Abhyuday Mandal

Publications and source records attributed to Abhyuday Mandal.

At least 19 recordsLinked to original sources

Musings on Constructions of Optimal Latin Hypercube Designs with Flexible Sizes

Latin hypercube designs (LHDs) play an important role in computer experiments, offering flexible and efficient space-filling properties under a variety of optimality criteria, including maximin distance, maximum projection, and orthogonality. Constructing optimal LHDs with flexible sizes is challenging due to the limited availability of theoretical results, namely, algebraic constructions with provable optimality guarantees, and the rapidly expanding search space encountered by search algorithms. For design sizes outside the scope of known algebraic constructions, search-based algorithms are widely used, but they often involve substantial computational effort and lack assurances of global optimality. This paper provides a comprehensive review and comparison of current popular algebraic constructions for optimal LHDs and widely adopted search algorithms for generating high-quality LHDs. By reviewing their theoretical properties and empirical performance across a range of criteria, we offer a unified perspective on their relative strengths and limitations. The comparisons presented herein aim to assist practitioners in selecting appropriate design strategies for different objectives and constraints. The insights from this work also highlight open challenges and may serve as a benchmark for future development in optimal LHD construction.

stat.ME↗

p-PSO: A Penalized Particle Swarm Optimization Technique for Finding D-Optimal Designs with Mixed Factors in Generalized Linear Models

Finding D-optimal designs for generalized linear models (GLMs) is challenging due to the dependence of the Fisher information matrix on unknown parameters and the lack of closed-form solutions, particularly when input factors include both discrete and continuous variables. Although classical algorithms and recent metaheuristic approaches have offered partial solutions, there remains a need for robust and computationally efficient methods. In this paper, we propose a penalized Particle Swarm Optimization (PSO) approach, named $p$-PSO. Here we introduce a new, general-purpose penalty formulation for constrained optimization and demonstrate its effectiveness in optimal design problems. The formulation is algorithm-agnostic and applicable to a broad class of black-box optimization methods. Results show that the method is highly efficient, with its primary contribution being a penalty formulation that enables the direct use of an off-the-shelf PSO algorithm and extends naturally to more general constrained optimization tasks.

stat.ME↗

Credible Distributions of Overall Ranking of Entities

Ranking, and inferences based on ranking of a set of entities, are important problems in numerous contexts. This is especially true in small area statistics where there may be only a limited amount of directly observed data from each entity or small area, while precise and accurate estimates of best or worst performing entities are needed for fund allocation, planning and policymaking, stakeholder advocacy, evaluation of welfare programs, and so on. However, ranks estimates constructed exclusively on point estimates of parameters lack uncertainty quantification, and may lead to imbalances and inequities when these are based on small sample sizes. We propose novel Bayesian approaches to address this problem. Our proposals result in partitions of the parameter space with posterior distribution driven partial ordering of the sets in a partition. This in turn translates to a coherent probability mass function over ranks for every entity, and a coherent probability mass function over entities for every rank. Our Bayesian algorithms significantly outperform the state-of-the-art non-Bayesian alternatives, and are amenable to inclusion of covariates in the model as well as borrowing strengths across small areas. We evaluate our proposed Bayesian algorithms in terms of accuracy and stability using a number of applications and a simulation study. Additionally, we develop a novel theoretical framework for inference and ranking problems involving a triangular array of Fay-Herriot models and data, and provide probabilistic guarantees of performances of the proposed Bayesian ranking algorithms.

stat.ME↗

ForLion: A New Algorithm for D-optimal Designs under General Parametric Statistical Models with Mixed Factors

In this paper, we address the problem of designing an experimental plan with both discrete and continuous factors under fairly general parametric statistical models. We propose a new algorithm, named ForLion, to search for locally optimal approximate designs under the D-criterion. The algorithm performs an exhaustive search in a design space with mixed factors while keeping high efficiency and reducing the number of distinct experimental settings. Its optimality is guaranteed by the general equivalence theorem. We present the relevant theoretical results for multinomial logit models (MLM) and generalized linear models (GLM), and demonstrate the superiority of our algorithm over state-of-the-art design algorithms using real-life experiments under MLM and GLM. Our simulation studies show that the ForLion algorithm could reduce the number of experimental settings by 25% or improve the relative efficiency of the designs by 17.5% on average. Our algorithm can help the experimenters reduce the time cost, the usage of experimental devices, and thus the total cost of their experiments while preserving high efficiencies of the designs.

stat.CO↗

A General Equivalence Theorem for Crossover Designs under Generalized Linear Models

With the help of Generalized Estimating Equations, we identify locally D-optimal crossover designs for generalized linear models. We adopt the variance of parameters of interest as the objective function, which is minimized using constrained optimization to obtain optimal crossover designs. In this case, the traditional general equivalence theorem could not be used directly to check the optimality of obtained designs. In this manuscript, we derive a corresponding general equivalence theorem for crossover designs under generalized linear models.

stat.ME↗

Solving an Inverse Problem for Time Series Valued Computer Simulators via Multiple Contour Estimation

Computer simulators are often used as a substitute of complex real-life phenomena which are either expensive or infeasible to experiment with. This paper focuses on how to efficiently solve the inverse problem for an expensive to evaluate time series valued computer simulator. The research is motivated by a hydrological simulator which has to be tuned for generating realistic rainfall-runoff measurements in Athens, Georgia, USA. Assuming that the simulator returns g(x,t) over L time points for a given input x, the proposed methodology begins with a careful construction of a discretization (time-) point set (DPS) of size $k << L$, achieved by adopting a regression spline approximation of the target response series at k optimal knots locations $\{t^*_1, t^*_2, ..., t^*_k\}$. Subsequently, we solve k scalar valued inverse problems for simulator $g(x,t^*_j)$ via the contour estimation method. The proposed approach, named MSCE, also facilitates the uncertainty quantification of the inverse solution. Extensive simulation study is used to demonstrate the performance comparison of the proposed method with the popular competitors for several test-function based computer simulators and a real-life rainfall-runoff measurement model.

stat.ME↗

Modeling and Active Learning for Experiments with Quantitative-Sequence Factors

A new type of experiment that aims to determine the optimal quantities of a sequence of factors is eliciting considerable attention in medical science, bioengineering, and many other disciplines. Such studies require the simultaneous optimization of both quantities and the sequence orders of several components which are called quantitative-sequence (QS) factors. Given the large and semi-discrete solution spaces in such experiments, efficiently identifying optimal or near-optimal solutions by using a small number of experimental trials is a nontrivial task. To address this challenge, we propose a novel active learning approach, called QS-learning, to enable effective modeling and efficient optimization for experiments with QS factors. QS-learning consists of three parts: a novel mapping-based additive Gaussian process (MaGP) model, an efficient global optimization scheme (QS-EGO), and a new class of optimal designs (QS-design). The theoretical properties of the proposed method are investigated, and optimization techniques using analytical gradients are developed. The performance of the proposed method is demonstrated via a real drug experiment on lymphoma treatment and several simulation studies.

stat.ME↗

EzGP: Easy-to-Interpret Gaussian Process Models for Computer Experiments with Both Quantitative and Qualitative Factors

Computer experiments with both quantitative and qualitative (QQ) inputs are commonly used in science and engineering applications. Constructing desirable emulators for such computer experiments remains a challenging problem. In this article, we propose an easy-to-interpret Gaussian process (EzGP) model for computer experiments to reflect the change of the computer model under the different level combinations of qualitative factors. The proposed modeling strategy, based on an additive Gaussian process, is flexible to address the heterogeneity of computer models involving multiple qualitative factors. We also develop two useful variants of the EzGP model to achieve computational efficiency for data with high dimensionality and large sizes. The merits of these models are illustrated by several numerical examples and a real data application.

stat.ME↗

Statistical Analysis of Complex Computer Models in Astronomy

We introduce statistical techniques required to handle complex computer models with potential applications to astronomy. Computer experiments play a critical role in almost all fields of scientific research and engineering. These computer experiments, or simulators, are often computationally expensive, leading to the use of emulators for rapidly approximating the outcome of the experiment. Gaussian process models, also known as Kriging, are the most common choice of emulator. While emulators offer significant improvements in computation over computer simulators, they require a selection of inputs along with the corresponding outputs of the computer experiment to function well. Thus, it is important to select inputs judiciously for the full computer simulation to construct an accurate emulator. Space-filling designs are efficient when the general response surface of the outcome is unknown, and thus they are a popular choice when selecting simulator inputs for building an emulator. In this tutorial we discuss how to construct these space filling designs, perform the subsequent fitting of the Gaussian process surrogates, and briefly indicate their potential applications to astronomy research.

astro-ph.IM↗

Inverse Problem for Dynamic Computer Simulators via Multiple Scalar-valued Contour Estimation

In this paper we consider a dynamic computer simulator that produces a time-series response $y_t(x)$ over $L$ time points, for every given input parameter $x$. We propose a method for solving inverse problems, which refer to the finding of a set of inputs that generates a pre-specified simulator output. Inspired by the sequential approach of contour estimation via expected improvement criterion developed by Ranjan et al. (2008, DOI: 10.1198/004017008000000541), our proposed method discretizes the target response series on $k \; (\ll L)$ time points, and then iteratively solves $k$ scalar-valued inverse problems with respect to the discretized targets. We also propose to use spline smoothing of the target response series to identify the optimal number of knots, $k$, and the actual location of the knots for discretization. The performance of the proposed methods is compared for several test-function based computer simulators and the motivating real application that uses a rainfall-runoff measurement model named Matlab-Simulink model.

stat.ME↗

LowCon: A design-based subsampling approach in a misspecified linear modeL

We consider a measurement constrained supervised learning problem, that is, (1) full sample of the predictors are given; (2) the response observations are unavailable and expensive to measure. Thus, it is ideal to select a subsample of predictor observations, measure the corresponding responses, and then fit the supervised learning model on the subsample of the predictors and responses. However, model fitting is a trial and error process, and a postulated model for the data could be misspecified. Our empirical studies demonstrate that most of the existing subsampling methods have unsatisfactory performances when the models are misspecified. In this paper, we develop a novel subsampling method, called "LowCon", which outperforms the competing methods when the working linear model is misspecified. Our method uses orthogonal Latin hypercube designs to achieve a robust estimation. We show that the proposed design-based estimator approximately minimizes the so-called "worst-case" bias with respect to many possible misspecification terms. Both the simulated and real-data analyses demonstrate the proposed estimator is more robust than several subsample least squares estimators obtained by state-of-the-art subsampling methods.

stat.ME↗

Optimal Crossover Designs for Generalized Linear Models

We identify locally $D$-optimal crossover designs for generalized linear models. We use generalized estimating equations to estimate the model parameters along with their variances. To capture the dependency among the observations coming from the same subject, we propose six different correlation structures. We identify the optimal allocations of units for different sequences of treatments. For two-treatment crossover designs, we show via simulations that the optimal allocations are reasonably robust to different choices of the correlation structures. We discuss a real example of multiple treatment crossover experiments using Latin square designs. Using a simulation study, we show that a two-stage design with our locally $D$-optimal design at the second stage is more efficient than the uniform design, especially when the responses from the same subject are correlated.

stat.ME↗

A-ComVar: A Flexible Extension of Common Variance Design

We consider nonregular fractions of factorial experiments for a class of linear models. These models have a common general mean and main effects, however they may have different 2-factor interactions. Here we assume for simplicity that 3-factor and higher order interactions are negligible. In the absence of a priori knowledge about which interactions are important, it is reasonable to prefer a design that results in equal variance for the estimates of all interaction effects to aid in model discrimination. Such designs are called common variance designs and can be quite challenging to identify without performing an exhaustive search of possible designs. In this work, we introduce an extension of common variance designs called approximate common variance, or A-ComVar designs. We develop a numerical approach to finding A-ComVar designs that is much more efficient than an exhaustive search. We present the types of A-ComVar designs that can be found for different number of factors, runs, and interactions. We further demonstrate the competitive performance of both common variance and A-ComVar designs using several comparisons to other popular designs in the literature.

stat.CO↗

A Hierarchical Bayes Unit-Level Small Area Estimation Model for Normal Mixture Populations

National statistical agencies are regularly required to produce estimates about various subpopulations, formed by demographic and/or geographic classifications, based on a limited number of samples. Traditional direct estimates computed using only sampled data from individual subpopulations are usually unreliable due to small sample sizes. Subpopulations with small samples are termed small areas or small domains. To improve on the less reliable direct estimates, model-based estimates, which borrow information from suitable auxiliary variables, have been extensively proposed in the literature. However, standard model-based estimates rely on the normality assumptions of the error terms. In this research we propose a hierarchical Bayesian (HB) method for the unit-level nested error regression model based on a normal mixture for the unit-level error distribution. To implement our proposal we use a uniform prior for the regression parameters, random effects variance parameter, and the mixing proportion, and we use a partially proper non-informative prior distribution for the two unit-level error variance components in the mixture. We apply our method to two examples to predict summary characteristics of farm products at the small area level. One of the examples is prediction of twelve county-level crop areas cultivated for corn in some Iowa counties. The other example involves total cash associated in farm operations in twenty-seven farming regions in Australia. We compare predictions of small area characteristics based on the proposed method with those obtained by applying the Datta and Ghosh (1991) and the Chakraborty et al. (2018) HB methods. Our simulation study comparing these three Bayesian methods showed the superiority of our proposed method, measured by prediction mean squared error, coverage probabilities and lengths of credible intervals for the small area means.

stat.ME↗

A History Matching Approach for Calibrating Hydrological Models

Calibration of hydrological time-series models is a challenging task since these models give a wide spectrum of output series and calibration procedures require significant amount of time. From a statistical standpoint, this model parameter estimation problem simplifies to finding an inverse solution of a computer model that generates pre-specified time-series output (i.e., realistic output series). In this paper, we propose a modified history matching approach for calibrating the time-series rainfall-runoff models with respect to the real data collected from the state of Georgia, USA. We present the methodology and illustrate the application of the algorithm by carrying a simulation study and the two case studies. Several goodness-of-fit statistics were calculated to assess the model performance. The results showed that the proposed history matching algorithm led to a significant improvement, of 30% and 14% (in terms of root mean squared error) and 26% and 118% (in terms of peak percent threshold statistics), for the two case-studies with Matlab-Simulink and SWAT models, respectively.

stat.AP↗

Robust Hierarchical Bayes Small Area Estimation for Nested Error Regression Model

National statistical institutes in many countries are now mandated to produce reliable statistics for important variables such as population, income, unemployment, health outcomes, etc. for small areas, defined by geography and/or demography. Due to small samples from these areas, direct sample-based estimates are often unreliable. Model-based small area estimation is now extensively used to generate reliable statistics by "borrowing strength" from other areas and related variables through suitable models. Outliers adversely influence standard model-based small area estimates. To deal with outliers, Sinha and Rao (2009) proposed a robust frequentist approach. In this article, we present a robust Bayesian alternative to the nested error regression model for unit-level data to mitigate outliers. We consider a two-component scale mixture of normal distributions for the unit-level error to model outliers and present a computational approach to produce Bayesian predictors of small area means under a noninformative prior for model parameters. A real example and extensive simulations convincingly show robustness of our Bayesian predictors to outliers. Simulations comparison of these two procedures with Bayesian predictors by Datta and Ghosh (1991) and M-quantile estimators by Chambers et al. (2014) shows that our proposed procedure is better than the others in terms of bias, variability, and coverage probability of prediction intervals, when there are outliers. The superior frequentist performance of our procedure shows its dual (Bayes and frequentist) dominance, and makes it attractive to all practitioners, both Bayesian and frequentist, of small area estimation.

stat.ME↗

D-optimal Designs with Ordered Categorical Data

Cumulative link models have been widely used for ordered categorical responses. Uniform allocation of experimental units is commonly used in practice, but often suffers from a lack of efficiency. We consider D-optimal designs with ordered categorical responses and cumulative link models. For a predetermined set of design points, we derive the necessary and sufficient conditions for an allocation to be locally D-optimal and develop efficient algorithms for obtaining approximate and exact designs. We prove that the number of support points in a minimally supported design only depends on the number of predictors, which can be much less than the number of parameters in the model. We show that a D-optimal minimally supported allocation in this case is usually not uniform on its support points. In addition, we provide EW D-optimal designs as a highly efficient surrogate to Bayesian D-optimal designs. Both of them can be much more robust than uniform designs.

math.ST↗

Using particle swarm optimization to search for locally $D$-optimal designs for mixed factor experiments with binary response

Identifying optimal designs for generalized linear models with a binary response can be a challenging task, especially when there are both continuous and discrete independent factors in the model. Theoretical results rarely exist for such models, and the handful that do exist come with restrictive assumptions. This paper investigates the use of particle swarm optimization (PSO) to search for locally $D$-optimal designs for generalized linear models with discrete and continuous factors and a binary outcome and demonstrates that PSO can be an effective method. We provide two real applications using PSO to identify designs for experiments with mixed factors: one to redesign an odor removal study and the second to find an optimal design for an electrostatic discharge study. In both cases we show that the $D$-efficiencies of the designs found by PSO are much better than the implemented designs. In addition, we show PSO can efficiently find $D$-optimal designs on a prototype or an irregularly shaped design space, provide insights on the existence of minimally supported optimal designs, and evaluate sensitivity of the $D$-optimal design to mis-specifications in the link function.

stat.AP↗