SearcharxivSearch

arXiv subjects

Shifeng Xiong

Publications and source records attributed to Shifeng Xiong.

15 recordsLinked to original sources

Discretization approximation: An alternative to Monte Carlo in Bayesian computation

In this paper we propose a new deterministic approximation method, called discretization approximation, for Bayesian computation. Discretization approximation is very simple to understand and to implement, It only requires calculating posterior density values as probability masses at pre-specified support points. The resulted discrete distribution can be a good approximation to the target posterior distribution. All posterior quantities, including means, standard deviations, and quantiles, can be approximated by those of this completely known discrete distribution. We establish the convergence rate of discretization approximation as the number of support points goes to infinity. If the support points are generated from quasi-Monte Carlo sequences, then the rate is actually the same as that in integration approximation, generally faster than the optimal statistical rate. In this sense, discretization approximation is superior to the popular Markov chain Monte Carlo method. We also provide random sampling and representation point construction methods from discretization approximation. Numerical examples including some benchmarks demonstrate that the proposed method performs quite well for both low-dimensional and high-dimensional cases.

stat.CO

Nonparametric inference with massive data via grouped empirical likelihood

To address the computational issue in empirical likelihood methods with massive data, this paper proposes a grouped empirical likelihood (GEL) method. It divides $N$ observations into $n$ groups, and assigns the same probability weight to all observations within the same group. GEL estimates the $n\ (\ll N)$ weights by maximizing the empirical likelihood ratio. The dimensionality of the optimization problem is thus reduced from $N$ to $n$, thereby lowering the computational complexity. We prove that GEL possesses the same first order asymptotic properties as the conventional empirical likelihood method under the estimating equation settings and the classical two-sample mean problem. A distributed GEL method is also proposed with several servers. Numerical simulations and real data analysis demonstrate that GEL can keep the same inferential accuracy as the conventional empirical likelihood method, and achieves substantial computational acceleration compared to the divide-and-conquer empirical likelihood method. We can analyze a billion data with GEL in tens of seconds on only one PC.

stat.ME

A unified framework of principal component analysis and factor analysis

Principal component analysis and factor analysis are fundamental multivariate analysis methods. In this paper a unified framework to connect them is introduced. Under a general latent variable model, we present matrix optimization problems from the viewpoint of loss function minimization, and show that the two methods can be viewed as solutions to the optimization problems with specific loss functions. Specifically, principal component analysis can be derived from a broad class of loss functions including the L2 norm, while factor analysis corresponds to a modified L0 norm problem. Related problems are discussed, including algorithms, penalized maximum likelihood estimation under the latent variable model, and a principal component factor model. These results can lead to new tools of data analysis and research topics.

stat.ME

Physical Parameter Calibration

Computer simulation models are widely used to study complex physical systems. A related fundamental topic is the inverse problem, also called calibration, which aims at learning about the values of parameters in the model based on observations. In most real applications, the parameters have specific physical meanings, and we call them physical parameters. To recognize the true underlying physical system, we need to effectively estimate such parameters. However, existing calibration methods cannot do this well due to the model identifiability problem. This paper proposes a semi-parametric model, called the discrepancy decomposition model, to describe the discrepancy between the physical system and the computer model. The proposed model possesses a clear interpretation, and more importantly, it is identifiable under mild conditions. Under this model, we present estimators of the physical parameters and the discrepancy, and then establish their asymptotic properties. Numerical examples show that the proposed method can better estimate the physical parameters than existing methods.

stat.ME

Unweighted estimation based on optimal sample under measurement constraints

To tackle massive data, subsampling is a practical approach to select the more informative data points. However, when responses are expensive to measure, developing efficient subsampling schemes is challenging, and an optimal sampling approach under measurement constraints was developed to meet this challenge. This method uses the inverses of optimal sampling probabilities to reweight the objective function, which assigns smaller weights to the more important data points. Thus the estimation efficiency of the resulting estimator can be improved. In this paper, we propose an unweighted estimating procedure based on optimal subsamples to obtain a more efficient estimator. We obtain the unconditional asymptotic distribution of the estimator via martingale techniques without conditioning on the pilot estimate, which has been less investigated in the existing subsampling literature. Both asymptotic results and numerical results show that the unweighted estimator is more efficient in parameter estimation.

stat.CO

Design and analysis of computer experiments with both numeral and distribution inputs

Nowadays stochastic computer simulations with both numeral and distribution inputs are widely used to mimic complex systems which contain a great deal of uncertainty. This paper studies the design and analysis issues of such computer experiments. First, we provide preliminary results concerning the Wasserstein distance in probability measure spaces. To handle the product space of the Euclidean space and the probability measure space, we prove that, through the mapping from a point in the Euclidean space to the mass probability measure at this point, the Euclidean space can be isomorphic to the subset of the probability measure space, which consists of all the mass measures, with respect to the Wasserstein distance. Therefore, the product space can be viewed as a product probability measure space. We derive formulas of the Wasserstein distance between two components of this product probability measure space. Second, we use the above results to construct Wasserstein distance-based space-filling criteria in the product space of the Euclidean space and the probability measure space. A class of optimal Latin hypercube-type designs in this product space are proposed. Third, we present a Wasserstein distance-based Gaussian process model to analyze data from computer experiments with both numeral and distribution inputs. Numerical examples and real applications to a metro simulation are presented to show the effectiveness of our methods.

stat.ME

Linear screening for high-dimensional computer experiments

In this paper we propose a linear variable screening method for computer experiments when the number of input variables is larger than the number of runs. This method uses a linear model to model the nonlinear data, and screens the important variables by existing screening methods for linear models. When the underlying simulator is nearly sparse, we prove that the linear screening method is asymptotically valid under mild conditions. To improve the screening accuracy, we also provide a two-stage procedure that uses different basis functions in the linear model. The proposed methods are very simple and easy to implement. Numerical results indicate that our methods outperform existing model-free screening methods.

stat.ME

The Reconstruction Approach: From Interpolation to Regression

This paper introduces an interpolation-based method, called the reconstruction approach, for nonparametric regression. Based on the fact that interpolation usually has negligible errors compared to statistical estimation, the reconstruction approach uses an interpolator to parameterize the regression function with its values at finite knots, and then estimates these values by (regularized) least squares. Some popular methods including kernel ridge regression can be viewed as its special cases. It is shown that, the reconstruction idea not only provides different angles to look into existing methods, but also produces new effective experimental design and estimation methods for nonparametric models. In particular, for some methods of complexity O(n3), where n is the sample size, this approach provides effective surrogates with much less computational burden. This point makes it very suitable for large datasets.

stat.ML

Better Solution Principle: A Facet of Concordance between Optimization and Statistics

Many statistical methods require solutions to optimization problems. When the global solution is hard to attain, statisticians always use the better if there are two solutions for chosen, where the word "better" is understood in the sense of optimization. This seems reasonable in that the better solution is more likely to be the global solution, whose statistical properties of interest usually have been well established. From the statistical perspective, we use the better solution because we intuitively believe the principle, called better solution principle (BSP) in this paper, that a better solution to a statistical optimization problem also has better statistical properties of interest. BSP displays some concordance between optimization and statistics, and is expected to widely hold. Since theoretical study on BSP seems to be neglected by statisticians, this paper aims to establish a framework for discussing BSP in various statistical optimization problems. We demonstrate several simple but effective comparison theorems as the key results of this paper, and apply them to verify BSP in commonly encountered statistical optimization problems, including maximum likelihood estimation, best subsample selection, and best subset regression. It can be seen that BSP for these problems holds under reasonable conditions, i.e., a better solution indeed has better statistical properties of interest. In addition, guided by the BSP theory, we develop a new best subsample selection method that performs well when there are clustered outliers.

math.ST

Personalized Optimization for Computer Experiments with Environmental Inputs

Optimization problems with both control variables and environmental variables arise in many fields. This paper introduces a framework of personalized optimization to han- dle such problems. Unlike traditional robust optimization, personalized optimization devotes to finding a series of optimal control variables for different values of environmental variables. Therefore, the solution from personalized optimization consists of optimal surfaces defined on the domain of the environmental variables. When the environmental variables can be observed or measured, personalized optimization yields more reasonable and better solution- s than robust optimization. The implementation of personalized optimization for complex computer models is discussed. Based on statistical modeling of computer experiments, we provide two algorithms to sequentially design input values for approximating the optimal surfaces. Numerical examples show the effectiveness of our algorithms.

stat.CO

Local optimization-based statistical inference

This paper introduces a local optimization-based approach to test statistical hypotheses and to construct confidence intervals. This approach can be viewed as an extension of bootstrap, and yields asymptotically valid tests and confidence intervals as long as there exist consistent estimators of unknown parameters. We present simple algorithms including a neighborhood bootstrap method to implement the approach. Several examples in which theoretical analysis is not easy are presented to show the effectiveness of the proposed approach.

stat.ME

Comparisons of penalized least squares methods by simulations

Penalized least squares methods are commonly used for simultaneous estimation and variable selection in high-dimensional linear models. In this paper we compare several prevailing methods including the lasso, nonnegative garrote, and SCAD in this area through Monte Carlo simulations. Criterion for evaluating these methods in terms of variable selection and estimation are presented. This paper focuses on the traditional n > p cases. For larger p, our results are still helpful to practitioners after the dimensionality is reduced by a screening method. K

stat.CO

OEM for least squares problems

We propose an algorithm, called OEM (a.k.a. orthogonalizing EM), intended for var- ious least squares problems. The first step, named active orthogonization, orthogonalizes an arbi- trary regression matrix by elaborately adding more rows. The second step imputes the responses of the new rows. The third step solves the least squares problem of interest for the complete orthog- onal design. The second and third steps have simple closed forms, and iterate until convergence. The algorithm works for ordinary least squares and regularized least squares with the lasso, SCAD, MCP and other penalties. It has several attractive theoretical properties. For the ordinary least squares with a singular regression matrix, an OEM sequence converges to the Moore-Penrose gen- eralized inverse-based least squares estimator. For the SCAD and MCP, an OEM sequence can achieve the oracle property after sufficient iterations for a fixed or diverging number of variables. For ordinary and regularized least squares with various penalties, an OEM sequence converges to a point having grouping coherence for fully aliased regression matrices. Convergence and convergence rate of the algorithm are examined. These convergence rate results show that for the same data set, OEM converges faster for regularized least squares than ordinary least squares. This provides a new theoretical comparison between these methods. Numerical examples are provided to illustrate the proposed algorithm.

stat.CO

Better subset regression

To find efficient screening methods for high dimensional linear regression models, this paper studies the relationship between model fitting and screening performance. Under a sparsity assumption, we show that a subset that includes the true submodel always yields smaller residual sum of squares (i.e., has better model fitting) than all that do not in a general asymptotic setting. This indicates that, for screening important variables, we could follow a "better fitting, better screening" rule, i.e., pick a "better" subset that has better model fitting. To seek such a better subset, we consider the optimization problem associated with best subset regression. An EM algorithm, called orthogonalizing subset screening, and its accelerating version are proposed for searching for the best subset. Although the two algorithms cannot guarantee that a subset they yield is the best, their monotonicity property makes the subset have better model fitting than initial subsets generated by popular screening methods, and thus the subset can have better screening performance asymptotically. Simulation results show that our methods are very competitive in high dimensional variable screening even for finite sample sizes.

stat.ME

On best subset regression

In this paper we discuss the variable selection method from \ell0-norm constrained regression, which is equivalent to the problem of finding the best subset of a fixed size. Our study focuses on two aspects, consistency and computation. We prove that the sparse estimator from such a method can retain all of the important variables asymptotically for even exponentially growing dimensionality under regularity conditions. This indicates that the best subset regression method can efficiently shrink the full model down to a submodel of a size less than the sample size, which can be analyzed by well-developed regression techniques for such cases in a follow-up study. We provide an iterative algorithm, called orthogonalizing subset selection (OSS), to address computational issues in best subset regression. OSS is an EM algorithm, and thus possesses the monotonicity property. For any sparse estimator, OSS can improve its fit of the model by putting it as an initial point. After this improvement, the sparsity of the estimator is kept. Another appealing feature of OSS is that, similarly to an effective algorithm for a continuous optimization problem, OSS can converge to the global solution to the \ell0-norm constrained regression problem if the initial point lies in a neighborhood of the global solution. An accelerating algorithm of OSS and its combination with forward stepwise selection are also investigated. Simulations and a real example are presented to evaluate the performances of the proposed methods.

stat.ME