SearcharxivSearch

arXiv subjects

Heping He

Publications and source records attributed to Heping He.

3 recordsLinked to original sources

A Random Sample Partition Data Model for Big Data Analysis

Big data sets must be carefully partitioned into statistically similar data subsets that can be used as representative samples for big data analysis tasks. In this paper, we propose the random sample partition (RSP) data model to represent a big data set as a set of non-overlapping data subsets, called RSP data blocks, where each RSP data block has a probability distribution similar to the whole big data set. Under this data model, efficient block level sampling is used to randomly select RSP data blocks, replacing expensive record level sampling to select sample data from a big distributed data set on a computing cluster. We show how RSP data blocks can be employed to estimate statistics of a big data set and build models which are equivalent to those built from the whole big data set. In this approach, analysis of a big data set becomes analysis of few RSP data blocks which have been generated in advance on the computing cluster. Therefore, the new method for data analysis based on RSP data blocks is scalable to big data.

cs.DC

Asymptotic properties of maximum likelihood estimators in models with multiple change points

Models with multiple change points are used in many fields; however, the theoretical properties of maximum likelihood estimators of such models have received relatively little attention. The goal of this paper is to establish the asymptotic properties of maximum likelihood estimators of the parameters of a multiple change-point model for a general class of models in which the form of the distribution can change from segment to segment and in which, possibly, there are parameters that are common to all segments. Consistency of the maximum likelihood estimators of the change points is established and the rate of convergence is determined; the asymptotic distribution of the maximum likelihood estimators of the parameters of the within-segment distributions is also derived. Since the approach used in single change-point models is not easily extended to multiple change-point models, these results require the introduction of those tools for analyzing the likelihood function in a multiple change-point model.

math.ST

Higher-order asymptotic normality of approximations to the modified signed likelihood ratio statistic for regular models

Approximations to the modified signed likelihood ratio statistic are asymptotically standard normal with error of order $n^{-1}$, where $n$ is the sample size. Proofs of this fact generally require that the sufficient statistic of the model be written as $(\hatθ,a)$, where $\hatθ$ is the maximum likelihood estimator of the parameter $θ$ of the model and $a$ is an ancillary statistic. This condition is very difficult or impossible to verify for many models. However, calculation of the statistics themselves does not require this condition. The goal of this paper is to provide conditions under which these statistics are asymptotically normally distributed to order $n^{-1}$ without making any assumption about the sufficient statistic of the model.

math.ST