Searcharxiv⌕ Search

arXiv subjects

Tatjana Pavlenko

Publications and source records attributed to Tatjana Pavlenko.

5 recordsLinked to original sources

Test for mean matrix in GMANOVA model under heteroscedasticity and non-normality for high-dimensional data

This paper is concerned with the testing bilateral linear hypothesis on the mean matrix in the context of the generalized multivariate analysis of variance (GMANOVA) model when the dimensions of the observed vector may exceed the sample size, the design may become unbalanced, the population may not be normal, or the true covariance matrices may be unequal. The suggested testing methodology can treat many problems such as the one- and two-way MANOVA tests, the test for parallelism in profile analysis, etc., as specific ones. We propose a bias-corrected estimator of the Frobenius norm for the mean matrix, which is a key component of the test statistic. The null and non-null distributions are derived under a general high-dimensional asymptotic framework that allows the dimensionality to arbitrarily exceed the sample size of a group, thereby establishing consistency for the testing criterion. The accuracy of the proposed test in a finite sample is investigated through simulations conducted for several high-dimensional scenarios and various underlying population distributions in combination with different within-group covariance structures. Finally, the proposed test is applied to a high-dimensional two-way MANOVA problem for DNA microarray data.

stat.ME↗

Graphical posterior predictive classifier: Bayesian model averaging with particle Gibbs

In this study, we present a multi-class graphical Bayesian predictive classifier that incorporates the uncertainty in the model selection into the standard Bayesian formalism. For each class, the dependence structure underlying the observed features is represented by a set of decomposable Gaussian graphical models. Emphasis is then placed on the Bayesian model averaging which takes full account of the class-specific model uncertainty by averaging over the posterior graph model probabilities. An explicit evaluation of the model probabilities is well known to be infeasible. To address this issue, we consider the particle Gibbs strategy of Olsson et al. (2018b) for posterior sampling from decomposable graphical models which utilizes the Christmas tree algorithm of Olsson et al. (2018a) as proposal kernel. We also derive a strong hyper Markov law which we call the hyper normal Wishart law that allow to perform the resultant Bayesian calculations locally. The proposed predictive graphical classifier reveals superior performance compared to the ordinary Bayesian predictive rule that does not account for the model uncertainty, as well as to a number of out-of-the-box classifiers.

stat.ML↗

A $U$-classifier for high-dimensional data under non-normality

A classifier for two or more samples is proposed when the data are high-dimensional and the underlying distributions may be non-normal. The classifier is constructed as a linear combination of two easily computable and interpretable components, the $U$-component and the $P$-component. The $U$-component is a linear combination of $U$-statistics which are averages of bilinear forms of pairwise distinct vectors from two independent samples. The $P$-component is the discriminant score and is a function of the projection of the $U$-component on the observation to be classified. Combined, the two components constitute an inherently bias-adjusted classifier valid for high-dimensional data. The simplicity of the classifier helps conveniently study its properties, including its asymptotic normal limit, and extend it to multi-sample case. The classifier is linear but its linearity does not rest on the assumption of homoscedasticity. Probabilities of misclassification and asymptotic properties of their empirical versions are discussed in detail. Simulation results are used to show the accuracy of the proposed classifier for sample sizes as small as 5 or 7 and any large dimensions. Applications on real data sets are also demonstrated.

math.ST↗

Goodness-of-fit tests based on sup-functionals of weighted empirical processes

A large class of goodness-of-fit test statistics based on sup-functionals of weighted empirical processes is proposed and studied. The weight functions employed are Erdős-Feller-Kolmogorov-Petrovski upper-class functions of a Brownian bridge. Based on the result of M. Csörgő, S. Csörgő, Horváth, and Mason obtained for this type of test statistics, we provide the asymptotic null distribution theory for the class of tests in hand, and present an algorithm for tabulating the limit distribution functions under the null hypothesis. A new family of nonparametric confidence bands is constructed for the true distribution function and it is found to perform very well. The results obtained, together with a new result on the convergence in distribution of the higher criticism statistic, introduced by Donoho and Jin, demonstrate the advantage of our approach over a common approach that utilizes a family of regularly varying weight functions. Furthermore, we show that, in various subtle problems of detecting sparse heterogeneous mixtures, the proposed test statistics achieve the detection boundary found by Ingster and, when distinguishing between the null and alternative hypotheses, perform optimally adaptively to unknown sparsity and size of the non-null effects.

math.ST↗

Bayesian Network Classifiers in a High Dimensional Framework

We present a growing dimension asymptotic formalism. The perspective in this paper is classification theory and we show that it can accommodate probabilistic networks classifiers, including naive Bayes model and its augmented version. When represented as a Bayesian network these classifiers have an important advantage: The corresponding discriminant function turns out to be a specialized case of a generalized additive model, which makes it possible to get closed form expressions for the asymptotic misclassification probabilities used here as a measure of classification accuracy. Moreover, in this paper we propose a new quantity for assessing the discriminative power of a set of features which is then used to elaborate the augmented naive Bayes classifier. The result is a weighted form of the augmented naive Bayes that distributes weights among the sets of features according to their discriminative power. We derive the asymptotic distribution of the sample based discriminative power and show that it is seriously overestimated in a high dimensional case. We then apply this result to find the optimal, in a sense of minimum misclassification probability, type of weighting.

cs.LG↗