SearcharxivSearch

arXiv subjects

Hildete P. Pinheiro

Publications and source records attributed to Hildete P. Pinheiro.

4 recordsLinked to original sources

Some homogeneity test statistics for DNA evolutionary models

We present a test statistic for the comparison of DNA sequences under some of the most popular evolutionary processes available in the literature. Theoretical properties for the test statistic as well as its empirical performance by stochastic simulations are presented. The proposed test statistic is a generalized $U$-statistics built for tests of distributional homogeneity under the null hypothesis. We show that a dicothomous situation exists here. Under the null hypothesis, the $U$-statistics kernel is first-order degenerated, this test statistic falls in the quasi $U$-statistics class and follows an asymptotic normal law, albeit of higher order than the standard case. Under heterogeneity, the asymptotic normality is attained on the more usual first-order asymptotics. Asymptotic normality is proven for the cases: high-dimension/large sample size, high-dimension/small sample size, low-dimension/large sample size. Moreover, the case of local alternatives is discussed, and the contiguity of the test statistic for them is established. Simulation studies are performed to assess some finite-dimensional properties of the test statistic, regarding issues such as balanced/unbalanced samples, dimension and sample size.

stat.ME

A Multivariate Methodology for Analysing Students' Performance Using Register Data

We present a new method for jointly modelling the students' results in the university's admission exams and their performance in subsequent courses at the university. The case considered involved all the students enrolled at the University of Campinas in 2014 to evening studies programs in educational branches related to exact sciences. We collected the number of attempts used for passing the university course of geometry and the results of the admission exams of those students in seven disciplines. The method introduced involved a combination of multivariate generalised linear mixed models (GLMM) and graphical models for representing the covariance structure of the random components. The models we used allowed us to discuss the association of quantities of very different nature. We used Gaussian GLMM for modelling the performance in the admission exams and a frailty discrete-time Cox proportional model, represented by a GLMM, to describe the number of attempts for passing Geometry. The analyses were stratified into two populations: the students who received a bonus giving advantages in the university's admission process to compensate social and racial inequalities and those who did not receive the compensation. The two populations presented different patterns. Using general properties of graphical models, we argue that, on the one hand, the predicted performance in the admission exam of Mathematics could solely be used as a predictor of the performance in geometry for the students who received the bonus. On the other hand, the Portuguese admission exam's predicted performance could be used as a single predictor of the performance in geometry for the students who did not receive the bonus.

stat.AP

A nonparametric approach to assess undergraduate performance

Nonparametric methodologies are proposed to assess college students' performance. Emphasis is given to gender and sector of High School. The application concerns the University of Campinas, a research university in Southeast Brazil. In Brazil college is based on a somewhat rigid set of subjects for each major. Thence a student's relative performance can not be accurately measured by the Grade Point Average or by any other single measure. We then define individual vectors of course grades. These vectors are used in pairwise comparisons of common subject grades for individuals that entered college in the same year. The relative college performances of any two students is compared to their relative performances on the Entrance Exam Score. A test based on generalized U-statistics is developed for homogeneity of some predefined groups. Asymptotic normality of the test statistic is true for both null and alternative hypotheses. Maximum power is attained by employing the union intersection principle.

stat.ME

An asymptotically normal test for the selective neutrality hypothesis

An important parameter in the study of population evolution is $θ=4Nν$, where $N$ is the effective population size and $ν$ is the rate of mutation per locus per generation. Therefore, $θ$ represents the mean number of mutations per site per generation. There are many estimators of $θ$, one of them being the mean number of pairwise nucleotide differences, which we call $\mathcal{T}_2$. Other estimators are $\mathcal{T}_1$, based on the number of segregating sites and $\mathcal{T}_3$, based on the number of singletons. The concept of selective neutrality can be interpreted as a differentiated nucleotide distribution for mutant sites when compared to the overall nucleotide distribution. Tajima (1989) has proposed the so-called Tajima's test of selective neutrality based on $\mathcal{T}_2-\mathcal{T}_1$. Its complex empirical behavior (Kiihl, 2005) motivates us to propose a test statistic solely based on $\mathcal{T}_2$. We are thus able to prove asymptotic normality under different assumptions on the number of sequences and number of sites via $U$-statistics theory.

math.ST