SearcharxivSearch

arXiv subjects

Aluísio Pinheiro

Publications and source records attributed to Aluísio Pinheiro.

9 recordsLinked to original sources

Some homogeneity test statistics for DNA evolutionary models

We present a test statistic for the comparison of DNA sequences under some of the most popular evolutionary processes available in the literature. Theoretical properties for the test statistic as well as its empirical performance by stochastic simulations are presented. The proposed test statistic is a generalized $U$-statistics built for tests of distributional homogeneity under the null hypothesis. We show that a dicothomous situation exists here. Under the null hypothesis, the $U$-statistics kernel is first-order degenerated, this test statistic falls in the quasi $U$-statistics class and follows an asymptotic normal law, albeit of higher order than the standard case. Under heterogeneity, the asymptotic normality is attained on the more usual first-order asymptotics. Asymptotic normality is proven for the cases: high-dimension/large sample size, high-dimension/small sample size, low-dimension/large sample size. Moreover, the case of local alternatives is discussed, and the contiguity of the test statistic for them is established. Simulation studies are performed to assess some finite-dimensional properties of the test statistic, regarding issues such as balanced/unbalanced samples, dimension and sample size.

stat.ME

Sparse Principal Component Analysis via Wavelets for Distributed Data

The large volume of data and concerns about data privacy have motivated the development of techniques for distributed data, a problem also known as federated learning. In this scenario, sub-samples of the data are divided across different machines, and statistics must be computed over that data without direct access to the full sample. Johnstone & Lu (2009, JASA) show that principal component analysis (PCA) is statistically inconsistent in the high-dimensional regime, and propose a way to recover consistency through wavelet-based sparsification and variable selection. Fan et al. (2019, AoS) show a way to perform this same estimation -- specifically, to estimate the eigenspace that would be obtained if all the data were pooled together, even though it remains effectively distributed -- without addressing the high-dimensional regime. This work incorporates the wavelet-based sparsification of Johnstone & Lu (2009) into the distributed PCA framework of Fan et al. (2019), aiming to reduce communication cost without compromising the quality of the eigenspace estimation. Simulations across $d \in [52, 5000]$ show that the proposed method overtakes Fan et al. (2019) in estimation error beyond a clear dimensional threshold ($d \geq 152$ for $λ=25$, $d \geq 252$ for $λ=50$), while transmitting systematically fewer coefficients throughout the entire range studied. This study was financed by the Sao Paulo Research Foundation (FAPESP), Brazil. Process Number #2023/02538-0 and Number #2025/21250-2.

stat.ME

A new wavelet-based variational family with copula dependence structures

Variational inference (VI) has become a widely used approach for scalable Bayesian inference, but its performance strongly depends on the flexibility of the chosen variational family. In this work, we propose a novel variational family that combines wavelet-based representations for marginal posterior densities with copula functions to model dependence structures. The marginal distributions are constructed using coefficients from the discrete wavelet transform, providing a flexible and adaptive framework capable of capturing complex features such as asymmetry. The joint distribution is then obtained through a copula, allowing for explicit modeling of dependence among parameters, including both independence and Gaussian copula structures. We develop an efficient estimation procedure based on Monte Carlo approximations of the evidence lower bound (ELBO) and automatic differentiation, enabling scalable optimization using gradient-based methods. Through extensive simulation studies, including logistic regression, sparse linear models, and hierarchical models, we demonstrate that the proposed approach achieves posterior mean estimates comparable to Markov chain Monte Carlo (MCMC) methods, while providing improved uncertainty quantification relative to standard variational approaches. Applications to hierarchical logistic regression and Bayesian conditional transformation models further illustrate the practical advantages of the method in complex, high dimensional settings. The proposed wavelet copula variational family offers a flexible and computationally efficient alternative for Bayesian inference.

stat.ME

Wavelet-based estimation of power densities of size-biased data

We propose a new wavelet-based method for density estimation when the data are size-biased. More specifically, we consider a power of the density of interest, where this power exceeds 1/2. Warped wavelet bases are employed, where warping is attained by some continuous cumulative distribution function. A special case is the conventional orthonormal wavelet estimation, where the warping distribution is the standard continuous uniform. We show that both linear and nonlinear wavelet estimators are consistent, with optimal and/or near-optimal rates. Monte Carlo simulations are performed to compare four special settings which are easy to interpret in practice. An application with a real dataset on fatal traffic accidents involving alcohol illustrates the method. We observe that warped bases provide more flexible and superior estimates for both simulated and real data. Moreover, we find that estimating the power of a density (for instance, its square root) further improves the results.

stat.ME

Wavelet Spatio-Temporal Change Detection on multi-temporal PolSAR images

We introduce WECS (Wavelet Energies Correlation Sreening), an unsupervised sparse procedure to detect spatio-temporal change points on multi-temporal SAR (POLSAR) images or even on sequences of very high resolution images. The procedure is based on wavelet approximation for the multi-temporal images, wavelet energy apportionment, and ultra-high dimensional correlation screening for the wavelet coefficients. We present two complimentary wavelet measures in order to detect sudden and/or cumulative changes, as well as for the case of stationary or non-stationary multi-temporal images. We show WECS performance on synthetic multi-temporal image data. We also apply the proposed method to a time series of 85 satellite images in the border region of Brazil and the French Guiana. The images were captured from November 08, 2015 to December 09 2017.

stat.AP

Nonparametric methods for detecting change in Multitemporal SAR/PolSAR Satellite Data

We employ nonparametric statistical procedures to analyse multitemporal SAR/PolSAR satellite images. The aim is two-fold. We seek parsimony in data representation as well as efficient change detection. For these, wavelets and geostatistical analyses are applied to the images (Morettin et al., 2017; Krainski et al., 2018). Following this representation, the dimension of the underlying generating process is estimated (Fonseca and Pinheiro, 2019), and a set of multivariate characteristics is extracted. Change-points are then detected via wavelets (Montoril et al., 2019).

stat.AP

A nonparametric approach to assess undergraduate performance

Nonparametric methodologies are proposed to assess college students' performance. Emphasis is given to gender and sector of High School. The application concerns the University of Campinas, a research university in Southeast Brazil. In Brazil college is based on a somewhat rigid set of subjects for each major. Thence a student's relative performance can not be accurately measured by the Grade Point Average or by any other single measure. We then define individual vectors of course grades. These vectors are used in pairwise comparisons of common subject grades for individuals that entered college in the same year. The relative college performances of any two students is compared to their relative performances on the Entrance Exam Score. A test based on generalized U-statistics is developed for homogeneity of some predefined groups. Asymptotic normality of the test statistic is true for both null and alternative hypotheses. Maximum power is attained by employing the union intersection principle.

stat.ME

Wavelet estimation of the dimensionality of curve time series

Functional data analysis is ubiquitous in most areas of sciences and engineering. Several paradigms are proposed to deal with the dimensionality problem which is inherent to this type of data. Sparseness, penalization, thresholding, among other principles, have been used to tackle this issue. We discuss here a solution based on a finite-dimensional functional space. We employ wavelet representation of the functionals to estimate this finite dimension, and successfully model a time series of curves. The proposed method is shown to have nice asymptotic properties. Moreover, the wavelet representation permits the use of several bootstrap procedures, and it results in faster computing algorithms. Besides the theoretical and computational properties, some simulation studies and an application to real data are provided.

stat.ME

An asymptotically normal test for the selective neutrality hypothesis

An important parameter in the study of population evolution is $θ=4Nν$, where $N$ is the effective population size and $ν$ is the rate of mutation per locus per generation. Therefore, $θ$ represents the mean number of mutations per site per generation. There are many estimators of $θ$, one of them being the mean number of pairwise nucleotide differences, which we call $\mathcal{T}_2$. Other estimators are $\mathcal{T}_1$, based on the number of segregating sites and $\mathcal{T}_3$, based on the number of singletons. The concept of selective neutrality can be interpreted as a differentiated nucleotide distribution for mutant sites when compared to the overall nucleotide distribution. Tajima (1989) has proposed the so-called Tajima's test of selective neutrality based on $\mathcal{T}_2-\mathcal{T}_1$. Its complex empirical behavior (Kiihl, 2005) motivates us to propose a test statistic solely based on $\mathcal{T}_2$. We are thus able to prove asymptotic normality under different assumptions on the number of sequences and number of sites via $U$-statistics theory.

math.ST