SearcharxivSearch

arXiv subjects

Gonzalo Contador

Publications and source records attributed to Gonzalo Contador.

6 recordsLinked to original sources

Sharp bounds for stochastic proximal and projection estimators via radial dominance

We study stochastic barycentric estimators for proximal points and metric projections obtained by exponentially reweighting Gaussian perturbations. Our main result is an abstract comparison theorem for probability measures with densities proportional to an exponential weight, under a radial dominance condition relative to a prescribed profile. This yields an explicit upper bound for the norm of the associated barycenter in terms of a one-dimensional comparison measure. We also provide tractable sufficient conditions for radial dominance, including strong convexity, addition of nonnegative convex terms, and star-shaped constraints. As a consequence, we obtain a refined convergence rate for stochastic proximal estimators of weakly convex functions, together with asymptotic sharpness of the constant. The same framework yields a corresponding rate for stochastic projection estimators onto closed convex sets. We further establish basic structural properties of the barycentric approximation operator, such as smoothness and cocoercivity. Numerical experiments illustrate the predicted rate, the dimensional scaling of the constant, and its asymptotic sharpness.

math.OC

Differentiability and Approximation of Probability Functions under Gaussian Mixture Models

In this work, we study probability functions associated with Gaussian mixture models. Our primary focus is on extending the use of spherical radial decomposition for multivariate Gaussian random vectors to the context of Gaussian mixture models, which are not inherently spherical, but conditionally so. Specifically, the conditional probability distribution, given a random parameter of the random vector, follows a Gaussian distribution, which allows us to rewrite the probability function as a tractable integrated Gaussian mixture. This assumption, together with spherical radial decomposition for Gaussian random vectors, enables us to represent the probability function as an integral over the Euclidean sphere. Using this representation, we establish sufficient conditions to ensure the differentiability of the probability function and provide an integral representation of its gradient. Furthermore, we approximate the probability function using random sampling over the parameter space and the Euclidean sphere. Finally, we present a numerical example that illustrates the advantages of this approach over classical approximations based on random vector sampling.

math.OC

Optimal Adjustment and Combination of Independent Discrete $p$-Values

Combining p-values from multiple independent tests is a fundamental task in statistical inference, but presents unique challenges when the p-values are discrete. We extend a recent optimal transport-based framework for combining discrete p-values, which constructs a continuous surrogate distribution by minimizing the Wasserstein distance between the transformed discrete null and its continuous analogue. We provide a unified approach for several classical combination methods, including Fisher's, Pearson's, George's, Stouffer's, and Edgington's statistics. Our theoretical analysis and extensive simulations show that accurate Type I error control is achieved when the variance of the adjusted discrete statistic closely matches that of the continuous case. We further demonstrate that, when the likelihood ratio test is a monotonic function of a combination statistic, the proposed approximation achieves power comparable to the uniformly most powerful (UMP) test. The methodology is illustrated with a genetic association study of rare variants using case-control data, and is implemented in the R package DPComb.

stat.ME

A minimum Wasserstein distance approach to Fisher's combination of independent discrete p-values

This paper introduces a comprehensive framework to adjust a discrete test statistic for improving its hypothesis testing procedure. The adjustment minimizes the Wasserstein distance to a null-approximating continuous distribution, tackling some fundamental challenges inherent in combining statistical significances derived from discrete distributions. The related theory justifies Lancaster's mid-p and mean-value chi-squared statistics for Fisher's combination as special cases. However, in order to counter the conservative nature of Lancaster's testing procedures, we propose an updated null-approximating distribution. It is achieved by further minimizing the Wasserstein distance to the adjusted statistics within a proper distribution family. Specifically, in the context of Fisher's combination, we propose an optimal gamma distribution as a substitute for the traditionally used chi-squared distribution. This new approach yields an asymptotically consistent test that significantly improves type I error control and enhances statistical power.

math.ST

Another reason why normalized gain should continue to be used to analyze concept inventories (and estimate learning rates)

A transformation called normalized gain (ngain) has been acknowledged as one of the most common measures of knowledge growth in pretest-posttest contexts in physics education research. Recent studies in math education have shown that ngains can also be applied to assess learners' ability to acquire unfamiliar knowledge, that is, to estimate their "learning rate". This quantity is estimated from learning data through two well-known methods: computing the average ngain of the group or computing the ngain of the average learner. These two methods commonly yield different results, and prior research has concluded that the difference between them is associated with a pretest-ngains correlation. Such a correlation would suggest a bias of this learning measurement because it implies its favoring of certain subgroups of students according to their performance in pretest measurements. The present study analyzes these two estimation methods by drawing on statistical models. Our results show that the two estimation methods are equivalent when no measurement errors exist. In contrast, when there are measurement errors, the first method provides a biased estimator, whereas the second one provides an unbiased estimator. Furthermore, these measurement errors induce a spurious correlation between the pretest and ngain scores. Our results seem consistent with prior research, except they show that measurement errors in pretest and posttest scores are the source of a spurious pretest-ngain correlation. Consequently, estimating learning rates might effectively provide unbiased estimates of knowledge change that control for the effect of prior knowledge even in the presence of pretest-ngain correlations.

stat.AP

Sampling distributions and estimation for multi-type Branching Processes

Consider a multi-dimensional supercritical branching process with offspring distribution in a parametric family. Here, each vector coordinate corresponds to the number of offspring of a given type. The process is observed under family-size sampling: a random sample is drawn, each individual reporting its vector of brood sizes. In this work, we show that the set in which no siblings are sampled (so that the sample can be considered independent) has probability converging to one under certain conditions on the sampling size. Furthermore, we show that the sampling distribution of the observed sizes converges to the product of identical distributions, hence developing a framework for which the process can be considered iid, and the usual methods for parameter estimation apply. We provide asymptotic distributions for the resulting estimators.

math.ST