SearcharxivSearch

arXiv subjects

Mariela Sued

Publications and source records attributed to Mariela Sued.

At least 19 recordsLinked to original sources

$K-$means with learned metrics

We study the Fr\'echet $k-$means of a metric measure space when both the measure and the distance are unknown and have to be estimated. We prove a general result that states that the $k-$means are continuous with respect to the measured Gromov-Hausdorff topology. In this situation, we also prove a stability result for the Voronoi clusters they determine. We do not assume uniqueness of the set of $k-$means, but when it is unique, the results are stronger. This framework provides a unified approach to proving consistency for a wide range of metric learning procedures. As concrete applications, we obtain new consistency results for several important estimators that were previously unestablished, even when $k=1$. These include $k-$means based on: (i) Isomap and Fermat geodesic distances on manifolds, (ii) difussion distances, (iii) Wasserstein distances computed with respect to learned ground metrics. Finally, we consider applications beyond the statistical inference paradigm like (iv) first passage percolation and (v) discrete approximations of length spaces.

math.ST

Invariant Feature Extraction Through Conditional Independence and the Optimal Transport Barycenter Problem: the Gaussian case

A methodology is developed to extract $d$ invariant features $W=f(X)$ that predict a response variable $Y$ without being confounded by variables $Z$ that may influence both $X$ and $Y$. The methodology's main ingredient is the penalization of any statistical dependence between $W$ and $Z$ conditioned on $Y$, replaced by the more readily implementable plain independence between $W$ and the random variable $Z_Y = T(Z,Y)$ that solves the [Monge] Optimal Transport Barycenter Problem for $Z\mid Y$. In the Gaussian case considered in this article, the two statements are equivalent. When the true confounders $Z$ are unknown, other measurable contextual variables $S$ can be used as surrogates, a replacement that involves no relaxation in the Gaussian case if the covariance matrix $\Sigma_{ZS}$ has full range. The resulting linear feature extractor adopts a closed form in terms of the first $d$ eigenvectors of a known matrix. The procedure extends with little change to more general, non-Gaussian / non-linear cases.

math.ST

Threshold detection under a semiparametric regression model

Linear regression models have been extensively considered in the literature. However, in some practical applications they may not be appropriate all over the range of the covariate. In this paper, a more flexible model is introduced by considering a regression model $Y=r(X)+\varepsilon$ where the regression function $r(\cdot)$ is assumed to be linear for large values in the domain of the predictor variable $X$. More precisely, we assume that $r(x)=α_0+β_0 x$ for $x> u_0$, where the value $u_0$ is identified as the smallest value satisfying such a property. A penalized procedure is introduced to estimate the threshold $u_0$. The considered proposal focusses on a semiparametric approach since no parametric model is assumed for the regression function for values smaller than $u_0$. Consistency properties of both the threshold estimator and the estimators of $(α_0,β_0)$ are derived, under mild assumptions. Through a numerical study, the small sample properties of the proposed procedure and the importance of introducing a penalization are investigated. The analysis of a real data set allows us to demonstrate the usefulness of the penalized estimators.

math.ST

Asymptotic results for nonparametric regression estimators after sufficient dimension reduction estimation

Prediction, in regression and classification, is one of the main aims in modern data science. When the number of predictors is large, a common first step is to reduce the dimension of the data. Sufficient dimension reduction (SDR) is a well established paradigm of reduction that keeps all the relevant information in the covariates X that is necessary for the prediction of Y . In practice, SDR has been successfully used as an exploratory tool for modelling after estimation of the sufficient reduction. Nevertheless, even if the estimated reduction is a consistent estimator of the population, there is no theory that supports this step when non-parametric regression is used in the imputed estimator. In this paper, we show that the asymptotic distribution of the non-parametric regression estimator is the same regardless if the true SDR or its estimator is used. This result allows making inferences, for example, computing confidence intervals for the regression function avoiding the curse of dimensionality.

stat.ME

Each student with her/his own data: understanding sampling distributions

Sampling distribution, a foundational concept in statistics, is difficult to understand, since we usually have only one realization of the estimator of interest. In this work, we present an innovative method for helping university students understand the variability of an estimator. In our approach, each student uses a different data set, getting diverse estimations. Then, sharing the results, we can empirically study the sampling distribution. After some "handmade" experiences, we have built a web page to deliver a personalized data set for each student. Through this web page, we can also reformulate the role of the student in the classroom, inviting him/her to became an active player, submitting the solution to different problems and checking whether they are correct. In this work we present a recent experience in such direction.

stat.OT

Double-robust and efficient methods for estimating the causal effects of a binary treatment

We consider the problem of estimating the effects of a binary treatment on a continuous outcome of interest from observational data in the absence of confounding by unmeasured factors. We provide a new estimator of the population average treatment effect (ATE) based on the difference of novel double-robust (DR) estimators of the treatment-specific outcome means. We compare our new estimator with previously estimators both theoretically and via simulation. DR-difference estimators may have poor finite sample behavior when the estimated propensity scores in the treated and untreated do not overlap. We therefore propose an alternative approach, which can be used even in this unfavorable setting, based on locally efficient double-robust estimation of a semiparametric regression model for the modification on an additive scale of the magnitude of the treatment effect by the baseline covariates $X$. In contrast with existing methods, our approach simultaneously provides estimates of: i) the average treatment effect in the total study population, ii) the average treatment effect in the random subset of the population with overlapping estimated propensity scores, and iii) the treatment effect at each level of the baseline covariates $X$. When the covariate vector $X$ is high dimensional, one cannot be certain, owing to lack of power, that given models for the propensity score and for the regression of the outcome on treatment and $X$ used in constructing our DR estimators are nearly correct, even if they pass standard goodness of fit tests. Therefore to select among candidate models, we propose a novel approach to model selection that leverages the DR-nature of our treatment effect estimator and that outperforms cross-validation in a small simulation study.

stat.ME

On semi-supervised learning

Semi-supervised learning deals with the problem of how, if possible, to take advantage of a huge amount of unclassified data, to perform a classification in situations when, typically, there is little labeled data. Even though this is not always possible (it depends on how useful, for inferring the labels, it would be to know the distribution of the unlabeled data), several algorithm have been proposed recently. %but in general they are not proved to outperform A new algorithm is proposed, that under almost necessary conditions, %and it is proved that it attains asymptotically the performance of the best theoretical rule as the amount of unlabeled data tends to infinity. The set of necessary assumptions, although reasonable, show that semi-supervised classification only works for very well conditioned problems. The focus is on understanding when and why semi-supervised learning works when the size of the initial training sample remains fixed and the asymptotic is on the size of the unlabeled data. The performance of the algorithm is assessed in the well known "Isolet" real-data of phonemes, where a strong dependence on the choice of the initial training sample is shown.

stat.ML

The spatial sign covariance operator: Asymptotic results and applications

Due to the increasing recording capability, functional data analysis has become an important research topic. For functional data the study of outlier detection and/or the development of robust statistical procedures has started recently. One robust alternative to the sample covariance operator is the sample spatial sign covariance operator. In this paper, we study the asymptotic behaviour of the sample spatial sign covariance operator when location is unknown. Among other possible applications of the obtained results, we derive the asymptotic distribution of the principal directions obtained from the sample spatial sign covariance operator and we develop test to detect differences between the scatter operators of two populations. In particular, the test performance is illustrated through a Monte Carlo study for small sample sizes.

math.ST

Semi-supervised learning

Semi-supervised learning deals with the problem of how, if possible, to take advantage of a huge amount of not classified data, to perform classification, in situations when, typically, the labelled data are few. Even though this is not always possible (it depends on how useful is to know the distribution of the unlabelled data in the inference of the labels), several algorithm have been proposed recently. A new algorithm is proposed, that under almost neccesary conditions, attains asymptotically the performance of the best theoretical rule, when the size of unlabeled data tends to infinity. The set of necessary assumptions, although reasonables, show that semi-parametric classification only works for very well conditioned problems.

math.ST

Robust Doubly Protected Estimators for Quantiles with Missing Data

Doubly protected estimators are widely used for estimating the population mean of an outcome Y from a sample where the response is missing in some individuals. To compensate for the missing responses, a vector X of covariates is observed at each individual, and the missing mechanism is assumed to be independent of the response, conditioned on X (missing at random). In recent years, many authors have moved from the mean to the median, and more generally, doubly protected estimators of the quantiles have been proposed, assuming a parametric regression model for the relationship between X and Y and a parametric form for the propensity score. In this work, we present doubly protected estimators for the quantiles that are also robust, in the sense that they are resistant to the presence of outliers in the sample. We also flexibilize the model for the relationship between X and Y . Thus we present robust doubly protected estimators for the quantiles of the response in the presence of missing observations, postulating a semiparametric regression model for the relationship between the response and the covariates and a parametric model for the propensity score.

stat.ME

Asymptotic theory for maximum likelihood estimates in reduced-rank multivariate generalised linear models

Reduced-rank regression is a dimensionality reduction method with many applications. The asymptotic theory for reduced rank estimators of parameter matrices in multivariate linear models has been studied extensively. In contrast, few theoretical results are available for reduced-rank multivariate generalised linear models. We develop M-estimation theory for concave criterion functions that are maximised over parameters spaces that are neither convex nor closed. These results are used to derive the consistency and asymptotic distribution of maximum likelihood estimators in reduced-rank multivariate generalised linear models, when the response and predictor vectors have a joint distribution. We illustrate our results in a real data classification problem with binary covariates.

math.ST

Models for the Propensity Score that Contemplate the Positivity Assumption and their Application to Missing Data and Causality

Generalized linear models are often assumed to fit propensity scores, which are used to compute inverse probability weighted (IPW) estimators. In order to derive the asymptotic properties of IPW estimators, the propensity score is supposed to be bounded away from cero. This condition is known in the literature as strict positivity (or positivity assumption) and, in practice, when it does not hold, IPW estimators are very unstable and have a large variability. Although strict positivity is often assumed, it is not upheld when some of the covariates are continuous. In this work, we attempt to conciliate between the strict positivity condition and the theory of generalized linear models by incorporating an extra parameter, which results in an explicit lower bound for the propensity scores.

stat.ME

Testing equality between several populations covariance operators

In many situations, when dealing with several populations, equality of the covariance operators is assumed. An important issue is to study if this assumption holds before making other inferences. In this paper, we develop a test for comparing covariance operators of several functional data samples. The proposed test is based on the Hilbert--Schmidt norm of the difference between estimated covariance operators. In particular, when dealing with two populations, the tests statistic is just the squared norm of the difference between the two covariance operators estimators. The asymptotic behaviour of the test statistic under the null and under local alternatives is obtained. Since the statistic null asymptotic distribution does not allow to obtain easily its quantiles, a bootstrap procedure to compute the critical values is considered. The performance of the test statistics for small sample sizes is illustrated through a Monte Carlo study.

math.ST

Some considerations on the back door theorem and conditional randomization

In this work we propose a different surgical modified model for the construction of counterfactual variables under non parametric structural equation models. This approach allows the simultaneous representation of counterfactual responses and observed treatment assignment, at least when the intervention is done in one node. Using the new proposal, the d-separation criterion is used verify conditions related with ignorability or conditional ignorability and a new proof of the back door theorem is provided under this framework.

math.ST

Continuity and differentiability of regression M functionals

This paper deals with the Fisher-consistency, weak continuity and differentiability of estimating functionals corresponding to a class of both linear and nonlinear regression high breakdown M estimates, which includes S and MM estimates. A restricted type of differentiability, called weak differentiability, is defined, which suffices to prove the asymptotic normality of estimates based on the functionals. This approach allows to prove the consistency, asymptotic normality and qualitative robustness of M estimates under more general conditions than those required in standard approaches. In particular, we prove that regression MM-estimates are asymptotically normal when the observations are $ϕ$-mixing.

math.ST

Robust location estimation with missing data

In a missing-data setting, we have a sample in which a vector of explanatory variables x_i is observed for every subject i, while scalar outcomes y_i are missing by happenstance on some individuals. In this work we propose robust estimates of the distribution of the responses assuming missing at random (MAR) data, under a semiparametric regression model. Our approach allows the consistent estimation of any weakly continuous functional of the response's distribution. In particular, strongly consistent estimates of any continuous location functional, such as the median or MM functionals, are proposed. A robust fit for the regression model combined with the robust properties of the location functional gives rise to a robust recipe for estimating the location parameter. Robustness is quantified through the breakdown point of the proposed procedure. The asymptotic distribution of the location estimates is also derived.

math.ST

First-principles molecular dynamics simulations at solid-liquid interfaces with a continuum solvent

Continuum solvent models have become a standard technique in the context of electronic structure calculations, yet, no implementations have been reported capable to perform molecular dynamics at solid-liquid interfaces. We propose here such a continuum approach in a DFT framework, using plane-waves basis sets and periodic boundary conditions. Our work stems from a recent model designed for Car-Parrinello simulations of quantum solutes in a dielectric medium [J. Chem. Phys. 124, 74103 (2006)], for which the permittivity of the solvent is defined as a function of the electronic density of the solute. This strategy turns out to be inadequate for systems extended in two dimensions, by introducing new term in the Kohn-Sham potential which becomes unphysically large at the interfacial region, seriously affecting the convergence. If the dielectric medium is properly redefined as a function of the atomic coordinates, a good convergence is obtained and the constant of motion is conserved during the molecular dynamics simulations. Moreover, a significant gain in efficiency can be achieved if the simulation box is partitioned in two, solving the Poisson problem separately for the "dry" region using fast Fourier transforms, and for the solvated or "wet" region using a multigrid method. Eventually both solutions are combined in a self-consistent procedure, and in this way Car-Parrinello molecular dynamics simulations of solid-liquid interfaces can be performed at a very moderate computational cost. This scheme is employed to investigate the acid-base equilibrium at the TiO2-water interface.

physics.chem-ph