SearcharxivSearch

arXiv subjects

Fadoua Balabdaoui

Publications and source records attributed to Fadoua Balabdaoui.

At least 19 recordsLinked to original sources

Semi-supervised learning in unmatched linear regression using an empirical likelihood approach

Knowing the link between observed predictive variables and outcomes is crucial for making inference in any regression model. When this link is missing, partially or completely, classical estimation methods fail in recovering the true regression function. Deconvolution approaches have been proposed and studied in detail in the unmatched setting where the predictive variables and responses are allowed to be independent. In this work, we consider linear regression in a semi-supervised learning setting where, beside a small sample of matched data, we have access to a relatively large unmatched sample. Using maximum likelihood estimation, we show that under some mild assumptions the semi-supervised learning empirical maximum likelihood estimator (SSLEMLE) is asymptotically normal and give explicitly its asymptotic covariance matrix as a function of the ratio of the matched/unmatched sample sizes and other parameters. Furthermore, we quantify the statistical gain achieved by having the additional large unmatched sample over having only the small matched sample. To illustrate the theory, we present the results of an extensive simulation study and apply our methodology to the "combined cycle power plant" data set.

math.ST

Deconvolution in unlinked linear models

Unlinked regression, in which covariates and responses are observed separately without known correspondence, has recently gained increasing attention. Deconvolution, on the other hand, is a fundamental and challenging problem in nonparametric statistics with the aim of estimating the distribution of a latent random variable $Z$ based on observations contaminated by some additive noise. The complexity of this task is heavily influenced by the smoothness of the noise distribution and often leads to slow estimation rates. In this paper, we combine the recent unlinked linear regression problem with the classical deconvolution framework. Specifically, we study nonparametric deconvolution under the assumption that $Z$ is a linear function of an observable multidimensional covariate. This structural constraint allows us to introduce a nonparametric estimator of the distribution of $Z$ which achieves the parametric rate of convergence in the Wasserstein distance of order 1, where the smoothness of the noise does not affect the rate. Furthermore, we introduce nonparametric estimators for the unconditional density of $Z$ and the conditional density of $Z$ given an observed response. This allows us to study the problem of estimating the value of the latent linear predictor, whose link to the observed response is not accessible. Through several simulations, we illustrate the fast convergence rate of our deconvolution estimator and the performance of the proposed conditional estimators of the latent predictor in different simulation scenarios.

math.ST

Balancing Evidentiary Value and Sample Size of Adaptive Designs with Application to Animal Experiments

Reducing the number of experimental units is one of the three pillars of the 3R principles (Replace, Reduce, Refine) in animal research. At the same time, statistical error rates need to be controlled to enable reliable inferences and decisions. This paper proposes to adopt diagnostic likelihood ratios and the diagnostic odds ratio to statistical hypothesis tests and to adjust it for sample size to obtain a novel measure to quantify for the evidentiary value of one experimental unit. The experimental unit information index (EUII) is based on power, Type-I error and sample size, and has attractive interpretations both in terms of frequentist error rates and Bayesian posterior odds. We introduce the EUII in simple statistical test settings and show that its asymptotic value depends only on the assumed relative effect size under the alternative. We then extend the definition to adaptive designs where early stopping for efficacy or futility may cause reductions in sample size. Application to group-sequential designs show the usefulness of the approach when the goal is to maximize the evidentiary value of one experimental unit. A reanalysis of 2738 animal experiments with simulated results from (post-hoc) interim analyses illustrates the possible savings in sample size.

stat.ME

Linear regression with known noise distribution up to a scale: The reward of not using the OLSE

While the ordinary least squares estimator (OLSE) is still the most used estimator in linear regression models, other estimators can be more efficient when the error distribution is not Gaussian. In this paper, our goal is to evaluate this efficiency in the case of the Maximum Likelihood estimator (MLE) when the noise distribution belongs to a scale family. Under some regularity conditions, we show that (\beta_n,s_n), the MLE of the unknown regression vector \beta_0 and the scale s_0 exists and give the expression of the asymptotic efficiency of \beta_n over the OLSE. For given three scale families of densities, we quantify the true statistical gain of the MLE as a function of their deviation from the Gaussian family. To illustrate the theory, we present simulation results for different settings and also compare the MLE to the OLSE for the real market fish dataset.

math.ST

Parametric convergence rate of a non-parametric estimator in multivariate mixtures of power series distributions under conditional independence

The conditional independence assumption has recently appeared in a growing body of literature on the estimation of multivariate mixtures. We consider here conditionally independent multivariate mixtures of power series distributions with infinite support, to which belong Poisson, Geometric or Negative Binomial mixtures. We show that for all these mixtures, the non-parametric maximum likelihood estimator converges to the truth at the rate $(\log (nd))^{1+d/2} n^{-1/2}$ in the Hellinger distance, where $n$ denotes the size of the observed sample and $d$ represents the dimension of the mixture. Using this result, we then construct a new non-parametric estimator based on the maximum likelihood estimator that converges with the parametric rate $n^{-1/2}$ in all $\ell_p$-distances, for $p \ge 1$. These convergences rates are supported by simulations and the theory is illustrated using the famous V\'{e}lib dataset of the bike sharing system of Paris. We also introduce a testing procedure for whether the conditional independence assumption is satisfied for a given sample. This testing procedure is applied for several multivariate mixtures, with varying levels of dependence, and is thereby shown to distinguish well between conditionally independent and dependent mixtures. Finally, we use this testing procedure to investigate whether conditional independence holds for V\'{e}lib dataset.

math.ST

Parametric convergence rate of some nonparametric estimators in mixtures of power series distributions

We consider the problem of estimating a mixture of power series distributions with infinite support, to which belong very well-known models such as Poisson, Geometric, Logarithmic or Negative Binomial probability mass functions. We consider the nonparametric maximum likelihood estimator (NPMLE) and show that, under very mild assumptions, it converges to the true mixture distribution $\pi_0$ at a rate no slower than $(\log n)^{3/2} n^{-1/2}$ in the Hellinger distance. Recent work on minimax lower bounds suggests that the logarithmic factor in the obtained Hellinger rate of convergence can not be improved, at least for mixtures of Poisson distributions. Furthermore, we construct nonparametric estimators that are based on the NPMLE and show that they converge to $\pi_0$ at the parametric rate $n^{-1/2}$ in the $\ell_p$-norm ($p \in [1, \infty]$ or $p \in [2, \infty])$: The weighted least squares and hybrid estimators. Simulations and a real data application are considered to assess the performance of all estimators we study in this paper and illustrate the practical aspect of the theory. The simulations results show that the NPMLE has the best performance in the Hellinger, $\ell_1$ and $\ell_2$ distances in all scenarios. Finally, to construct confidence intervals of the true mixture probability mass function, both the nonparametric and parametric bootstrap procedures are considered. Their performances are compared with respect to the coverage and length of the resulting intervals.

math.ST

Identifiability in Unlinked Linear Regression: Some Results and Open Problems

A tacit assumption in classical linear regression problems is the full knowledge of the existing link between the covariates and responses. In Unlinked Linear Regression (ULR) this link is either partially or completely missing. While the reasons causing such missingness can be different, a common challenge in statistical inference is the potential non-identifiability of the regression parameter. In this note, we review the existing literature on identifiability when the $d \ge 2$ components of the vector of covariates are independent and identically distributed. When these components have different distributions, we show that it is not possible to prove similar theorems in the general case. Nevertheless, we prove some identifiability results, either under additional parametric assumptions for $d \ge 2$ or conditions on the fourth moments in the case $d=2$. Finally, we draw some interesting connections between the ULR and the well established field of Independent Component Analysis (ICA).

math.ST

Asymptotic theory for nonparametric testing of $k$-monotonicity in discrete distributions

In shape-constrained nonparametric inference, it is often necessary to perform preliminary tests to verify whether a probability mass function (p.m.f.) satisfies qualitative constraints such as monotonicity, convexity, or in general $k$-monotonicity. In this paper, we are interested in nonparametric testing of $k$-monotonicity of a finitely supported discrete distribution. We consider a unified testing framework based on a natural statistic which is directly derived from the very definition of $k$-monotonicity. The introduced framework allows us to design a new consistent method to select the unknown knot points that are required to consistently approximate the limit distribution of several test statistics based either on the empirical measure or the shape-constrained estimators of the p.m.f. We show that the resulting tests are asymptotically valid and consistent for any fixed alternative. Additionally, for the test based solely on the empirical measure, we study the asymptotic power under contiguous alternatives and derive a quantitative separation result that provides sufficient conditions to achieve a given power. We employ this test to design an estimator for the largest parameter $k \in \mathbb N_0$ such that the p.m.f. is $j$-monotone for all $j = 0, \ldots, k$, and show that the estimator is different from the true parameter with probability which is asymptotically smaller than the nominal level of the test. Finally, we conduct an extensive simulation study to validate the theory, assess the finite-sample performance of the proposed methods, and illustrate them on several real datasets.

math.ST

Estimation and convergence rates in the distributional single index model

The distributional single index model is a semiparametric regression model in which the conditional distribution functions $P(Y \leq y | X = x) = F_0(θ_0(x), y)$ of a real-valued outcome variable $Y$ depend on $d$-dimensional covariates $X$ through a univariate, parametric index function $θ_0(x)$, and increase stochastically as $θ_0(x)$ increases. We propose least squares approaches for the joint estimation of $θ_0$ and $F_0$ in the important case where $θ_0(x) = α_0^{\top}x$ and obtain convergence rates of $n^{-1/3}$, thereby improving an existing result that gives a rate of $n^{-1/6}$. A simulation study indicates that the convergence rate for the estimation of $α_0$ might be faster. Furthermore, we illustrate our methods in a real data application that demonstrates the advantages of shape restrictions in single index models.

math.ST

Linear regression with unmatched data: a deconvolution perspective

Consider the regression problem where the response $Y\in\mathbb{R}$ and the covariate $X\in\mathbb{R}^d$ for $d\geq 1$ are \textit{unmatched}. Under this scenario, we do not have access to pairs of observations from the distribution of $(X, Y)$, but instead, we have separate datasets $\{Y_i\}_{i=1}^n$ and $\{X_j\}_{j=1}^m$, possibly collected from different sources. We study this problem assuming that the regression function is linear and the noise distribution is known or can be estimated. We introduce an estimator of the regression vector based on deconvolution and demonstrate its consistency and asymptotic normality under an identifiability assumption. In the general case, we show that our estimator (DLSE: Deconvolution Least Squared Estimator) is consistent in terms of an extended $\ell_2$ norm. Using this observation, we devise a method for semi-supervised learning, i.e., when we have access to a small sample of matched pairs $(X_k, Y_k)$. Several applications with synthetic and real datasets are considered to illustrate the theory.

math.ST

Assessing replicability with the sceptical p-value: Type-I error control and sample size planning

We study a statistical framework for replicability based on a recently proposed quantitative measure of replication success, the sceptical $p$-value. A recalibration is proposed to obtain exact overall Type-I error control if the effect is null in both studies and additional bounds on the partial and conditional Type-I error rate, which represent the case where only one study has a null effect. The approach avoids the double dichotomization for significance of the two-trials rule and has larger project power to detect existing effects over both studies in combination. It can also be used for power calculations and requires a smaller replication sample size than the two-trials rule for already convincing original studies. We illustrate the performance of the proposed methodology in an application to data from the Experimental Economics Replication Project.

stat.ME

Profile least squares estimators in the monotone single index model

We consider least squares estimators of the finite regression parameter $α$ in the single index regression model $Y=ψ(α^T X)+ε$, where $X$ is a $d$-dimensional random vector, $\E(Y|X)=ψ(α^T X)$, and where $ψ$ is monotone. It has been suggested to estimate $α$ by a profile least squares estimator, minimizing $\sum_{i=1}^n(Y_i-ψ(α^T X_i))^2$ over monotone $ψ$ and $α$ on the boundary $S_{d-1}$of the unit ball. Although this suggestion has been around for a long time, it is still unknown whether the estimate is $\sqrt{n}$ convergent. We show that a profile least squares estimator, using the same pointwise least squares estimator for fixed $α$, but using a different global sum of squares, is $\sqrt{n}$-convergent and asymptotically normal. The difference between the corresponding loss functions is studied and also a comparison with other methods is given.

math.ST

Dynamic gravitational excitation of structural resonances in the hertz regime using two rotating bars

With the planning of new ambitious gravitational wave (GW) observatories, fully controlled laboratory experiments on dynamic gravitation become more and more important. Such new experiments can provide new insights in potential dynamic effects such as gravitational shielding or energy flow and might contribute to bringing light into the mystery still surrounding gravity. Here we present a laboratory-based transmitter-detector experiment using two rotating bars as transmitter and a 42 Hz, high-Q bending beam resonator as detector. Using a highly precise phase control to synchronize the rotating bars, a dynamic gravitational field emerges that excites the bending motion with amplitudes up to 100 nm/s or 370 pm, which is a factor of 500 above the thermal noise. The two-transmitter design enables the investigation of different setup configurations. The detector movement is measured optically, using three commercial interferometers. Acoustical, mechanical, and electrical isolation, a temperature-stable environment, and lock-in detection are central elements of the setup. The moving load response of the detector is numerically calculated based on Newton's law of gravitation via discrete volume integration, showing excellent agreement between measurement and theory both in amplitude and phase. The near field gravitational energy transfer is 10$^{25}$ times higher than what is expected from GW analysis.

physics.ins-det

Measurement and theory of gravitational coupling between resonating beams

Recent spectacular results of gravitational waves obtained by the LIGO system, with frequencies in the 100 Hz regime, make corresponding laboratory experiments with full control over cause and effect of great importance. Dynamic measurements of gravitation in the laboratory have to date been scarce, due to difficulties in assessing non-gravitational crosstalk and the intrinsically weak nature of gravitational forces. In fact, fully controlled quantitative experiments have so far been limited to frequencies in the mHz regime. New experiments in gravity might also yield new physics, thereby opening avenues towards a theory that explains all of physics within one coherent framework. Here we introduce a new, fully-characterized experiment at three orders of magnitude higher frequencies. It allows experimenters to quantitatively determine the dynamic gravitational interaction between two parallel beams vibrating at 42 Hz in bending motion. The large amplitude vibration of the transmitter beam produces gravitationally-induced motion with amplitudes up to 1E-11 m of the resonant detector beam. The reliable measurement with sub-pm displacement resolution is made possible by a set-up which combines acoustical, mechanical and electrical isolation, a temperature-stable environment, heterodyne laser interferometry and lock-in detection. The interaction is quantitatively modelled based on Newton's theory. Our initial results agree with the theory to within about three percent in amplitude. Based on a power balance analysis, we determined the near-field gravitational energy flow from the transmitter to the detector to be 2.5 E-20 J/s, and to decay with distance as d-4. We expect our experiment to make significant progress in directions where current experimental evidence for dynamic gravitation is limited, such as the dynamic determination of G, inverse-square law, and gravitational shielding.

physics.app-ph

Unlinked monotone regression

We consider so-called univariate unlinked (sometimes ``decoupled,'' or ``shuffled'') regression when the unknown regression curve is monotone. In standard monotone regression, one observes a pair $(X,Y)$ where a response $Y$ is linked to a covariate $X$ through the model $Y= m_0(X) + ε$, with $m_0$ the (unknown) monotone regression function and $ε$ the unobserved error (assumed to be independent of $X$). In the unlinked regression setting one gets only to observe a vector of realizations from both the response $Y$ and from the covariate $X$ where now $Y \stackrel{d}{=} m_0(X) + ε$. There is no (observed) pairing of $X$ and $Y$. Despite this, it is actually still possible to derive a consistent non-parametric estimator of $m_0$ under the assumption of monotonicity of $m_0$ and knowledge of the distribution of the noise $ε$. In this paper, we establish an upper bound on the rate of convergence of such an estimator under minimal assumption on the distribution of the covariate $X$. We discuss extensions to the case in which the distribution of the noise is unknown. We develop a second order algorithm for its computation, and we demonstrate its use on synthetic data. Finally, we apply our method (in a fully data driven way, without knowledge of the error distribution) on longitudinal data from the US Consumer Expenditure Survey.

stat.ME

Testing for spherical and elliptical symmetry

We construct new testing procedures for spherical and elliptical symmetry based on the characterization that a random vector $X$ with finite mean has a spherical distribution if and only if $\Ex[u^\top X | v^\top X] = 0$ holds for any two perpendicular vectors $u$ and $v$. Our test is based on the Kolmogorov-Smirnov statistic, and its rejection region is found via the spherically symmetric bootstrap. We show the consistency of the spherically symmetric bootstrap test using a general Donsker theorem which is of some independent interest. For the case of testing for elliptical symmetry, the Kolmogorov-Smirnov statistic has an asymptotic drift term due to the estimated location and scale parameters. Therefore, an additional standardization is required in the bootstrap procedure. In a simulation study, the size and the power properties of our tests are assessed for several distributions and the performance is compared to that of several competing procedures.

math.ST

Score estimation in the monotone single index model

We consider estimation in the single index model where the link function is monotone. For this model a profile least squares estimator has been proposed to estimate the unknown link function and index. Although it is natural to propose this procedure, it is still unknown whether it produces index estimates which converge at the parametric rate. We show that this holds if we solve a score equation corresponding to this least squares problem. Using a Lagrangian formulation, we show how one can solve this score equation without any reparametrization. This makes it easy to solve the score equations in high dimensions. We also compare our method with the Effective Dimension Reduction (EDR) and the Penalized Least Squares Estimator (PLSE) methods, both available on CRAN as R packages, and compare with link-free methods, where the covariates are ellipticallly symmetric.

math.ST

Linear regression estimation in non-linear single index models

In this article, we consider the problem of estimating the index parameter $α_0$ in the single index model $E[Y |X] = f_0(α_0^T X)$ with $f_0$ the unknown ridge function defined on $\mathbb{R}$, $X$ a d-dimensional covariate and $Y$ the response. We show that when $X$ is Gaussian, then $α_0$ can be consistently estimated by regressing the observed responses $Y_i$, $i = 1, . . ., n$ on the covariates $X_1, . . ., X_n$ after centering and rescaling. The method works without any additional smoothness assumptions on $f_0$ and only requires that $cov(f_0(α_0^T X),α_0^TX) \neq 0$, which is always satisfied by monotone and non-constant functions $f_0$. We show that our estimator is asymptotically normal and give the expression with its asymptotic variance. The approach is illustrated through a simulation study.

math.ST