SearcharxivSearch

arXiv subjects

Luca Insolia

Publications and source records attributed to Luca Insolia.

9 recordsLinked to original sources

Towards Open Science: Monitoring Crustal Deformations in North America

The study of the Earth's behavior has greatly benefited from the widespread deployment of Global Navigation Satellite Systems (GNSS), enabling large-scale monitoring of crustal deformation and long-term geophysical trends. In this work, we focus on the North American region, where complex tectonic activity, particularly along the western margin, requires methods capable of processing and analyzing large collections of GNSS time series distributed across extensive spatial domains. Analyzing the full GNSS network provides a coherent view of deformation across multiple scales, allowing detection of long-wavelength signals, subtle intraplate strain and regionally consistent velocity fields that are difficult to capture through local or subsampled analyses. Despite the availability of such data, existing methodologies remain computationally prohibitive for large-scale analyses across this region (or others). To address this limitation, we introduce a highly scalable framework (implemented in open-source software) that enables inference on crustal deformation across the full North American GNSS network using standard computational resources. The proposed method achieves substantial computational gains while maintaining inferential performance comparable to existing approaches and confirms existing tectonic trends, thereby supporting efficient large-scale monitoring and contributing to ongoing efforts toward "Open Science".

physics.geo-ph

Equivalence Testing Under Privacy Constraints

Protecting individual privacy is essential across research domains, from socio-economic surveys to big-tech user data. This need is particularly acute in healthcare, where analyses often involve sensitive patient information. A typical example is comparing treatment efficacy across hospitals or ensuring consistency in diagnostic laboratory calibrations, both requiring privacy-preserving statistical procedures. However, standard equivalence testing procedures for differences in proportions or means, commonly used to assess average equivalence, can inadvertently disclose sensitive information. To address this problem, we develop differentially private equivalence testing procedures that rely on simulation-based calibration, as the finite-sample distribution is analytically intractable. Our approach introduces a unified framework, termed DP-TOST, for conducting differentially private equivalence testing of both means and proportions. Through numerical simulations and real-world applications, we demonstrate that the proposed method maintains type-I error control at the nominal level and achieves power comparable to its non-private counterpart as the privacy budget and/or sample size increases, while ensuring strong privacy guarantees. These findings establish a reliable and practical framework for privacy-preserving equivalence testing in high-stakes fields such as healthcare, among others.

stat.AP

Bridging the gap between experimental burden and statistical power for quantiles equivalence testing

Testing the equivalence of multiple quantiles between two populations is important in many scientific applications, such as clinical trials, where conventional mean-based methods may be inadequate. This is particularly relevant in bridging studies that compare drug responses across different experimental conditions or patient populations. These studies often aim to assess whether a proposed dose for a target population achieves pharmacokinetic levels comparable to those of a reference population where efficacy and safety have been established. The focus is on extreme quantiles which directly inform both efficacy and safety assessments. When analyzing heterogeneous Gaussian samples, where a single quantile of interest is estimated, the existing Two One-Sided Tests method for quantile equivalence testing (qTOST) tends to be overly conservative. To mitigate this behavior, we introduce $\alpha$-qTOST, a finite-sample adjustment that achieves uniformly higher power compared to qTOST while maintaining the test size at the nominal level. Moreover, we extend the quantile equivalence framework to simultaneously assess equivalence across multiple quantiles. Through theoretical guarantees and an extensive simulation study, we demonstrate that $\alpha$-qTOST offers substantial improvements, especially when testing extreme quantiles under heteroskedasticity and with small, unbalanced sample sizes. We illustrate these advantages through two case studies, one in HIV drug development, where a bridging clinical trial examines exposure distributions between male and female populations with unbalanced sample sizes, and another in assessing the reproducibility of an identical experimental protocol performed by different operators for generating biodistribution profiles of topically administered and locally acting products.

stat.ME

Bioequivalence Assessment for Locally Acting Drugs: A Framework for Feasible and Efficient Evaluation

Equivalence testing plays a key role in several domains, such as the development of generic medical products, which are therapeutically equivalent to brand-name drugs but with reduced cost and increased accessibility. Promoting access to generics is a critical public health issue with substantial societal implications, but establishing equivalence is particularly challenging in multivariate settings. A notable example refers to locally acting drugs designed to exert their therapeutic effects at a localized area where they are administered rather than being absorbed into the bloodstream, where complex experimental protocols lead to reduced sample sizes and substantial experimental noise. Traditional approaches, such as the Two One-Sided Tests (TOST), cannot adequately tackle the complex multivariate nature of such data. In this work, we develop an adjustment for the TOST procedure by simultaneously correcting its significance level and equivalence margins to ensure control of the test size and increase its power. In large samples, this approach leads to an optimal adjustment for the univariate TOST procedure. In multivariate settings, where we show that an optimal adjustment does not exist, our proposal maintains equal marginal test sizes and overall size control while maximizing power in important cases. Through extensive simulation studies and a case study on multivariate bioequivalence assessment for two antifungal topical products, we demonstrate the superior performance of our method across various scenarios encountered in practice.

stat.ME

Multivariate Adjustments for Average Equivalence Testing

Multivariate (average) equivalence testing is widely used to assess whether the means of two conditions of interest are `equivalent' for different outcomes simultaneously. The multivariate Two One-Sided Tests (TOST) procedure is typically used in this context by checking if, outcome by outcome, the marginal $100(1-2\alpha$)\% confidence intervals for the difference in means between the two conditions of interest lie within pre-defined lower and upper equivalence limits. This procedure, known to be conservative in the univariate case, leads to a rapid power loss when the number of outcomes increases, especially when one or more outcome variances are relatively large. In this work, we propose a finite-sample adjustment for this procedure, the multivariate $\alpha$-TOST, that consists in a correction of $\alpha$, the significance level, taking the (arbitrary) dependence between the outcomes of interest into account and making it uniformly more powerful than the conventional multivariate TOST. We present an iterative algorithm allowing to efficiently define $\alpha^{\star}$, the corrected significance level, a task that proves challenging in the multivariate setting due to the inter-relationship between $\alpha^{\star}$ and the sets of values belonging to the null hypothesis space and defining the test size. We study the operating characteristics of the multivariate $\alpha$-TOST both theoretically and via an extensive simulation study considering cases relevant for real-world analyses -- i.e.,~relatively small sample sizes, unknown and heterogeneous variances, and different correlation structures -- and show the superior finite-sample properties of the multivariate $\alpha$-TOST compared to its conventional counterpart. We finally re-visit a case study on ticlopidine hydrochloride and compare both methods when simultaneously assessing bioequivalence for multiple pharmacokinetic parameters.

stat.ME

Inference for Large Scale Regression Models with Dependent Errors

The exponential growth in data sizes and storage costs has brought considerable challenges to the data science community, requiring solutions to run learning methods on such data. While machine learning has scaled to achieve predictive accuracy in big data settings, statistical inference and uncertainty quantification tools are still lagging. Priority scientific fields collect vast data to understand phenomena typically studied with statistical methods like regression. In this setting, regression parameter estimation can benefit from efficient computational procedures, but the main challenge lies in computing error process parameters with complex covariance structures. Identifying and estimating these structures is essential for inference and often used for uncertainty quantification in machine learning with Gaussian Processes. However, estimating these structures becomes burdensome as data scales, requiring approximations that compromise the reliability of outputs. These approximations are even more unreliable when complexities like long-range dependencies or missing data are present. This work defines and proves the statistical properties of the Generalized Method of Wavelet Moments with Exogenous variables (GMWMX), a highly scalable, stable, and statistically valid method for estimating and delivering inference for linear models using stochastic processes in the presence of data complexities like latent dependence structures and missing data. Applied examples from Earth Sciences and extensive simulations highlight the advantages of the GMWMX.

stat.ME

Tk-merge: Computationally Efficient Robust Clustering Under General Assumptions

We address general-shaped clustering problems under very weak parametric assumptions with a two-step hybrid robust clustering algorithm based on trimmed k-means and hierarchical agglomeration. The algorithm has low computational complexity and effectively identifies the clusters also in presence of data contamination. We also present natural generalizations of the approach as well as an adaptive procedure to estimate the amount of contamination in a data-driven fashion. Our proposal outperforms state-of-the-art robust, model-based methods in our numerical simulations and real-world applications related to color quantization for image analysis, human mobility patterns based on GPS data, biomedical images of diabetic retinopathy, and functional data across weather stations.

stat.ME

Doubly Robust Feature Selection with Mean and Variance Outlier Detection and Oracle Properties

We propose a general approach to handle data contaminations that might disrupt the performance of feature selection and estimation procedures for high-dimensional linear models. Specifically, we consider the co-occurrence of mean-shift and variance-inflation outliers, which can be modeled as additional fixed and random components, respectively, and evaluated independently. Our proposal performs feature selection while detecting and down-weighting variance-inflation outliers, detecting and excluding mean-shift outliers, and retaining non-outlying cases with full weights. Feature selection and mean-shift outlier detection are performed through a robust class of nonconcave penalization methods. Variance-inflation outlier detection is based on the penalization of the restricted posterior mode. The resulting approach satisfies a robust oracle property for feature selection in the presence of data contamination -- which allows the number of features to exponentially increase with the sample size -- and detects truly outlying cases of each type with asymptotic probability one. This provides an optimal trade-off between a high breakdown point and efficiency. Computationally efficient heuristic procedures are also presented. We illustrate the finite-sample performance of our proposal through an extensive simulation study and a real-world application.

stat.ME

Simultaneous Feature Selection and Outlier Detection with Optimality Guarantees

Sparse estimation methods capable of tolerating outliers have been broadly investigated in the last decade. We contribute to this research considering high-dimensional regression problems contaminated by multiple mean-shift outliers which affect both the response and the design matrix. We develop a general framework for this class of problems and propose the use of mixed-integer programming to simultaneously perform feature selection and outlier detection with provably optimal guarantees. We characterize the theoretical properties of our approach, i.e. a necessary and sufficient condition for the robustly strong oracle property, which allows the number of features to exponentially increase with the sample size; the optimal estimation of the parameters; and the breakdown point of the resulting estimates. Moreover, we provide computationally efficient procedures to tune integer constraints and to warm-start the algorithm. We show the superior performance of our proposal compared to existing heuristic methods through numerical simulations and an application investigating the relationships between the human microbiome and childhood obesity.

stat.ME