SearcharxivSearch

arXiv subjects

Lucas Resende

Publications and source records attributed to Lucas Resende.

6 recordsLinked to original sources

Improved Concentration for Mean Estimators via Shrinkage

We study a class of robust mean estimators $\widehat{\mu}$ obtained by adaptively shrinking the weights of sample points far from a base estimator $\widehat{\kappa}$. Given a data-dependent scaling factor $\widehat{\alpha}$ and a weighting function $w:[0, \infty) \to [0,1]$, we let $\widehat{\mu}=\widehat{\kappa} + \frac{1}{n}\sum_{i=1}^n(X_i - \widehat{\kappa})w(\widehat{\alpha}|X_i-\widehat{\kappa}|)$. We prove that, under mild assumptions over $w$, these estimators achieve stronger concentration bounds than the base estimate $\widehat{\kappa}$, including sub-Gaussian guarantees. This framework unifies and extends several existing approaches to robust mean estimation in $\R$, and can also be generalized to the multivariate setting. Through numerical experiments, we show that our shrinking approach translates to faster concentration, even for small sample sizes.

math.ST

Statistical Inference in Large Multi-way Networks

We propose the Polyads estimator, a new method to estimate structural parameters in weighted multi-way networks while controlling for rich, arbitrary structures of fixed effects. The method is based on a series of classification tasks and is agnostic to both the number and structure of fixed effects. Unlike full maximum likelihood, our estimator does not suffer from the incidental parameter problem: it is consistent and satisfies a Central Limit Theorem with no asymptotic bias, even when some dimensions of the network are short. For sparsely connected networks, it is also computationally faster than PPML. We provide experimental evidence that our estimator yields more reliable confidence intervals, i.e., better empirical coverage, than PPML and its bias-correction strategies. These improvements hold even under model misspecification and are more pronounced in sparse settings. While PPML remains competitive in dense, low-dimensional data, our approach offers a robust alternative for multi-way models that scales efficiently with sparsity. We apply the method to French health insurance claims data to study how a 2017 physician fee reform affected the geography and gender composition of doctor-patient connections.

econ.EM

Robust high-dimensional Gaussian and bootstrap approximations for trimmed sample means

Robust mean estimation has largely focused on concentration guarantees under heavy tails and contamination. We study robustness from a different perspective: high-dimensional Gaussian and bootstrap approximations. We show that trimmed sample means admit Gaussian and bootstrap approximations under finite p-th moment assumptions, even in high-dimensional regimes and in the presence of adversarial contamination. Our bounds recover, up to the dependence on the moment parameter, the rates available for the empirical mean under light tails, while requiring substantially weaker moment assumptions. We further extend the Gaussian approximation to VC-subgraph classes and apply it to robust vector mean estimation under arbitrary norms, obtaining bounds with optimal Gaussian-width complexity. Finally, we develop uniform confidence intervals based on the bootstrap approximation and show empirically that they maintain coverage under heavy tails and adversarial contamination.

math.ST

Deep Hashing via Householder Quantization

Hashing is at the heart of large-scale image similarity search, and recent methods have been substantially improved through deep learning techniques. Such algorithms typically learn continuous embeddings of the data. To avoid a subsequent costly binarization step, a common solution is to employ loss functions that combine a similarity learning term (to ensure similar images are grouped to nearby embeddings) and a quantization penalty term (to ensure that the embedding entries are close to binarized entries, e.g., -1 or 1). Still, the interaction between these two terms can make learning harder and the embeddings worse. We propose an alternative quantization strategy that decomposes the learning problem in two stages: first, perform similarity learning over the embedding space with no quantization; second, find an optimal orthogonal transformation of the embeddings so each coordinate of the embedding is close to its sign, and then quantize the transformed embedding through the sign function. In the second step, we parametrize orthogonal transformations using Householder matrices to efficiently leverage stochastic gradient descent. Since similarity measures are usually invariant under orthogonal transformations, this quantization strategy comes at no cost in terms of performance. The resulting algorithm is unsupervised, fast, hyperparameter-free and can be run on top of any existing deep hashing or metric learning algorithm. We provide extensive experimental results showing that this approach leads to state-of-the-art performance on widely used image datasets, and, unlike other quantization strategies, brings consistent improvements in performance to existing deep hashing algorithms.

cs.CV

Trimmed sample means for robust uniform mean estimation and regression

It is well-known that trimmed sample means are robust against heavy tails and data contamination. This paper analyzes the performance of trimmed means and related methods in two novel contexts. The first one consists of estimating expectations of functions in a given family, with uniform error bounds; this is closely related to the problem of estimating the mean of a random vector under a general norm. The second problem considered is that of regression with quadratic loss. In both cases, trimmed-mean-based estimators are the first to obtain optimal dependence on the (adversarial) contamination level. Moreover, they also match or improve upon the state of the art in terms of heavy tails. Experiments with synthetic data show that a natural ``trimmed mean linear regression'' method often performs better than both ordinary least squares and alternative methods based on median-of-means.

math.ST

Quantifying protocols for safe school activities

By the peak of COVID-19 restrictions on April 8, 2020, up to 1.5 billion students across 188 countries were by the suspension of physical attendance in schools. Schools were among the first services to reopen as vaccination campaigns advanced. With the emergence of new variants and infection waves, the question now is to find safe protocols for the continuation of school activities. We need to understand how reliable these protocols are under different levels of vaccination coverage, as many countries have a meager fraction of their population vaccinated, including Uganda where the coverage is about 8\%. We investigate the impact of face-to-face classes under different protocols and quantify the surplus number of infected individuals in a city. Using the infection transmission when schools were closed as a baseline, we assess the impact of physical school attendance in classrooms with poor air circulation. We find that (i) resuming school activities with people only wearing low-quality masks leads to a near fivefold city-wide increase in the number of cases even if all staff is vaccinated, (ii) resuming activities with students wearing good-quality masks and staff wearing N95s leads to about a threefold increase, (iii) combining high-quality masks and active monitoring, activities may be carried out safely even with low vaccination coverage. These results highlight the effectiveness of good mask-wearing. Compared to ICU costs, high-quality masks are inexpensive and can help curb the spreading. Classes can be carried out safely, provided the correct set of measures are implemented.

physics.soc-ph