SearcharxivSearch

arXiv subjects

Gustavo Amorim

Publications and source records attributed to Gustavo Amorim.

3 recordsLinked to original sources

Generalized raking and stabilized weights for regression modeling in two-phase samples

In regression models fitted to data from complex survey designs, sampling weights often incorporate non-essential variation, inflating variance estimates. Stabilized weights mitigate this issue by adjusting sampling weights to account for variation explained by covariates. In the context of two-phase sampling, we evaluate the performance of optimal stabilized weights and propose combining the stabilized weight estimator with generalized raking, a class of efficient design-based estimators. This combination improves efficiency by reducing unnecessary weight variation and leveraging information from auxiliary variables. We show this combination can be implemented using the standard statistical package that handles two-phase samples and generalized raking. Simulation studies demonstrate that the proposed estimator enhances precision under realistic two-phase designs, though efficiency gains may be limited in highly informative designs. The developed methods were applied to a large multinational two-phase study of Kaposi sarcoma among people living with HIV.

stat.ME

Three-phase generalized raking and multiple imputation estimators to address error-prone data

Validation studies are often used to obtain more reliable information in settings with error-prone data. Validated data on a subsample of subjects can be used together with error-prone data on all subjects to improve estimation. In practice, more than one round of data validation may be required, and direct application of standard approaches for combining validation data into analyses may lead to inefficient estimators since the information available from intermediate validation steps is only partially considered or even completely ignored. In this paper, we present two novel extensions of multiple imputation and generalized raking estimators that make full use of all available data. We show through simulations that incorporating information from intermediate steps can lead to substantial gains in efficiency. This work is motivated by and illustrated in a study of contraceptive effectiveness among 82,957 women living with HIV whose data were originally extracted from electronic medical records, of whom 4855 had their charts reviewed, and a subsequent 1203 also had a telephone interview to validate key study variables.

stat.ME

Fitting Probabilistic Index Models on Large Datasets

Recently, Thas et al. (2012) introduced a new statistical model for the probability index. This index is defined as $P(Y \leq Y^*|X, X^*)$ where Y and Y* are independent random response variables associated with covariates X and X* [...] Crucially to estimate the parameters of the model, a set of pseudo-observations is constructed. For a sample size n, a total of $n(n-1)/2$ pairwise comparisons between observations is considered. Consequently for large sample sizes, it becomes computationally infeasible or even impossible to fit the model as the set of pseudo-observations increases nearly quadratically. In this dissertation, we provide two solutions to fit a probabilistic index model. The first algorithm consists of splitting the entire data set into unique partitions. On each of these, we fit the model and then aggregate the estimates. A second algorithm is a subsampling scheme in which we select $K << n$ observations without replacement and after B iterations aggregate the estimates. In Monte Carlo simulations, we show how the partitioning algorithm outperforms the latter [...] We illustrate the partitioning algorithm and the interpretation of the probabilistic index model on a real data set (Przybylski and Weinstein, 2017) of n = 116,630 where we compare it against the ordinary least squares method. By modelling the probabilistic index, we give an intuitive and meaningful quantification of the effect of the time adolescents spend using digital devices such as smartphones on self-reported mental well-being. We show how moderate usage is associated with an increased probability of reporting a higher mental well-being compared to random adolescents who do not use a smartphone. On the other hand, adolescents who excessively use their smartphone are associated with a higher probability of reporting a lower mental well-being than randomly chosen peers who do not use a smartphone.[...]

stat.CO