SearcharxivSearch

arXiv subjects

Florian Stijven

Publications and source records attributed to Florian Stijven.

3 recordsLinked to original sources

Estimating the Wasserstein barycenter of one-dimensional distributions under sparse sampling

We study distributional data under sparse sampling where each unit is represented by a probability distribution on the real line observed only through a small i.i.d.~sample. A natural notion of central tendency for one-dimensional distributional data is the Wasserstein barycenter, whose quantile function is the pointwise average of the unit-level quantile functions. We focus on pointwise estimation of the Wasserstein barycenter quantile function: at a given quantile level, the target is the population mean of the corresponding unit-level quantiles. A naive plug-in estimator is the empirical Wasserstein barycenter, which treats observed unit-level empirical distributions as the true latent unit-level distributions. Under sparse sampling, however, this estimator can be severely biased. We propose an approach that avoids directly estimating either the unit-level distributions or the full population law of distributions. We start with the more ambitious goal of characterizing the distribution of latent unit-level quantiles at a given quantile level. We show that this distribution can be written in terms of the marginal distributions of the unit-level CDF values, which can be estimated using binomial mixture methods. This motivates our estimator, the marginal-constructed barycenter (MCB) estimator, obtained by taking the mean of the estimated distribution of latent unit-level quantiles. We establish conditions under which the MCB estimator is pointwise consistent and asymptotically normal, and show through simulations that it can substantially outperform the empirical Wasserstein barycenter under sparse sampling. We illustrate the method in an analysis of HIV-1 sequence data from the HVTN 502/503 vaccine efficacy trials, using the barycenter to summarize and compare within-participant distributions of viral sequence features when only a small number of sequences are available per participant.

stat.ME

Evaluation of Surrogate Endpoints Based on Meta-Analysis with Surrogate Indices

The meta-analytic (MA) framework is the gold standard for evaluating putative surrogate endpoints but it is not well-suited for complex surrogates. We address this limitation by considering real-valued summaries of complex surrogates as alternative, univariate putative surrogates. We focus on the surrogate index as a specific summary measure with desirable properties. We first formalize the data-generating mechanism underlying the MA framework, making explicit the assumptions required for valid inferences in any evaluation of trial-level surrogacy. These assumptions are often left implicit in the MA framework. Building on this formalization, we show that, under certain conditions, the surrogate index maximizes the trial-level surrogacy among all real-valued summaries. When the surrogate index is estimated, we show that valid inference about the trial-level surrogacy of this estimated surrogate index is possible. The proposed approach can be implemented using standard software and is illustrated using COVID-19 vaccine efficacy trials, where antibody markers are evaluated as candidate surrogate endpoints.

stat.ME

A Reflection on the Impact of Misspecifying Unidentifiable Causal Inference Models in Surrogate Endpoint Evaluation

Surrogate endpoints are often used in place of expensive, delayed, or rare true endpoints in clinical trials. However, regulatory authorities require thorough evaluation to accept these surrogate endpoints as reliable substitutes. One evaluation approach is the information-theoretic causal inference framework, which quantifies surrogacy using the individual causal association (ICA). Like most causal inference methods, this approach relies on models that are only partially identifiable. For continuous outcomes, a normal model is often used. Based on theoretical elements and a Monte Carlo procedure we studied the impact of model misspecification across two scenarios: 1) the true model is based on a multivariate t-distribution, and 2) the true model is based on a multivariate log-normal distribution. In the first scenario, the misspecification has a negligible impact on the results, while in the second, it has a significant impact when the misspecification is detectable using the observed data. Finally, we analyzed two data sets using the normal model and several D-vine copula models that were indistinguishable from the normal model based on the data at hand. We observed that the results may vary when different models are used.

stat.ME