SearcharxivSearch

arXiv subjects

Vivian Viallon

Publications and source records attributed to Vivian Viallon.

At least 19 recordsLinked to original sources

Data Shared Neighbourhood Selection for multi-condition network inference

External stresses may affect both the circulating levels of specific biomarkers and disturb the correlation structures across molecular entities. The contribution of both types of dysregulations to the subsequent risk of disease are yet to be evaluated. We propose Data Shared Neighbourhood Selection (DSNS), a joint network inference method for estimating preserved and altered conditional association structures across related conditions. DSNS combines neighbourhood selection with the Data Shared Lasso decomposition, representing each nodewise regression coefficient as the sum of a shared component and a sparse condition-specific deviation. This provides an interpretable decomposition of molecular associations while retaining the computational advantages of neighbourhood selection. We also adapt the Stability Approach to Regularisation Selection (StARS) to this two-parameter joint estimation setting. In simulations involving sparse, hub-based and rewiring perturbation mechanisms, DSNS matched the best joint estimation methods for two conditions and outperformed them as the number of conditions increased, while remaining substantially faster than graphical lasso frameworks. Applied to prediagnostic inflammatory proteomic data from future lung cancer cases and matched controls in the EPIC-Italy and NOWAC cohorts, DSNS highlighted altered associations involving CDCP1 and IL10, two established lung cancer risk markers, as well as differential associations involving proteins not selected by risk models.

stat.AP

Considering causality in the construction of molecular signatures of lifestyle exposures

Molecular signatures derived from omics data are increasingly used in epidemiological studies to characterize lifestyle exposures, either as proxies of exposure or to provide insight into disease mechanisms. These signatures are typically constructed by regressing the exposure on high-dimensional omics features. In the literature, an initial univariate screening step has sometimes been applied prior to multivariate modelling, but the causal implications of this choice have not yet been considered. Focusing on settings where the exposure causally influences molecular features (and not the reverse), we use directed acyclic graphs (DAGs) and $d$-separation arguments to show that collider bias may arise when the screening step is ignored, leading to the inclusion of non-causal features in the signature. We further demonstrate that the screening step can mitigate this bias. Our simulation studies illustrate that screening reduces the inclusion of non-causal features, albeit at the cost of lower sensitivity and reduced correlation between the exposure and the resulting signature. Overall, we recommend applying univariate screening prior to signature construction, particularly when the inclusion of non-causal features is undesirable, such as in mechanistic studies.

stat.ME

On the estimation of inclusion probabilities for weighted analyses of nested case control studies

Nested case-control (NCC) studies are a widely adopted design in epidemiology to investigate exposure-disease relationships. This paper examines weighted analyses in NCC studies, focusing on two prominent weighting methods: Kaplan-Meier (KM) weights and Generalized Additive Model (GAM) weights. We consider three target estimands: log-hazard ratios, conditional survival, and associations between exposures. While KM- and GAM-weights are generally robust, we identify specific scenarios where they can lead to biased estimates. We demonstrate that KM-weights can lead to biased estimates when a proportion of the originating cohort is effectively ineligible for NCC selection, particularly with small case proportions or numerous matching factors. Instead, GAM-weights can yield biased results if interactions between matching factors influence disease risk and are not adequately incorporated into weight calculation. Using Directed Acyclic Graphs (DAGs), we develop a framework to systematically determine which variables should be included in weight calculations. We show that the optimal set of variables depends on the target estimand and the causal relationships between matching factors, exposures, and disease risk. We illustrate our findings with both synthetic and real data from the European Prospective Investigation into Cancer and nutrition (EPIC) study. Additionally, we extend the application of GAM-weights to "untypical" NCC studies, where only a subset of cases are included. Our work provides crucial insights for conducting accurate and robust weighted analyses in NCC studies.

stat.ME

Optimal transport for automatic alignment of untargeted metabolomic data

Untargeted metabolomic profiling through liquid chromatography-mass spectrometry (LC-MS) measures a vast array of metabolites within biospecimens, advancing drug development, disease diagnosis, and risk prediction. However, the low throughput of LC-MS poses a major challenge for biomarker discovery, annotation, and experimental comparison, necessitating the merging of multiple datasets. Current data pooling methods encounter practical limitations due to their vulnerability to data variations and hyperparameter dependence. Here we introduce GromovMatcher, a flexible and user-friendly algorithm that automatically combines LC-MS datasets using optimal transport. By capitalizing on feature intensity correlation structures, GromovMatcher delivers superior alignment accuracy and robustness compared to existing approaches. This algorithm scales to thousands of features requiring minimal hyperparameter tuning. Manually curated datasets for validating alignment algorithms are limited in the field of untargeted metabolomics, and hence we develop a dataset split procedure to generate pairs of validation datasets to test the alignments produced by GromovMatcher and other methods. Applying our method to experimental patient studies of liver and pancreatic cancer, we discover shared metabolic features related to patient alcohol intake, demonstrating how GromovMatcher facilitates the search for biomarkers associated with lifestyle risk factors linked to several cancer types.

q-bio.QM

On some limitations of probabilistic models for dimension-reduction: Illustration in the case of probabilistic formulations of partial least squares

Partial Least Squares (PLS) refer to a class of dimension-reduction techniques aiming at the identification of two sets of components with maximal covariance, to model the relationship between two sets of observed variables $x\in\mathbb{R}^p$ and $y\in\mathbb{R}^q$, with $p\geq 1, q\geq 1$. Probabilistic formulations have recently been proposed for several versions of the PLS. Focusing first on the probabilistic formulation of the PLS-SVD proposed by el Bouhaddani et al., we establish that the constraints on their model parameters are too restrictive and define particular distributions for $(x,y)$, under which components with maximal covariance (solutions of PLS-SVD) are also necessarily of respective maximal variances (solutions of principal components analyses of $x$ and $y$, respectively). We propose an alternative probabilistic formulation of PLS-SVD, no longer restricted to these particular distributions. We then present numerical illustrations of the limitation of the original model of el Bouhaddani et al. We also briefly discuss similar limitations in another latent variable model for dimension-reduction.

stat.ME

Causal inference under over-simplified longitudinal causal models

Many causal models of interest in epidemiology involve longitudinal exposures, confounders and mediators. However, repeated measurements are not always available or used in practice, leading analysts to overlook the time-varying nature of exposures and work under over-simplified causal models. Our objective is to assess whether - and how - causal effects identified under such misspecified causal models relates to true causal effects of interest. We derive sufficient conditions ensuring that the quantities estimated in practice under over-simplified causal models can be expressed as weighted averages of longitudinal causal effects of interest. Unsurprisingly, these sufficient conditions are very restrictive, and our results state that the quantities estimated in practice should be interpreted with caution in general, as they usually do not relate to any longitudinal causal effect of interest. Our simulations further illustrate that the bias between the quantities estimated in practice and the weighted averages of longitudinal causal effects of interest can be substantial. Overall, our results confirm the need for repeated measurements to conduct proper analyses and/or the development of sensitivity analyses when they are not available.

stat.ME

Sparse estimation for case-control studies with multiple subtypes of cases

The analysis of case-control studies with several subtypes of cases is increasingly common, e.g. in cancer epidemiology. For matched designs, we show that a natural strategy is based on a stratified conditional logistic regression model. Then, to account for the potential homogeneity among the subtypes of cases, we adapt the ideas of data shared lasso, which has been recently proposed for the estimation of regression models in a stratified setting. For unmatched designs, we compare two standard methods based on L1-norm penalized multinomial logistic regression. We describe formal connections between these two approaches, from which practical guidance can be derived. We show that one of these approaches, which is based on a symmetric formulation of the multinomial logistic regression model, actually reduces to a data shared lasso version of the other. Consequently, the relative performance of the two approaches critically depends on the level of homogeneity that exists among the subtypes of cases: more precisely, when homogeneity is moderate to high, the non-symmetric formulation with controls as the reference is not recommended. Empirical results obtained from synthetic data are presented, which confirm the benefit of properly accounting for potential homogeneity under both matched and unmatched designs. We also present preliminary results from the analysis a case-control study nested within the EPIC cohort, where the objective is to identify metabolites associated with the occurrence of subtypes of breast cancer.

stat.ME

Which practical interventions does the do-operator refer to in causal inference? Illustration on the example of obesity and cancer

For exposures $X$ like obesity, no precise and unambiguous definition exists for the hypothetical intervention $do(X = x_0)$. This has raised concerns about the relevance of causal effects estimated from observational studies for such exposures. Under the framework of structural causal models, we study how the effect of $do(X = x_0)$ relates to the effect of interventions on causes of $X$. We show that for interventions focusing on causes of $X$ that affect the outcome through $X$ only, the effect of $do(X = x_0)$ equals the effect of the considered intervention. On the other hand, for interventions on causes $W$ of $X$ that affect the outcome not only through $X$, we show that the effect of $do(X = x_0)$ only partly captures the effect of the intervention. In particular, under simple causal models (e.g., linear models with no interaction), the effect of $do(X = x_0)$ can be seen as an indirect effect of the intervention on $W$.

stat.ME

Causal inference to detect selection bias in road safety epidemiology

In the field of road safety, it is common to use responsibility analyses to assess the effect of a given factor on the risk of being responsible for an accident, among drivers involved in an accident only. Even if this design is now widely adopted in the field, the question of selection bias is often raised. The structural Causal Model framework now provides valuable tools to assess causal effects from observational data and identify selection bias. In this article, we briefly review recent results regarding the recoverability of causal effects from selection biased data, and apply them to the case of responsibility analyses. Our objective is to formally determine whether causal effects can be unbiasedly estimated through this type of analyses, when available data are restricted to severe accidents, as it is commonly the case in practice. However, because speed has a direct effect on the severity of the accident, we show that causal odds-ratios are not estimable from responsibility analyses. We present numerical results to illustrate our argument, the magnitude of the bias and to discuss recent results from real data.

stat.ME

Structure estimation of binary graphical models on stratified data: application to the description of injury tables for victims of road accidents

Graphical models are used in many applications such as medical diagnostic, computer security, etc. More and more often, the estimation of such models has to be performed on several predefined strata of the whole population. For instance, in epidemiology and clinical research, strata are often defined according to age, gender, treatment or disease type, etc. In this article, we propose new approaches aimed at estimating binary graphical models on such strata. Our approaches are obtained by combining well-known methods when estimating one single binary graphical model, with penalties encouraging structured sparsity, and which have recently been shown appropriate when dealing with stratified data. Empirical comparions on synthetic data highlight that our approaches generally outperform the competitors we considered. An application is provided where we study associations among injuries suffered by victims of road accidents according to road user type.

stat.ME

A SAEM Algorithm for Fused Lasso Penalized Non Linear Mixed Effect Models: Application to Group Comparison in Pharmacokinetic

Non linear mixed effect models are classical tools to analyze non linear longitudinal data in many fields such as population Pharmacokinetic. Groups of observations are usually compared by introducing the group affiliations as binary covariates with a reference group that is stated among the groups. This approach is relatively limited as it allows only the comparison of the reference group to the others. In this work, we propose to compare the groups using a penalized likelihood approach. Groups are described by the same structural model but with parameters that are group specific. The likelihood is penalized with a fused lasso penalty that induces sparsity on the differences between groups for both fixed effects and variances of random effects. A penalized Stochastic Approximation EM algorithm is proposed that is coupled to Alternating Direction Method Multipliers to solve the maximization step. An extensive simulation study illustrates the performance of this algorithm when comparing more than two groups. Then the approach is applied to real data from two pharmacokinetic drug-drug interaction trials.

stat.CO

Can collider bias fully explain the obesity paradox?

The "obesity paradox" has been reported in several observational studies, where obesity was shown to be associated to a decreased mortality in individuals suffering from a chronic disease, such as diabetes or heart failure. Causal arguments have recently been given to explain this apparently paradoxical fact: because the chronic disease is caused by obesity, the observed "protective effect" of obesity among patients with, say, diabetes, actually has no causal value. Recently, Sperrin et al. (2016) relaunched the debate and claimed that the resulting bias, the so-called collider bias, was unlikely to be the main explanation for the obesity paradox. However, a number of issues in their work make their conclusions questionable. In this article, we first study the bias between (i) the association between obesity and early death among patients suffering from the chronic disease $Δ_{AS}$ and (ii) the causal effect considered by Sperrin et al. Under the usual framework of structural causal models, we explain why this bias can be much higher than what these authors reported. We further consider alternative causal effects of potential interest and study their difference with $Δ_{AS}$. Numerical examples are presented to illustrate the magnitude of these differences under realistic scenarios. We show that it is possible to have a negative $Δ_{AS}$, while the causal effects we considered are all positive. Therefore, even under the very simple generative model we considered, collider bias can be the sole cause of the obesity paradox.

stat.ME

Regression modeling on stratified data with the lasso

We consider the estimation of regression models on strata defined using a categorical covariate, in order to identify interactions between this categorical covariate and the other predictors. A basic approach requires the choice of a reference stratum. We show that the performance of a penalized version of this approach depends on this arbitrary choice. We propose a refined approach that bypasses this arbitrary choice, at almost no additional computational cost. Regarding model selection consistency, our proposal mimics the strategy based on an optimal and covariate-specific choice for the reference stratum. Results from an empirical study confirm that our proposal generally outperforms the basic approach in the identification and description of the interactions. An illustration on gene expression data is provided.

math.ST

Joint estimation of $K$ related regression models with simple $L_1$-norm penalties

We propose a new approach, along with refinements, based on $L_1$ penalties and aimed at jointly estimating several related regression models. Its main interest is that it can be rewritten as a weighted lasso on a simple transformation of the original data set. In particular, it does not need new dedicated algorithms and is ready to implement under a variety of regression models, {\em e.g.}, using standard R packages. Moreover, asymptotic oracle properties are derived along with preliminary non-asymptotic results, suggesting good theoretical properties. Our approach is further compared with state-of-the-art competitors under various settings on synthetic data: these empirical results confirm that our approach performs at least similarly to its competitors. As a final illustration, an analysis of road safety data is provided.

stat.ME

Time-dependent AUC with right-censored data: a survey study

The ROC curve and the corresponding AUC are popular tools for the evaluation of diagnostic tests. They have been recently extended to assess prognostic markers and predictive models. However, due to the many particularities of time-to-event outcomes, various definitions and estimators have been proposed in the literature. This review article aims at presenting the ones that accommodate to right-censoring, which is common when evaluating such prognostic markers.

stat.ME

Safe Feature Elimination in Sparse Supervised Learning

We investigate fast methods that allow to quickly eliminate variables (features) in supervised learning problems involving a convex loss function and a $l_1$-norm penalty, leading to a potentially substantial reduction in the number of variables prior to running the supervised learning algorithm. The methods are not heuristic: they only eliminate features that are {\em guaranteed} to be absent after solving the learning problem. Our framework applies to a large class of problems, including support vector machine classification, logistic regression and least-squares. The complexity of the feature elimination step is negligible compared to the typical computational effort involved in the sparse supervised learning problem: it grows linearly with the number of features times the number of examples, with much better count if data is sparse. We apply our method to data sets arising in text classification and observe a dramatic reduction of the dimensionality, hence in computational effort required to solve the learning problem, especially when very sparse classifiers are sought. Our method allows to immediately extend the scope of existing algorithms, allowing us to run them on data sets of sizes that were out of their reach before.

cs.LG

Safe Feature Elimination for the LASSO and Sparse Supervised Learning Problems

We describe a fast method to eliminate features (variables) in l1 -penalized least-square regression (or LASSO) problems. The elimination of features leads to a potentially substantial reduction in running time, specially for large values of the penalty parameter. Our method is not heuristic: it only eliminates features that are guaranteed to be absent after solving the LASSO problem. The feature elimination step is easy to parallelize and can test each feature for elimination independently. Moreover, the computational effort of our method is negligible compared to that of solving the LASSO problem - roughly it is the same as single gradient step. Our method extends the scope of existing LASSO algorithms to treat larger data sets, previously out of their reach. We show how our method can be extended to general l1 -penalized convex problems and present preliminary results for the Sparse Support Vector Machine and Logistic Regression problems.

cs.LG

An empirical comparative study of approximate methods for binary graphical models; application to the search of associations among causes of death in French death certificates

Looking for associations among multiple variables is a topical issue in statistics due to the increasing amount of data encountered in biology, medicine and many other domains involving statistical applications. Graphical models have recently gained popularity for this purpose in the statistical literature. Following the ideas of the LASSO procedure designed for the linear regression framework, recent developments dealing with graphical model selection have been based on $\ell_1$-penalization. In the binary case, however, exact inference is generally very slow or even intractable because of the form of the so-called log-partition function. Various approximate methods have recently been proposed in the literature and the main objective of this paper is to compare them. Through an extensive simulation study, we show that a simple modification of a method relying on a Gaussian approximation achieves good performance and is very fast. We present a real application in which we search for associations among causes of death recorded on French death certificates.

stat.ML