SearcharxivSearch

arXiv subjects

Saverio Ranciati

Publications and source records attributed to Saverio Ranciati.

8 recordsLinked to original sources

Dynamic prediction intervals for survival times

Most work on survival prediction focuses on estimating survival probabilities rather than predicting individual event times. Recent conformal methods have made it possible to construct prediction intervals for survival times with right-censored outcomes, but existing approaches are restricted to settings with covariates only measured at baseline and do not address dynamic prediction with longitudinal data. We study prediction intervals for survival times in a dynamic prediction framework with longitudinal covariates. Our approach uses Penalized Regression Calibration (PRC) as a working dynamic prediction model, combining linear mixed models for the longitudinal histories with a Cox model for post-landmark survival, and then applies a conformal calibration step to obtain prediction intervals. We compare naive intervals obtained by direct inversion of the survival function estimated by PRC to our dynamic conformal method. A Monte Carlo simulation study evaluates empirical coverage and interval length across sample sizes, censoring levels, landmark times, and non-proportional hazards (NPH) scenarios. We illustrate the proposed methodology by computing dynamic prediction intervals for the time until a dementia diagnosis in the ADNI dataset. The results show that naive inversion is often unreliable, whereas the proposed dynamic conformal method yields more stable predictive performance.

stat.ME

On the application of Gaussian graphical models to paired data problems

Gaussian graphical models are nowadays commonly applied to the comparison of groups sharing the same variables, by jointy learning their independence structures. We consider the case where there are exactly two dependent groups and the association structure is represented by a family of coloured Gaussian graphical models suited to deal with paired data problems. To learn the two dependent graphs, together with their across-graph association structure, we implement a fused graphical lasso penalty. We carry out a comprehensive analysis of this approach, with special attention to the role played by some relevant submodel classes. In this way, we provide a broad set of tools for the application of Gaussian graphical models to paired data problems. These include results useful for the specification of penalty values in order to obtain a path of lasso solutions and an ADMM algorithm that solves the fused graphical lasso optimization problem. Finally, we present an application of our method to cancer genomics where it is of interest to compare cancer cells with a control sample from histologically normal tissues adjacent to the tumor. All the methods described in this article are implemented in the $\texttt{R}$ package $\texttt{pdglasso}$ availabe at: https://github.com/savranciati/pdglasso.

stat.ME

Identifying Brexit voting patterns in the British House of Commons: an analysis based on Bayesian mixture models with flexible concomitant covariate effects

Brexit and its implications are an ongoing topic of interest since the Brexit referendum in 2016. In 2019 the House of commons held a number of "indicative" and "meaningful" votes as part of the Brexit approval process. The voting behaviour of members of the parliament in these votes is investigated to gain insight into the Brexit approval process. In particular, a mixture model with concomitant covariates is developed to identify groups of members of parliament who share similar voting behaviour while also considering characteristics of the members of parliament. The novelty of the method lies in the flexible structure used to model the effect of concomitant covariates on the component weights of the mixture, with the (potentially nonlinear) terms represented as a smooth function of the covariates. Results show this approach allows to quantify the effect of the age of members of parliament, as well as preferences and competitiveness in the constituencies they represent, on their position towards Brexit. This helps grouping the aforementioned politicians into homogeous clusters, whose composition departs sensibly from that of the parties.

stat.ME

Fused graphical lasso for brain networks with symmetries

Neuroimaging is the growing area of neuroscience devoted to produce data with the goal of capturing processes and dynamics of the human brain. We consider the problem of inferring the brain connectivity network from time dependent functional magnetic resonance imaging (fMRI) scans. To this aim we propose the symmetric graphical lasso, a penalized likelihood method with a fused type penalty function that takes into explicit account the natural symmetrical structure of the brain. Symmetric graphical lasso allows one to learn simultaneously both the network structure and a set of symmetries across the two hemispheres. We implement an alternating directions method of multipliers algorithm to solve the corresponding convex optimization problem. Furthermore, we apply our methods to estimate the brain networks of two subjects, one healthy and the other affected by a mental disorder, and to compare them with respect to their symmetric structure. The method applies once the temporal dependence characterising fMRI data has been accounted for and we compare the impact on the analysis of different detrending techniques on the estimated brain networks.

stat.ME

Identifying overlapping terrorist cells from the Noordin Top actor-event network

Actor-event data are common in sociological settings, whereby one registers the pattern of attendance of a group of social actors to a number of events. We focus on 79 members of the Noordin Top terrorist network, who were monitored attending 45 events. The attendance or non-attendance of the terrorist to events defines the social fabric, such as group coherence and social communities. The aim of the analysis of such data is to learn about the affiliation structure. Actor-event data is often transformed to actor-actor data in order to be further analysed by network models, such as stochastic block models. This transformation and such analyses lead to a natural loss of information, particularly when one is interested in identifying, possibly overlapping, subgroups or communities of actors on the basis of their attendances to events. In this paper we propose an actor-event model for overlapping communities of terrorists, which simplifies interpretation of the network. We propose a mixture model with overlapping clusters for the analysis of the binary actor-event network data, called {\tt manet}, and develop a Bayesian procedure for inference. After a simulation study, we show how this analysis of the terrorist network has clear interpretative advantages over the more traditional approaches of affiliation network analysis.

stat.AP

Mixtures of multivariate generalized linear models with overlapping clusters

With the advent of ubiquitous monitoring and measurement protocols, studies have started to focus more and more on complex, multivariate and heterogeneous datasets. In such studies, multivariate response variables are drawn from a heterogeneous population often in the presence of additional covariate information. In order to deal with this intrinsic heterogeneity, regression analyses have to be clustered for different groups of units. Up until now, mixture model approaches assigned units to distinct and non-overlapping groups. However, not rarely these units exhibit more complex organization and clustering. It is our aim to define a mixture of generalized linear models with overlapping clusters of units. This involves crucially an overlap function, that maps the coefficients of the parent clusters into the the coefficient of the multiple allocation units. We present a computationally efficient MCMC scheme that samples the posterior distribution of the parameters in the model. An example on a two-mode network study shows details of the implementation in the case of a multivariate probit regression setting. A simulation study shows the overall performance of the method, whereas an illustration of the voting behaviour on the US supreme court shows how the 9 justices split in two overlapping sets of justices.

stat.ME

Bayesian Smooth-and-Match strategy for ordinary differential equations models that are linear in the parameters

In many fields of application, dynamic processes that evolve through time are well described by systems of ordinary differential equations (ODEs). The analytical solution of the ODEs is often not available and different methods have been proposed to infer these quantities: from numerical optimization to regularized (penalized) models, these procedures aim to estimate indirectly the parameters without solving the system. We focus on the class of techniques that use smoothing to avoid direct integration and, in particular, on a Bayesian Smooth-and-Match strategy that allows to obtain the ODEs' solution while performing inference on models that are linear in the parameters. We incorporate in the strategy two main sources of uncertainty: the noise level in the measurements and the model error. We assess the performance of the proposed approach in three different simulation studies and we compare the results on a dataset on neuron electrical activity.

stat.ME

Mixture model with multiple allocations for clustering spatially correlated observations in the analysis of ChIP-Seq data

Model-based clustering is a technique widely used to group a collection of units into mutually exclusive groups. There are, however, situations in which an observation could in principle belong to more than one cluster. In the context of Next-Generation Sequencing (NGS) experiments, for example, the signal observed in the data might be produced by two (or more) different biological processes operating together and a gene could participate in both (or all) of them. We propose a novel approach to cluster NGS discrete data, coming from a ChIP-Seq experiment, with a mixture model, allowing each unit to belong potentially to more than one group: these multiple allocation clusters can be flexibly defined via a function combining the features of the original groups without introducing new parameters. The formulation naturally gives rise to a `zero-inflation group' in which values close to zero can be allocated, acting as a correction for the abundance of zeros that manifest in this type of data. We take into account the spatial dependency between observations, which is described through a latent Conditional Auto-Regressive process that can reflect different dependency patterns. We assess the performance of our model within a simulation environment and then we apply it to ChIP-seq real data.

stat.AP