SearcharxivSearch

arXiv subjects

Anthony CC Coolen

Publications and source records attributed to Anthony CC Coolen.

14 recordsLinked to original sources

An Accurate and Single-Communication Federated Inference Algorithm

Joint analyses across multiple institutions are increasingly important in biomedical and epidemiological research, particularly for rare diseases where datasets are typical small. However, privacy regulations and institutional policies often prevent the sharing of individual-level patient data. In this paper we present an accurate and single-communication federated inference algorithm. Single-communication federated inference enables statistical analyses through a single exchange of summary statistics between participating centers and a coordinating server, preserving privacy while reducing communication and computational costs compared with iterative federated learning. We extend a recently proposed single-communication federated inference strategy that is based on second-order Taylor expansions by using third-order expansions to better approximate local log-likelihood functions. The proposed method is evaluated through simulation studies based on real data and compared with existing federated inference strategies. The simulation studies assess the performance of the proposed method, with a particular focus on scenarios involving small local sample sizes, where quadratic approximations may fail to capture skewness and other higher-order characteristics of the log-likelihood function. They demonstrate that incorporating higher-order information of the log-likelihood function improves the accuracy while preserving the privacy, communication efficiency, and scalability required for collaborative biomedical and epidemiological research.

stat.ME

Bayesian Federated Inference for regression models based on non-shared multicenter data sets from heterogeneous populations

To estimate accurately the parameters of a regression model, the sample size must be large enough relative to the number of possible predictors for the model. In practice, sufficient data is often lacking, which can lead to overfitting of the model and, as a consequence, unreliable predictions of the outcome of new patients. Pooling data from different data sets collected in different (medical) centers would alleviate this problem, but is often not feasible due to privacy regulation or logistic problems. An alternative route would be to analyze the local data in the centers separately and combine the statistical inference results with the Bayesian Federated Inference (BFI) methodology. The aim of this approach is to compute from the inference results in separate centers what would have been found if the statistical analysis was performed on the combined data. We explain the methodology under homogeneity and heterogeneity across the populations in the separate centers, and give real life examples for better understanding. Excellent performance of the proposed methodology is shown. An R-package to do all the calculations has been developed and is illustrated in this paper. The mathematical details are given in the Appendix.

stat.AP

Bayesian Federated Inference for Survival Models

In cancer research, overall survival and progression free survival are often analyzed with the Cox model. To estimate accurately the parameters in the model, sufficient data and, more importantly, sufficient events need to be observed. In practice, this is often a problem. Merging data sets from different medical centers may help, but this is not always possible due to strict privacy legislation and logistic difficulties. Recently, the Bayesian Federated Inference (BFI) strategy for generalized linear models was proposed. With this strategy the statistical analyses are performed in the local centers where the data were collected (or stored) and only the inference results are combined to a single estimated model; merging data is not necessary. The BFI methodology aims to compute from the separate inference results in the local centers what would have been obtained if the analysis had been based on the merged data sets. In this paper we generalize the BFI methodology as initially developed for generalized linear models to survival models. Simulation studies and real data analyses show excellent performance; i.e., the results obtained with the BFI methodology are very similar to the results obtained by analyzing the merged data. An R package for doing the analyses is available.

stat.ME

Bayesian Federated Inference for estimating Statistical Models based on Non-shared Multicenter Data sets

Identifying predictive factors for an outcome of interest via a multivariable analysis is often difficult when the data set is small. Combining data from different medical centers into a single (larger) database would alleviate this problem, but is in practice challenging due to regulatory and logistic problems. Federated Learning (FL) is a machine learning approach that aims to construct from local inferences in separate data centers what would have been inferred had the data sets been merged. It seeks to harvest the statistical power of larger data sets without actually creating them. The FL strategy is not always efficient and precise. Therefore, in this paper we refine and implement an alternative Bayesian Federated Inference (BFI) framework for multicenter data with the same aim as FL. The BFI framework is designed to cope with small data sets by inferring locally not only the optimal parameter values, but also additional features of the posterior parameter distribution, capturing information beyond what is used in FL. BFI has the additional benefit that a single inference cycle across the centers is sufficient, whereas FL needs multiple cycles. We quantify the performance of the proposed methodology on simulated and real life data.

stat.AP

Exact results on high-dimensional linear regression via statistical physics

It is clear that conventional statistical inference protocols need to be revised to deal correctly with the high-dimensional data that are now common. Most recent studies aimed at achieving this revision rely on powerful approximation techniques, that call for rigorous results against which they can be tested. In this context, the simplest case of high-dimensional linear regression has acquired significant new relevance and attention. In this paper we use the statistical physics perspective on inference to derive a number of new exact results for linear regression in the high-dimensional regime.

math.ST

Transitions in loopy random graphs with fixed degrees and arbitrary degree distributions

We analyze maximum entropy random graph ensembles with constrained degrees, drawn from arbitrary degree distributions, and a tuneable number of 3-loops (triangles). We find that such ensembles generally exhibit two transitions, a clustering and a shattering transition, separating three distinct regimes. At the clustering transition, the graphs change from typically having only isolated loops to forming loop clusters. At the shattering transition the graphs break up into extensively many small cliques to achieve the desired loop density. The locations of both transitions depend nontrivially on the system size. We derive a general formula for the loop density in the regime of isolated loops, for graphs with degree distributions that have finite and second moments. For bounded degree distributions we present further analytical results on loop densities and phase transition locations, which, while non-rigorous, are all validated via MCMC sampling simulations. We show that the shattering transition is of an entropic nature, occurring for all loop density values, provided the system is large enough.

cond-mat.dis-nn

Replica analysis of Bayesian data clustering

We use statistical mechanics to study model-based Bayesian data clustering. In this approach, each partition of the data into clusters is regarded as a microscopic system state, the negative data log-likelihood gives the energy of each state, and the data set realisation acts as disorder. Optimal clustering corresponds to the ground state of the system, and is hence obtained from the free energy via a low `temperature' limit. We assume that for large sample sizes the free energy density is self-averaging, and we use the replica method to compute the asymptotic free energy density. The main order parameter in the resulting (replica symmetric) theory, the distribution of the data over the clusters, satisfies a self-consistent equation which can be solved by a population dynamics algorithm. From this order parameter one computes the average free energy, and all relevant macroscopic characteristics of the problem. The theory describes numerical experiments perfectly, and gives a significant improvement over the mean-field theory that was used to study this model in past.

cond-mat.dis-nn

Imaginary replica analysis of loopy regular random graphs

We present an analytical approach for describing spectrally constrained maximum entropy ensembles of finitely connected regular loopy graphs, valid in the regime of weak loop-loop interactions. We derive an expression for the leading two orders of the expected eigenvalue spectrum, through the use of infinitely many replica indices taking imaginary values. We apply the method to models in which the spectral constraint reduces to a soft constraint on the number of triangles, which exhibit `shattering' transitions to phases with extensively many disconnected cliques, to models with controlled numbers of triangles and squares, and to models where the spectral constraint reduces to a count of the number of adjacency matrix eigenvalues in a given interval. Our predictions are supported by MCMC simulations based on edge swaps with nontrivial acceptance probabilities.

cond-mat.dis-nn

Mean-field theory of Bayesian clustering

We show that model-based Bayesian clustering, the probabilistically most systematic approach to the partitioning of data, can be mapped into a statistical physics problem for a gas of particles, and as a result becomes amenable to a detailed quantitative analysis. A central role in the resulting statistical physics framework is played by an entropy function. We demonstrate that there is a relevant parameter regime where mean-field analysis of this function is exact, and that, under natural assumptions, the lowest entropy state of the hypothetical gas corresponds to the optimal clustering of data. The byproduct of our analysis is a simple but effective clustering algorithm, which infers both the most plausible number of clusters in the data and the corresponding partitions. Describing Bayesian clustering in statistical mechanical terms is found to be natural and surprisingly effective.

cond-mat.dis-nn

Exactly Solvable Random Graph Ensemble with Extensively Many Short Cycles

We introduce and analyse ensembles of 2-regular random graphs with a tuneable distribution of short cycles. The phenomenology of these graphs depends critically on the scaling of the ensembles' control parameters relative to the number of nodes. A phase diagram is presented, showing a second order phase transition from a connected to a disconnected phase. We study both the canonical formulation, where the size is large but fixed, and the grand canonical formulation, where the size is sampled from a discrete distribution, and show their equivalence in the thermodynamical limit. We also compute analytically the spectral density, which consists of a discrete set of isolated eigenvalues, representing short cycles, and a continuous part, representing cycles of diverging size.

cond-mat.dis-nn

Statistical mechanics of clonal expansion in lymphocyte networks modelled with slow and fast variables

We study the Langevin dynamics of the adaptive immune system, modelled by a lymphocyte network in which the B cells are interacting with the T cells and antigen. We assume that B clones and T clones are evolving in different thermal noise environments and on different timescales. We derive stationary distributions and use statistical mechanics to study clonal expansion of B clones in this model when the B and T clone sizes are assumed to be the slow and fast variables respectively and vice versa. We derive distributions of B clone sizes and use general properties of ferromagnetic systems to predict characteristics of these distributions, such as the average B cell concentration, in some regimes where T cells can be modelled as binary variables. This analysis is independent of network topologies and its results are qualitatively consistent with experimental observations. In order to obtain full distributions we assume that the network topologies are random and locally equivalent to trees. The latter allows us to employ the Bethe-Peierls approach and to develop a theoretical framework which can be used to predict the distributions of B clone sizes. As an example we use this theory to compute distributions for the models of immune system defined on random regular networks.

cond-mat.dis-nn

Spin systems on hypercubic Bethe lattices: A Bethe-Peierls approach

We study spin systems on Bethe lattices constructed from d-dimensional hypercubes. Although these lattices are not tree-like, and therefore closer to real cubic lattices than Bethe lattices or regular random graphs, one can still use the Bethe-Peierls method to derive exact equations for the magnetization and other thermodynamic quantities. We compute phase diagrams for ferromagnetic Ising models on hypercubic Bethe lattices with dimension d=2, 3, and 4. Our results are in good agreement with the results of the same models on d-dimensional cubic lattices, for low and high temperatures, and offer an improvement over the conventional Bethe lattice with connectivity k=2d.

cond-mat.dis-nn

Generic solution of the heterogeneity-induced competing risk problem in survival analysis

Most papers implicitly assume competing risks to be induced by residual cohort heterogeneity, i.e. heterogeneity that is not captured by the recorded covariates. Based on this observation we develop a generic statistical description of competing risks that unifies the main schools of thought. Assuming heterogeneity-induced competing risks is much weaker than assuming risk independence. However, we show that it still imposes sufficient constraints to solve the competing risk problem, and derive exact formulae for decontaminated primary risk hazard rates and cause-specific survival functions. The canonical description is in terms of a cohort's covariate-constrained functional distribution of individual hazard rates of all risks. Assuming proportional hazards at the level of individuals leads to a natural parametrisation of this distribution, from which Cox regression, frailty and random effects models, and latent class models can all be recovered in special limits, and which also generates parametrised cumulative incidence functions (the language of Fine and Gray). We demonstrate with synthetic data how the generic method can uncover and map a cohort's substructure, if such substructure exists, and remove heterogeneity-induced false protectivity and false exposure effects. Application to real survival data from the ULSAM study, with prostate cancer as the primary risk, is found to give plausible alternative explanations for previous counter-intuitive inferences.

stat.AP

Transfer operator analysis of the parallel dynamics of disordered Ising chains

We study the synchronous stochastic dynamics of the random field and random bond Ising chain. For this model the generating functional analysis methods of De Dominicis leads to a formalism with transfer operators, similar to transfer matrices in equilibrium studies, but with dynamical paths of spins and (conjugate) fields as arguments, as opposed to replicated spins. In the thermodynamic limit the macroscopic dynamics is captured by the dominant eigenspace of the transfer operator, leading to a relative simple and transparent set of equations that are easy to solve numerically. Our results are supported excellently by numerical simulations.

cond-mat.dis-nn