SearcharxivSearch

arXiv subjects

Tom Marshall

Publications and source records attributed to Tom Marshall.

3 recordsLinked to original sources

Variational Consensus Monte Carlo for Bayesian Mixture

Motivated by the privacy, sensitivity and sharing limitations of health data, we present a comprehensive pipeline for inference of Bayesian mixture models within a federated learning setting, i.e. when data cannot be fully shared or pooled across compute nodes. We adopt a Consensus Monte Carlo (CMC) approach, in which an MCMC algorithm is run independently within each data silo to estimate local posterior distributions, which are then aggregated to approximate the posterior over the full data. The variational CMC approach of Rabinovich, Angelino and Jordan (2015) [1] frames the aggregation step as a variational inference problem, but their application to mixtures assumes the number of clusters and key mixture parameters to be known. Our main methodological contributions are: (i) an extension of variational CMC to over-fitted Bayesian mixture models that infer the number of clusters and all model parameters, without requiring conjugacy; (ii) novel cluster-matching algorithms suitable for cross-silo settings in which not every cluster appears in each local dataset; (iii) a number of inference strategies for the aggregation step, matched to different federated learning constraints; and (iv) guidelines for choosing among these in practice. A comprehensive simulation study validates the framework and allows us to compare to state-of-the-art federated learning alternatives. Notably, we show that when the composition of local datasets reflects the underlying clustering structure in the data, our approach can recover small clusters with greater accuracy than standard MCMC applied to the pooled data. We illustrate the framework on large-scale electronic health record data, identifying multi-morbidity patterns in a British geriatric population.

stat.ML

Federated Variational Inference for Bayesian Mixture Models

We present a federated learning approach for Bayesian model-based clustering of large-scale binary and categorical datasets. We introduce a principled 'divide and conquer' inference procedure using variational inference with local merge and delete moves within batches of the data in parallel, followed by 'global' merge moves across batches to find global clustering structures. We show that these merge moves require only summaries of the data in each batch, enabling federated learning across local nodes without requiring the full dataset to be shared. Empirical results on simulated and benchmark datasets demonstrate that our method performs well in comparison to existing clustering algorithms. We validate the practical utility of the method by applying it to large scale electronic health record (EHR) data.

stat.ML

Risk Fluctuation Characteristics of Internet Finance: Combining Industry Characteristics with Ecological Value

The Internet plays a key role in society and is vital to economic development. Due to the pressure of competition, most technology companies, including Internet finance companies, continue to explore new markets and new business. Funding subsidies and resource inputs have led to significant business income tendencies in financial statements. This tendency of business income is often manifested as part of the business loss or long-term unprofitability. We propose a risk change indicator (RFR) and compare the risk indicator of fourteen representative companies. This model combines extreme risk value with slope, and the combination method is simple and effective. The results of experiment show the potential of this model. The risk volatility of technology enterprises including Internet finance enterprises is highly cyclical, and the risk volatility of emerging Internet fintech companies is much higher than that of other technology companies.

econ.EM