SearcharxivSearch

arXiv subjects

Cecilia Balocchi

Publications and source records attributed to Cecilia Balocchi.

10 recordsLinked to original sources

Scalable Bayesian Spatial Mixture Modelling for Remote Sensing Image Segmentation

Accurate and scalable land cover classification is essential for global conservation monitoring and policy-making. While remote sensing images provide a cost-effective alternative to ground surveys, current methods often lack principled uncertainty quantification and require substantial labelled data, limiting their usability and reliability in new regions with distribution shifts. We propose a Bayesian spatial mixture modelling approach for image segmentation, extending the classical Potts model by allowing for a generalised spatial dependence structure and incorporating informative priors estimated from pre-existing labelled data. Our framework, called POTTERS (Potts Model for Enhanced Remote Sensing), enables robust uncertainty quantification, accounts for class interactions, and can detect new clusters in the target region of interest. Crucially, our model does not require labelled data from the target region; instead, it incorporates prior information about the labels from pre-existing externally labelled images. To ensure scalability to large remote sensing images, we develop an efficient variational inference algorithm for posterior approximation. We demonstrate the benefits of our approach in simulation studies and apply it to land cover classification in a case study in Scotland, leveraging publicly available remote sensing data from England.

stat.ME

Multiomics Tissue Segmentation via Spatially-Informed Nested Biclustering Methods

Matrix-Assisted Laser Desorption/Ionisation Mass Spectrometry Imaging (MSI) is a powerful technique for spatially resolved molecular profiling and cancer biomarker discovery. Recent advances, including a novel multiomics workflow, enable multiple rounds of MSI on the same tissue section, extracting diverse molecular classes, e.g., lipids, peptides, and N-glycans, while preserving spatial resolution. This innovation is particularly valuable for studies with limited tissue sample availability, such as rare diseases or small tumors. However, the resulting data are high-dimensional, spatially structured, and the various molecular types share the same pixel grid. To address these challenges, we propose Poseidon, a Bayesian nonparametric nested biclustering model that simultaneously segments the common tissue pixels and clusters molecular signals within each molecular class, leveraging the shared spatial structure. A separately exchangeable framework is first considered, and then extended to handle spatial data via hidden Markov random fields. For scalability, we implement an efficient mean-field variational inference algorithm tailored for multi-dataset analysis. After validating the efficacy of our method on simulated scenarios, we demonstrate the applicability of our model in a real-world case study, where multiomics measurements were performed on kidney tissue affected by clear cell renal cell carcinoma. The nested, hierarchical structure of Poseidon, combined with its principled inferential framework, allows the extraction of interesting biological insights, such as clear tissue segmentation and biomarker detection.

stat.ME

Understanding uncertainty in Bayesian cluster analysis

The Bayesian approach to clustering is often appreciated for its ability to provide uncertainty in the partition structure. However, summarizing the posterior distribution over the clustering structure can be challenging, due the discrete, unordered nature and massive dimension of the space. While recent advancements provide a single clustering estimate to represent the posterior, this ignores uncertainty and may even be unrepresentative in instances where the posterior is multimodal. To enhance our understanding of uncertainty, we propose a WASserstein Approximation for Bayesian clusterIng (WASABI), which summarizes the posterior samples with not one, but multiple clustering estimates, each corresponding to a different part of the partition space that receives substantial posterior mass. Specifically, we find such clustering estimates by approximating the posterior distribution in a Wasserstein distance sense, equipped with a suitable metric on the partition space. An interesting byproduct is that a locally optimal solution can be found using a k-medoids-like algorithm on the partition space to divide the posterior samples into groups, each represented by one of the clustering estimates. Using synthetic and real datasets, we show that WASABI helps to improve the understanding of uncertainty, particularly when clusters are not well separated or when the employed model is misspecified.

stat.CO

VCBART: Bayesian trees for varying coefficients

The linear varying coefficient models posits a linear relationship between an outcome and covariates in which the covariate effects are modeled as functions of additional effect modifiers. Despite a long history of study and use in statistics and econometrics, state-of-the-art varying coefficient modeling methods cannot accommodate multivariate effect modifiers without imposing restrictive functional form assumptions or involving computationally intensive hyperparameter tuning. In response, we introduce VCBART, which flexibly estimates the covariate effect in a varying coefficient model using Bayesian Additive Regression Trees. With simple default settings, VCBART outperforms existing varying coefficient methods in terms of covariate effect estimation, uncertainty quantification, and outcome prediction. We illustrate the utility of VCBART with two case studies: one examining how the association between later-life cognition and measures of socioeconomic position vary with respect to age and socio-demographics and another estimating how temporal trends in urban crime vary at the neighborhood level. An R package implementing VCBART is available at https://github.com/skdeshpande91/VCBART

stat.ME

Quantifying patient and neighborhood risks for stillbirth and preterm birth in Philadelphia with a Bayesian spatial model

Stillbirth and preterm birth are major public health challenges. Using a Bayesian spatial model, we quantified patient-specific and neighborhood risks of stillbirth and preterm birth in the city of Philadelphia. We linked birth data from electronic health records at Penn Medicine hospitals from 2010 to 2017 with census-tract-level data from the United States Census Bureau. We found that both patient-level characteristics (e.g. self-identified race/ethnicity) and neighborhood-level characteristics (e.g. violent crime) were significantly associated with patients' risk of stillbirth or preterm birth. Our neighborhood analysis found that higher-risk census tracts had 2.68 times the average risk of stillbirth and 2.01 times the average risk of preterm birth compared to lower-risk census tracts. Higher neighborhood rates of women in poverty or on public assistance were significantly associated with greater neighborhood risk for these outcomes, whereas higher neighborhood rates of college-educated women or women in the labor force were significantly associated with lower risk. Several of these neighborhood associations were missed by the patient-level analysis. These results suggest that neighborhood-level analyses of adverse pregnancy outcomes can reveal nuanced relationships and, thus, should be considered by epidemiologists. Our findings can potentially guide place-based public health interventions to reduce stillbirth and preterm birth rates.

stat.AP

Bayesian Nonparametric Inference for "Species-sampling" Problems

Given an observed sample from a population of individuals belonging to species, "species-sampling" problems (SSPs) call for estimating some features of the unknown species composition of additional unobservable samples from the same population. Within SSPs, the problems of estimating coverage probabilities, the number of unseen species and coverages of prevalences have emerged in the past three decades for being the subject of numerous methodological and applied works, mostly in biological sciences but also in statistical machine learning, electrical engineering, theoretical computer science, information theory and forensic statistics. In this paper, we focus on these popular SSPs, and present an overview of their Bayesian nonparametric (BNP) analysis under the Pitman--Yor process (PYP) prior. While reviewing the literature, we improve on computation and interpretability of existing posterior inferences, typically expressed through complicated combinatorial numbers, by establishing novel posterior representations in terms of simple compound Binomial and Hypergeometric distributions. We also consider the problem of estimating the discount and scale parameters of the PYP prior, showing a property of Bayesian consistency with respect to estimation through the hierarchical Bayes and empirical Bayes approaches, that is: the discount parameter can be estimated consistently, whereas the scale parameter cannot be estimated consistently, thus advising caution in posterior inference. We conclude our work by discussing some generalizations of SSPs, mostly in the field of biological sciences, which deal with "feature-sampling", multiple populations of individuals sharing species and classes of Markov chains.

math.ST

A Bayesian Nonparametric Approach to Species Sampling Problems with Ordering

Species-sampling problems (SSPs) refer to a vast class of statistical problems calling for the estimation of (discrete) functionals of the unknown species composition of an unobservable population. A common feature of SSPs is their invariance with respect to species labeling, which is at the core of the Bayesian nonparametric (BNP) approach to SSPs under the popular Pitman-Yor process (PYP) prior. In this paper, we develop a BNP approach to SSPs that are not "invariant" to species labeling, in the sense that an ordering or ranking is assigned to species' labels. Inspired by the population genetics literature on age-ordered alleles' compositions, we study the following SSP with ordering: given an observable sample from an unknown population of individuals belonging to species (alleles), with species' labels being ordered according to weights (ages), estimate the frequencies of the first $r$ order species' labels in an enlarged sample obtained by including additional unobservable samples. By relying on an ordered PYP prior, we obtain an explicit posterior distribution of the first $r$ order frequencies, with estimates being of easy implementation and computationally efficient. We apply our approach to the analysis of genetic variation, showing its effectiveness in estimating the frequency of the oldest allele, and then we discuss other potential applications.

stat.ME

Clustering Areal Units at Multiple Levels of Resolution to Model Crime in Philadelphia

Estimation of the spatial heterogeneity in crime incidence across an entire city is an important step towards reducing crime and increasing our understanding of the physical and social functioning of urban environments. This is a difficult modeling endeavor since crime incidence can vary smoothly across space and time but there also exist physical and social barriers that result in discontinuities in crime rates between different regions within a city. A further difficulty is that there are different levels of resolution that can be used for defining regions of a city in order to analyze crime. To address these challenges, we develop a Bayesian non-parametric approach for the clustering of urban areal units at different levels of resolution simultaneously. Our approach is evaluated with an extensive synthetic data study and then applied to the estimation of crime incidence at various levels of resolution in the city of Philadelphia.

stat.ME

Crime in Philadelphia: Bayesian Clustering with Particle Optimization

Accurate estimation of the change in crime over time is a critical first step towards better understanding of public safety in large urban environments. Bayesian hierarchical modeling is a natural way to study spatial variation in urban crime dynamics at the neighborhood level, since it facilitates principled ``sharing of information'' between spatially adjacent neighborhoods. Typically, however, cities contain many physical and social boundaries that may manifest as spatial discontinuities in crime patterns. In this situation, standard prior choices often yield overly-smooth parameter estimates, which can ultimately produce mis-calibrated forecasts. To prevent potential over-smoothing, we introduce a prior that partitions the set of neighborhoods into several clusters and encourages spatial smoothness within each cluster. In terms of model implementation, conventional stochastic search techniques are computationally prohibitive, as they must traverse a combinatorially vast space of partitions. We introduce an ensemble optimization procedure that simultaneously identifies several high probability partitions by solving one optimization problem using a new local search strategy. We then use the identified partitions to estimate crime trends in Philadelphia between 2006 and 2017. On simulated and real data, our proposed method demonstrates good estimation and partition selection performance.

stat.AP

Spatial Modeling of Trends in Crime over Time in Philadelphia

Understanding the relationship between change in crime over time and the geography of urban areas is an important problem for urban planning. Accurate estimation of changing crime rates throughout a city would aid law enforcement as well as enable studies of the association between crime and the built environment. Bayesian modeling is a promising direction since areal data require principled sharing of information to address spatial autocorrelation between proximal neighborhoods. We develop several Bayesian approaches to spatial sharing of information between neighborhoods while modeling trends in crime counts over time. We apply our methodology to estimate changes in crime throughout Philadelphia over the 2006-15 period, while also incorporating spatially-varying economic and demographic predictors. We find that the local shrinkage imposed by a conditional autoregressive model has substantial benefits in terms of out-of-sample predictive accuracy of crime. We also explore the possibility of spatial discontinuities between neighborhoods that could represent natural barriers or aspects of the built environment.

stat.AP