SearcharxivSearch

arXiv subjects

Marco Mingione

Publications and source records attributed to Marco Mingione.

11 recordsLinked to original sources

Environmental Risk Assessment via Nonhomogeneous Hidden Semi-Markov Models with Penalized Vector Auto-Regression

Motivated by the study of pollution trends in the city of Bergen, we introduce a flexible statistical framework for modeling multivariate air pollution data via a nonhomogeneous Hidden Semi-Markov Vector Auto-Regression. The hidden process captures unobserved environmental conditions, while the vector autoregressive structure accounts for temporal autocorrelation and cross-pollutant dependencies. The model further allows time-varying environmental conditions to influence both the average levels of pollutant concentrations and the duration of different transient states. Parameters are estimated via maximum likelihood using a tailored Expectation-Maximization (EM) algorithm, integrated with state-specific $\ell_1$ regularization to control overfitting and automatically select relevant temporal lags. The proposal is tested on simulated data under different scenarios and then applied to daily concentrations of nitrogens and particulate matter recorded in a urban area. Environmental risk is assessed by a Shapley value-based decomposition that attribute marginal risk contributions. This approach offers a comprehensive framework for multivariate environmental risk modeling, enabling better identification of high-pollution episodes and informing policy interventions.

stat.ME

A new hierarchical distribution on arbitrary sparse precision matrices

We introduce a general strategy for defining distributions over the space of sparse symmetric positive definite matrices. Our method utilizes the Cholesky factorization of the precision matrix, imposing sparsity through constraints on its elements while preserving their independence and avoiding the numerical evaluation of normalization constants. In particular, we develop the S-Bartlett as a modified Bartlett decomposition, recovering the standard Wishart as a particular case. By incorporating a Spike-and-Slab prior to model graph sparsity, our approach facilitates Bayesian estimation through a tailored MCMC routine based on a Dual Averaging Hamiltonian Monte Carlo update. This framework extends naturally to the Generalized Linear Model setting, enabling applications to non-Gaussian outcomes via latent Gaussian variables. We test and compare the proposed S-Bartelett prior with the G-Wishart both on simulated and real data. Results highlight that the S-Bartlett prior offers a flexible alternative for estimating sparse precision matrices, with potential applications across diverse fields.

stat.ME

Does wind affect the orientation of vegetation stripes? A copula-based mixture model for axial and circular data

Motivated by a case study of vegetation patterns, we introduce a mixture model with concomitant variables to examine the association between the orientation of vegetation stripes and wind direction. The proposal relies on a novel copula-based bivariate distribution for mixed axial and circular observations and provides a parsimonious and computationally tractable approach to examine the dependence of two environmental variables observed in a complex manifold. The findings suggest that dominant winds shape the orientation of vegetation stripes through a mechanism of neighbouring plants providing wind shelter to downwind individuals.

stat.AP

Copula-based models for correlated circular data

We exploit Gaussian copulas to specify a class of multivariate circular distributions and obtain parametric models for the analysis of correlated circular data. This approach provides a straightforward extension of traditional multivariate normal models to the circular setting, without imposing restrictions on the marginal data distribution nor requiring overwhelming routines for parameter estimation. The proposal is illustrated on two case studies of animal orientation and sea currents, where we propose an autoregressive model for circular time series and a geostatistical model for circular spatial series.

stat.ME

Nonhomogeneous hidden semi-Markov models for toroidal data

A nonhomogeneous hidden semi-Markov model is proposed to segment toroidal time series according to a finite number of latent regimes and, simultaneously, estimate the influence of time-varying covariates on the process' survival under each regime. The model is a mixture of toroidal densities, whose parameters depend on the evolution of a semi-Markov chain, which is in turn modulated by time-varying covariates through a proportional hazards assumption. Parameter estimates are obtained using an EM algorithm that relies on an efficient augmentation of the latent process. The proposal is illustrated on a time series of wind and wave directions recorded during winter.

stat.AP

Finite mixtures in capture-recapture surveys for modelling residency patterns in marine wildlife populations

In this work, the goal is to estimate the abundance of an animal population using data coming from capture-recapture surveys. We leverage the prior knowledge about the population's structure to specify a parsimonious finite mixture model tailored to its behavioral pattern. Inference is carried out under the Bayesian framework, where we discuss suitable priors' specification that could alleviate label-switching and non-identifiability issues affecting finite mixtures. We conduct simulation experiments to show the competitive advantage of our proposal over less specific alternatives. Finally, the proposed model is used to estimate the common bottlenose dolphins' population size at the Tiber River estuary (Mediterranean Sea), using data collected via photo-identification from 2018 to 2020. Results provide novel insights on the population's size and structure, and shed light on some of the ecological processes governing the population dynamics.

stat.AP

Monitoring the West-Nile virus outbreaks in Italy using open-access data

This paper introduces a comprehensive and original database on West Nile virus (WNV) outbreaks that have occurred in Italy from September 2012 to November 2022. We have digitized bulletins published by the Italian National Institute of Health (ISS) to demonstrate the potential utilization of this data for the research community. Our aim is to establish a centralized open-access repository that facilitates analysis and monitoring of the disease. We have collected and curated data on the type of infected host, along with additional information whenever available, including the type of infection, age, and geographic details at different levels of spatial aggregation. By combining our data with other sources of information such as weather data, it becomes possible to assess potential relationships between WNV outbreaks and environmental factors. We strongly believe in supporting public oversight of government epidemic management, and we emphasize that open data plays a crucial role in generating reliable results by enabling greater transparency.

stat.AP

Bayesian Hierarchical Modeling and Analysis for Actigraph Data from Wearable Devices

The majority of Americans fail to achieve recommended levels of physical activity, which leads to numerous preventable health problems such as diabetes, hypertension, and heart diseases. This has generated substantial interest in monitoring human activity to gear interventions toward environmental features that may relate to higher physical activity. Wearable devices, such as wrist-worn sensors that monitor gross motor activity (actigraph units) continuously record the activity levels of a subject, producing massive amounts of high-resolution measurements. Analyzing actigraph data needs to account for spatial and temporal information on trajectories or paths traversed by subjects wearing such devices. Inferential objectives include estimating a subject's physical activity levels along a given trajectory; identifying trajectories that are more likely to produce higher levels of physical activity for a given subject; and predicting expected levels of physical activity in any proposed new trajectory for a given set of health attributes. Here, we devise a Bayesian hierarchical modeling framework for spatial-temporal actigraphy data to deliver fully model-based inference on trajectories while accounting for subject-level health attributes and spatial-temporal dependencies. We undertake a comprehensive analysis of an original dataset from the Physical Activity through Sustainable Transport Approaches in Los Angeles (PASTA-LA) study to ascertain spatial zones and trajectories exhibiting significantly higher levels of physical activity while accounting for various sources of heterogeneity.

stat.AP

Spatial modelling of COVID-19 incident cases using Richards' curve: an application to the Italian regions

We introduce an extended generalised logistic growth model for discrete outcomes, in which a network structure can be specified to deal with spatial dependence and time dependence is dealt with using an Auto-Regressive approach. A major challenge concerns the specification of the network structure, crucial to consistently estimate the canonical parameters of the generalised logistic curve, e.g. peak time and height. Parameters are estimated under the Bayesian framework, using the {\texttt{ Stan}} probabilistic programming language. The proposed approach is motivated by the analysis of the first and second wave of COVID-19 in Italy, i.e. from February 2020 to July 2020 and from July 2020 to December 2020, respectively. We analyse data at the regional level and, interestingly enough, prove that substantial spatial and temporal dependence occurred in both waves, although strong restrictive measures were implemented during the first wave. Accurate predictions are obtained, improving those of the model where independence across regions is assumed.

stat.AP

Nowcasting COVID-19 incidence indicators during the Italian first outbreak

A novel parametric regression model is proposed to fit incidence data typically collected during epidemics. The proposal is motivated by real-time monitoring and short-term forecasting of the main epidemiological indicators within the first outbreak of COVID-19 in Italy. Accurate short-term predictions, including the potential effect of exogenous or external variables are provided; this ensures to accurately predict important characteristics of the epidemic (e.g., peak time and height), allowing for a better allocation of health resources over time. Parameters estimation is carried out in a maximum likelihood framework. All computational details required to reproduce the approach and replicate the results are provided.

stat.AP

Adversarial Out-domain Examples for Generative Models

Deep generative models are rapidly becoming a common tool for researchers and developers. However, as exhaustively shown for the family of discriminative models, the test-time inference of deep neural networks cannot be fully controlled and erroneous behaviors can be induced by an attacker. In the present work, we show how a malicious user can force a pre-trained generator to reproduce arbitrary data instances by feeding it suitable adversarial inputs. Moreover, we show that these adversarial latent vectors can be shaped so as to be statistically indistinguishable from the set of genuine inputs. The proposed attack technique is evaluated with respect to various GAN images generators using different architectures, training processes and for both conditional and not-conditional setups.

cs.LG