SearcharxivSearch

arXiv subjects

Federico Bassetti

Publications and source records attributed to Federico Bassetti.

At least 19 recordsLinked to original sources

Large deviation principles for convolutional Bayesian neural networks

While suitably scaled CNNs with Gaussian initialization are known to converge to Gaussian processes as the number of channels diverges, little is known beyond this Gaussian limit. We establish a large deviation principle (LDP) for convolutional neural networks in the infinite-channel regime. We consider a broad class of multidimensional CNN architectures characterized by general receptive fields encoded through a patch-extractor function satisfying mild structural assumptions. Our main result establishes a large deviation principle (LDP) for the sequence of conditional covariance matrices under Gaussian prior distribution on the weights. We further derive an LDP for the posterior distribution obtained by conditioning on a finite number of observations. In addition, we provide a streamlined proof of the concentration of the conditional covariances and of the Gaussian equivalence of the network. To the best of our knowledge, this is the first large deviation principle established for convolutional neural networks.

math.PR

A Caveat on Metrizing Convergence in Distribution on Hilbert Spaces

We consider Sobolev-type distances on probability measures over separable Hilbert spaces involving the Schatten-$p$ norms, which include as special cases a distance first introduced by Bourguin and Campese (2020) when $p=2$, and a distance introduced by Gin\'e and Leon (1980) when $p=\infty$. Our analysis shows that, unless $p=\infty$, these distances fail to metrize convergence in distribution in infinite dimensions. This clarifies several inconsistencies and misconceptions in the recent literature that arose from confusion between different types of distances.

math.PR

LDP for the covariance process in fully connected neural networks

In this work, we study large deviation properties of the covariance process in fully connected Gaussian deep neural networks. More precisely, we establish a large deviation principle (LDP) for the covariance process in a functional framework, viewing it as a process in the space of continuous functions. As key applications of our main results, we obtain posterior LDPs under Gaussian likelihood in both the infinite-width and mean-field regimes. The proof is based on an LDP for the covariance process as a Markov process valued in the space of non-negative, symmetric trace-class operators equipped with the trace norm.

math.PR

Proportional infinite-width infinite-depth limit for deep linear neural networks

We study the distributional properties of linear neural networks with random parameters in the context of large networks, where the number of layers diverges in proportion to the number of neurons per layer. Prior works have shown that in the infinite-width regime, where the number of neurons per layer grows to infinity while the depth remains fixed, neural networks converge to a Gaussian process, known as the Neural Network Gaussian Process. However, this Gaussian limit sacrifices descriptive power, as it lacks the ability to learn dependent features and produce output correlations that reflect observed labels. Motivated by these limitations, we explore the joint proportional limit in which both depth and width diverge but maintain a constant ratio, yielding a non-Gaussian distribution that retains correlations between outputs. Our contribution extends previous works by rigorously characterizing, for linear activation functions, the limiting distribution as a nontrivial mixture of Gaussians.

stat.ML

Feature learning in finite-width Bayesian deep linear networks with multiple outputs and convolutional layers

Deep linear networks have been extensively studied, as they provide simplified models of deep learning. However, little is known in the case of finite-width architectures with multiple outputs and convolutional layers. In this manuscript, we provide rigorous results for the statistics of functions implemented by the aforementioned class of networks, thus moving closer to a complete characterization of feature learning in the Bayesian setting. Our results include: (i) an exact and elementary non-asymptotic integral representation for the joint prior distribution over the outputs, given in terms of a mixture of Gaussians; (ii) an analytical formula for the posterior distribution in the case of squared error loss function (Gaussian likelihood); (iii) a quantitative description of the feature learning infinite-width regime, using large deviation theory. From a physical perspective, deep architectures with multiple outputs or convolutional layers represent different manifestations of kernel shape renormalization, and our work provides a dictionary that translates this physics intuition and terminology into rigorous Bayesian statistics.

stat.ML

A Spatiotemporal Gamma Shot Noise Cox Process

A new discrete-time shot noise Cox process for spatiotemporal data is proposed. The random intensity is driven by a dependent sequence of latent gamma random measures. Some properties of the latent process are derived, such as an autoregressive representation and the Laplace functional. A simulation method based on the Inverse L\'evy Measure algorithm is provided. Moreover, these results are used to derive the moment, predictive, and pair correlation measures for the proposed shot noise Cox process. The model is flexible yet tractable, allowing it to capture persistence, global trends, and latent spatial and temporal factors. A Bayesian inference approach is adopted, and efficient Markov Chain Monte Carlo algorithms based on conditional Sequential Monte Carlo and adaptive Metropolis-Hastings are proposed. An application to georeferenced wildfire data illustrates the properties of the model and inference.

stat.ME

First-order integer-valued autoregressive processes with Generalized Katz innovations

A new integer--valued autoregressive process (INAR) with Generalised Lagrangian Katz (GLK) innovations is defined. This process family provides a flexible modelling framework for count data, allowing for under and over--dispersion, asymmetry, and excess of kurtosis and includes standard INAR models such as Generalized Poisson and Negative Binomial as special cases. We show that the GLK--INAR process is discrete semi--self--decomposable, infinite divisible, stable by aggregation and provides stationarity conditions. Some extensions are discussed, such as the Markov--Switching and the zero--inflated GLK--INARs. A Bayesian inference framework and an efficient posterior approximation procedure are introduced. The proposed models are applied to 130 time series from Google Trend, which proxy the worldwide public concern about climate change. New evidence is found of heterogeneity across time, countries and keywords in the persistence, uncertainty, and long--run public awareness level.

stat.ME

Remote teaching data-driven physical modeling through a COVID-19 open ended data challenge

Physics can be seen as a conceptual approach to scientific problems, a method for discovery, but teaching this aspect of our discipline can be a challenge. We report on a first-time remote teaching experience for a computational physics third-year physics laboratory class taught in the first part of the 2020 COVID-19 pandemic (March-May 2020). To convey a ``physics of data" approach to data analysis and data-driven physical modeling we used interdisciplinary data sources, with an openended ``COVID-19 data challenge" project as the core of the course. COVID-19 epidemiological data provided an ideal setting for motivating the students to deal with complex problems, where there is no unique or preconceived solution. Our results indicate that such problems yield qualitatively different improvements compared to close-ended projects, as well as point to critical aspects in using these problems as a teaching strategy. By breaking the students' expectations of unidirectionality, remote teaching provided unexpected opportunities to promote active work and active learning.

physics.ed-ph

Clustering structure for species sampling sequences with general base measure

We investigate the clustering structure of species sampling sequences $(ξ_n)_n$, with general base measure. Such sequences are exchangeable with a species sampling random probability as directing measure. The clustering properties of these sequences are interesting for Bayesian nonparametrics applications, where mixed base measures are used, for example, to accommodate sharp hypotheses in regression problems and provide sparsity. In this paper, we prove a stochastic representation for $(ξ_n)_n$ in terms of a latent exchangeable random partition. We provide explicit expression of the EPPF of the partition generated by $(ξ_n)_n$ in terms of the EPPF of the latent partition. We investigate the asymptotic behaviour of the total number of blocks and of the number of blocks with fixed cardinality in the partition generated by $(ξ_n)_n$.

math.PR

On the Computation of Kantorovich-Wasserstein Distances between 2D-Histograms by Uncapacitated Minimum Cost Flows

In this work, we present a method to compute the Kantorovich-Wasserstein distance of order one between a pair of two-dimensional histograms. Recent works in Computer Vision and Machine Learning have shown the benefits of measuring Wasserstein distances of order one between histograms with $n$ bins, by solving a classical transportation problem on very large complete bipartite graphs with $n$ nodes and $n^2$ edges. The main contribution of our work is to approximate the original transportation problem by an uncapacitated min cost flow problem on a reduced flow network of size $O(n)$ that exploits the geometric structure of the cost function. More precisely, when the distance among the bin centers is measured with the 1-norm or the $\infty$-norm, our approach provides an optimal solution. When the distance among bins is measured with the 2-norm: (i) we derive a quantitative estimate on the error between optimal and approximate solution; (ii) given the error, we construct a reduced flow network of size $O(n)$. We numerically show the benefits of our approach by computing Wasserstein distances of order one on a set of grey scale images used as benchmark in the literature. We show how our approach scales with the size of the images with 1-norm, 2-norm and $\infty$-norm ground distances, and we compare it with other two methods which are largely used in the literature.

math.OC

Computing Kantorovich-Wasserstein Distances on $d$-dimensional histograms using $(d+1)$-partite graphs

This paper presents a novel method to compute the exact Kantorovich-Wasserstein distance between a pair of $d$-dimensional histograms having $n$ bins each. We prove that this problem is equivalent to an uncapacitated minimum cost flow problem on a $(d+1)$-partite graph with $(d+1)n$ nodes and $dn^{\frac{d+1}{d}}$ arcs, whenever the cost is separable along the principal $d$-dimensional directions. We show numerically the benefits of our approach by computing the Kantorovich-Wasserstein distance of order 2 among two sets of instances: gray scale images and $d$-dimensional biomedical histograms. On these types of instances, our approach is competitive with state-of-the-art optimal transport algorithms.

math.OC

Hierarchical Species Sampling Models

This paper introduces a general class of hierarchical nonparametric prior distributions. The random probability measures are constructed by a hierarchy of generalized species sampling processes with possibly non-diffuse base measures. The proposed framework provides a general probabilistic foundation for hierarchical random measures with either atomic or mixed base measures and allows for studying their properties, such as the distribution of the marginal and total number of clusters. We show that hierarchical species sampling models have a Chinese Restaurants Franchise representation and can be used as prior distributions to undertake Bayesian nonparametric inference. We provide a method to sample from the posterior distribution together with some numerical illustrations. Our class of priors includes some new hierarchical mixture priors such as the hierarchical Gnedin measures, and other well-known prior distributions such as the hierarchical Pitman-Yor and the hierarchical normalized random measures.

stat.ME

Cell-to-cell variability and robustness in S-phase duration from genome replication kinetics

Genome replication, a key process for a cell, relies on stochastic initiation by replication origins, causing a variability of replication timing from cell to cell. While stochastic models of eukaryotic replication are widely available, the link between the key parameters and overall replication timing has not been addressed systematically.We use a combined analytical and computational approach to calculate how positions and strength of many origins lead to a given cell-to-cell variability of total duration of the replication of a large region, a chromosome or the entire genome.Specifically, the total replication timing can be framed as an extreme-value problem, since it is due to the last region that replicates in each cell. Our calculations identify two regimes based on the spread between characteristic completion times of all inter-origin regions of a genome. For widely different completion times, timing is set by the single specific region that is typically the last to replicate in all cells. Conversely, when the completion time of all regions are comparable,an extreme-value estimate shows that the cell-to-cell variability of genome replication timing has universal properties. Comparison with available data shows that the replication program of three yeast species falls in this extreme-value regime.

q-bio.GN

Bayesian Nonparametric Calibration and Combination of Predictive Distributions

We introduce a Bayesian approach to predictive density calibration and combination that accounts for parameter uncertainty and model set incompleteness through the use of random calibration functionals and random combination weights. Building on the work of Ranjan, R. and Gneiting, T. (2010) and Gneiting, T. and Ranjan, R. (2013), we use infinite beta mixtures for the calibration. The proposed Bayesian nonparametric approach takes advantage of the flexibility of Dirichlet process mixtures to achieve any continuous deformation of linearly combined predictive distributions. The inference procedure is based on Gibbs sampling and allows accounting for uncertainty in the number of mixture components, mixture weights, and calibration parameters. The weak posterior consistency of the Bayesian nonparametric calibration is provided under suitable conditions for unknown true density. We study the methodology in simulation examples with fat tails and multimodal densities and apply it to density forecasts of daily S&P returns and daily maximum wind speed at the Frankfurt airport.

stat.AP

Mean field dynamics of collisional processes with duplication, loss and copy

In this paper we introduce and discuss kinetic equations for the evolution of the probability distribution of the number of particles in a population subject to binary interactions. The microscopic binary law of interaction is assumed to be dependent on fixed-in-time random parameters which describe both birth and death of particles, and the migration rule. These assumptions lead to a Boltzmann-type equation that in the case in which the mean number of the population is preserved, can be fully studied, by obtaining in some case the analytic description of the steady profile. In all cases, however, a simpler kinetic description can be derived, by considering the limit of quasi-invariant interactions. This procedure allows to describe the evolution process in terms of a linear kinetic transport-type equation. Among the various processes that can be described in this way, one recognizes the Lea-Coulson model of mutation processes in bacteria, a variation of the original model proposed by Luria and Delbrück.

math.PR

Generalized Species Sampling Priors with Latent Beta reinforcements

Many popular Bayesian nonparametric priors can be characterized in terms of exchangeable species sampling sequences. However, in some applications, exchangeability may not be appropriate. We introduce a {novel and probabilistically coherent family of non-exchangeable species sampling sequences characterized by a tractable predictive probability function with weights driven by a sequence of independent Beta random variables. We compare their theoretical clustering properties with those of the Dirichlet Process and the two parameters Poisson-Dirichlet process. The proposed construction provides a complete characterization of the joint process, differently from existing work. We then propose the use of such process as prior distribution in a hierarchical Bayes modeling framework, and we describe a Markov Chain Monte Carlo sampler for posterior inference. We evaluate the performance of the prior and the robustness of the resulting inference in a simulation study, providing a comparison with popular Dirichlet Processes mixtures and Hidden Markov Models. Finally, we develop an application to the detection of chromosomal aberrations in breast cancer by leveraging array CGH data.

math.ST

Infinite energy solutions to inelastic homogeneous Boltzmann equation

This paper is concerned with the existence, shape and dynamical stability of infinite-energy equilibria for a general class of spatially homogeneous kinetic equations in space dimensions $d \geq 3$. Our results cover in particular Bobylëv's model for inelastic Maxwell molecules. First, we show under certain conditions on the collision kernel, that there exists an index $α\in(0,2)$ such that the equation possesses a nontrivial stationary solution, which is a scale mixture of radially symmetric $α$-stable laws. We also characterize the mixing distribution as the fixed point of a smoothing transformation. Second, we prove that any transient solution that emerges from the NDA of some (not necessarily radial symmetric) $α$-stable distribution converges to an equilibrium. The key element of the convergence proof is an application of the central limit theorem to a representation of the transient solution as a weighted sum of i.i.d. random vectors.

math-ph

Speed of convergence to equilibrium in Wasserstein metrics for Kac-s like kinetic equations

This work deals with a class of one-dimensional measure-valued kinetic equations, which constitute extensions of the Kac caricature. It is known that if the initial datum belongs to the domain of normal attraction of an α-stable law, the solution of the equation converges weakly to a suitable scale mixture of centered α-stable laws. In this paper we present explicit exponential rates for the convergence to equilibrium in Kantorovich-Wasserstein distances of order p>α, under the natural assumption that the distance between the initial datum and the limit distribution is finite. For α=2 this assumption reduces to the finiteness of the absolute moment of order p of the initial datum. On the contrary, when α<2, the situation is more problematic due to the fact that both the limit distribution and the initial datum have infinite absolute moment of any order p >α. For this case, we provide sufficient conditions for the finiteness of the Kantorovich-Wasserstein distance.

math.PR