SearcharxivSearch

arXiv subjects

Paola Stolfi

Publications and source records attributed to Paola Stolfi.

4 recordsLinked to original sources

Bayesian methods for the identification of model parameters for water transport in porous media

The structure of the nonlinear inverse problem arising from capillarity-driven imbibition in porous media is investigated, considering a degenerate parabolic PDE with compactly supported diffusivity and boundary-driven fluxes as the governing forward model. The inverse problem -- inferring hydraulic model parameters from sparse integral absorption measurements -- is inherently ill-posed: the nonlinear forward operator induces anisotropic parameter sensitivity and structured correlations that render the calibration landscape non-convex and partially unidentifiable. To characterise this structure rigorously, Approximate Bayesian Computation with Sequential Monte Carlo (ABC-SMC) is adopted as a likelihood-free inferential framework, bypassing the analytical intractability of the likelihood while providing full posterior distributions over the parameter space. Two physically motivated parameterisations of the diffusivity function are analysed -- the Natalini-Nitsch (NN) and the BkP formulations. It is shown that the posterior geometry obtained via ABC-SMC encodes, in directly readable form, the sensitivity structure of the nonlinear forward operator: the principal component decomposition of the posterior covariance provides a natural hierarchy of parameter sensitivity, with low-variance eigendirections identifying the parameter combinations to which the forward map is most responsive. This geometric decomposition constitutes a principled and computationally efficient alternative to classical sensitivity analysis, arising as a byproduct of the calibration procedure. These findings are established through both synthetic experiments, confirming accurate parameter recovery, and real laboratory imbibition data from materials of cultural heritage relevance.

math.DS

Normal approximation of Random Gaussian Neural Networks

In this paper we provide explicit upper bounds on some distances between the (law of the) output of a random Gaussian NN and (the law of) a random Gaussian vector. Our results concern both shallow random Gaussian neural networks with univariate output and fully connected and deep random Gaussian neural networks, with a rather general activation function. The upper bounds show how the widths of the layers, the activation functions and other architecture parameters affect the Gaussian approximation of the ouput. Our techniques, relying on Stein's method and integration by parts formulas for the Gaussian law, yield estimates on distances which are indeed integral probability metrics, and include the total variation and the convex distances. These latter metrics are defined by testing against indicator functions of suitable measurable sets, and so allow for accurate estimates of the probability that the output is localized in some region of the space. Such estimates have a significant interest both from a practitioner's and a theorist's perspective.

math.PR

NIAPU: network-informed adaptive positive-unlabeled learning for disease gene identification

Gene-disease associations are fundamental for understanding disease etiology and developing effective interventions and treatments. Identifying genes not yet associated with a disease due to a lack of studies is a challenging task in which prioritization based on prior knowledge is an important element. The computational search for new candidate disease genes may be eased by positive-unlabeled learning, the machine learning setting in which only a subset of instances are labeled as positive while the rest of the data set is unlabeled. In this work, we propose a set of effective network-based features to be used in a novel Markov diffusion-based multi-class labeling strategy for putative disease gene discovery. The performances of the new labeling algorithm and the effectiveness of the proposed features have been tested on ten different disease data sets using three machine learning algorithms. The new features have been compared against classical topological and functional/ontological features and a set of network- and biological-derived features already used in gene discovery tasks. The predictive power of the integrated methodology in searching for new disease genes has been found to be competitive against state-of-the-art algorithms.

cs.LG

The Sparse Multivariate Method of Simulated Quantiles

In this paper the method of simulated quantiles (MSQ) of Dominicy and Veredas (2013) and Dominick et al. (2013) is extended to a general multivariate framework (MMSQ) and to provide a sparse estimator of the scale matrix (sparse-MMSQ). The MSQ, like alternative likelihood-free procedures, is based on the minimisation of the distance between appropriate statistics evaluated on the true and synthetic data simulated from the postulated model. Those statistics are functions of the quantiles providing an effective way to deal with distributions that do not admit moments of any order like the $α$-Stable or the Tukey lambda distribution. The lack of a natural ordering represents the major challenge for the extension of the method to the multivariate framework. Here, we rely on the notion of projectional quantile recently introduced by Hallin etal. (2010) and Kong Mizera (2012). We establish consistency and asymptotic normality of the proposed estimator. The smoothly clipped absolute deviation (SCAD) $\ell_1$--penalty of Fan and Li (2001) is then introduced into the MMSQ objective function in order to achieve sparse estimation of the scaling matrix which is the major responsible for the curse of dimensionality problem. We extend the asymptotic theory and we show that the sparse-MMSQ estimator enjoys the oracle properties under mild regularity conditions. The method is illustrated and its effectiveness is tested using several synthetic datasets simulated from the Elliptical Stable distribution (ESD) for which alternative methods are recognised to perform poorly. The method is then applied to build a new network-based systemic risk measurement framework. The proposed methodology to build the network relies on a new systemic risk measure and on a parametric test of statistical dominance.

stat.ME