SearcharxivSearch

arXiv subjects

Davide Pigoli

Publications and source records attributed to Davide Pigoli.

13 recordsLinked to original sources

Autoregressive Processes on Riemannian Manifolds

This paper introduces a Riemannian autoregressive (R-AR) model of order one for manifold-valued time series. The model is specified through an autoregressive process in the tangent space at a reference point, which is mapped to the manifold via the exponential map. It is characterised by two parameters: a reference point $\mu$, representing the intrinsic central tendency through the Fr\'echet mean, and an autoregressive parameter $\phi$, governing the dependence structure and stationarity properties. When these parameters are unknown, their estimation introduces geometric and probabilistic challenges arising from the intrinsic estimation of $\mu$ and from estimating $\phi$ as the base point varies over the manifold and the associated tangent spaces change. The estimation theory is developed using a strong law of large numbers for the sample Fr\'echet mean set of ergodic Markov chains, together with ergodic arguments for the strong consistency of the autoregressive parameter estimator. The framework is validated through numerical simulations in the hyperbolic plane and an application to aerosol size distributions on the Fisher-Rao manifold, demonstrating how the proposed model can characterise mean-reverting dynamics in nonlinear geometries.

math.ST

Optimal design of experiments for functional linear models with dynamic factors

In this work we build optimal experimental designs for precise estimation of the functional coefficient of a function-on-function linear regression model where both the response and the factors are continuous functions of time. After obtaining the variance-covariance matrix of the estimator of the functional coefficient which minimizes the integrated sum of square of errors, we extend the classical definition of optimal design to this estimator, and we provide the expression of the A-optimal and of the D-optimal designs. Examples of optimal designs for dynamic experimental factors are then computed through a suitable algorithm, and we discuss different scenarios in terms of the set of basis functions used for their representation. Finally, we present an example with simulated data to illustrate the feasibility of our methodology.

stat.ME

A Riemannian covariance for manifold-valued data

The extension of bivariate measures of dependence to non-Euclidean spaces is a challenging problem. The non-linear nature of these spaces makes the generalisation of classical measures of linear dependence (such as the covariance) not trivial. In this paper, we propose a novel approach to measure stochastic dependence between two random variables taking values in a Riemannian manifold, with the aim of both generalising the classical concepts of covariance and correlation and building a connection to Fr\'echet moments of random variables on manifolds. We introduce generalised local measures of covariance and correlation and we show that the latter is a natural extension of Pearson correlation. We then propose suitable estimators for these quantities and we prove strong consistency results. Finally, we demonstrate their effectiveness through simulated examples and a real-world application.

math.ST

Statistics in Phonetics

Phonetics is the scientific field concerned with the study of how speech is produced, heard and perceived. It abounds with data, such as acoustic speech recordings, neuroimaging data, or articulatory data. In this paper, we provide an introduction to different areas of phonetics (acoustic phonetics, sociophonetics, speech perception, articulatory phonetics, speech inversion, sound change, and speech technology), an overview of the statistical methods for analyzing their data, and an introduction to the signal processing methods commonly applied to speech recordings. A major transition in the statistical modeling of phonetic data has been the shift from fixed effects to random effects regression models, the modeling of curve data (for instance via GAMMs or FDA methods), and the use of Bayesian methods. This shift has been driven in part by the increased focus on large speech corpora in phonetics, which has been driven by machine learning methods such as forced alignment. We conclude by identifying opportunities for future research.

stat.AP

A large-scale and PCR-referenced vocal audio dataset for COVID-19

The UK COVID-19 Vocal Audio Dataset is designed for the training and evaluation of machine learning models that classify SARS-CoV-2 infection status or associated respiratory symptoms using vocal audio. The UK Health Security Agency recruited voluntary participants through the national Test and Trace programme and the REACT-1 survey in England from March 2021 to March 2022, during dominant transmission of the Alpha and Delta SARS-CoV-2 variants and some Omicron variant sublineages. Audio recordings of volitional coughs, exhalations, and speech were collected in the 'Speak up to help beat coronavirus' digital survey alongside demographic, self-reported symptom and respiratory condition data, and linked to SARS-CoV-2 test results. The UK COVID-19 Vocal Audio Dataset represents the largest collection of SARS-CoV-2 PCR-referenced audio recordings to date. PCR results were linked to 70,794 of 72,999 participants and 24,155 of 25,776 positive cases. Respiratory symptoms were reported by 45.62% of participants. This dataset has additional potential uses for bioacoustics research, with 11.30% participants reporting asthma, and 27.20% with linked influenza PCR test results.

cs.SD

Audio-based AI classifiers show no evidence of improved COVID-19 screening over simple symptoms checkers

Recent work has reported that AI classifiers trained on audio recordings can accurately predict severe acute respiratory syndrome coronavirus 2 (SARSCoV2) infection status. Here, we undertake a large scale study of audio-based deep learning classifiers, as part of the UK governments pandemic response. We collect and analyse a dataset of audio recordings from 67,842 individuals with linked metadata, including reverse transcription polymerase chain reaction (PCR) test outcomes, of whom 23,514 tested positive for SARS CoV 2. Subjects were recruited via the UK governments National Health Service Test-and-Trace programme and the REal-time Assessment of Community Transmission (REACT) randomised surveillance survey. In an unadjusted analysis of our dataset AI classifiers predict SARS-CoV-2 infection status with high accuracy (Receiver Operating Characteristic Area Under the Curve (ROCAUC) 0.846 [0.838, 0.854]) consistent with the findings of previous studies. However, after matching on measured confounders, such as age, gender, and self reported symptoms, our classifiers performance is much weaker (ROC-AUC 0.619 [0.594, 0.644]). Upon quantifying the utility of audio based classifiers in practical settings, we find them to be outperformed by simple predictive scores based on user reported symptoms.

cs.SD

Statistical Design and Analysis for Robust Machine Learning: A Case Study from COVID-19

Since early in the coronavirus disease 2019 (COVID-19) pandemic, there has been interest in using artificial intelligence methods to predict COVID-19 infection status based on vocal audio signals, for example cough recordings. However, existing studies have limitations in terms of data collection and of the assessment of the performances of the proposed predictive models. This paper rigorously assesses state-of-the-art machine learning techniques used to predict COVID-19 infection status based on vocal audio signals, using a dataset collected by the UK Health Security Agency. This dataset includes acoustic recordings and extensive study participant meta-data. We provide guidelines on testing the performance of methods to classify COVID-19 infection status based on acoustic features and we discuss how these can be extended more generally to the development and assessment of predictive methods based on public health datasets.

cs.SD

Functional Principal Component Analysis as a Versatile Technique to Understand and Predict the Electric Consumption Patterns

Understanding and predicting the electric consumption patterns in the short-, mid- and long-term, at the distribution and transmission level, is a fundamental asset for smart grids infrastructure planning, dynamic network reconfiguration, dynamic energy pricing and savings, and thus energy efficiency. This work introduces the Functional Principal Component Analysis (FPCA) as a versatile method to both investigate and predict, at different level of spatial aggregation, the consumption patterns. The method was applied to a unique and sensitive dataset that includes electric consumption and contractual information of Milan metropolitan area. The decomposition of the load patterns into principal functions was found to be a powerful method to identify the physical and behavioral causes underlying the daily consumptions, given knowledge of exogenous variables such as calendar and meteorological data. The effectiveness of long-term predictions based on principal functions was proved on Milan's metropolitan area data and assessed on a publicly-available dataset.

eess.SY

Kriging Riemannian Data via Random Domain Decompositions

Data taking value on a Riemannian manifold and observed over a complex spatial domain are becoming more frequent in applications, e.g. in environmental sciences and in geoscience. The analysis of these data needs to rely on local models to account for the non stationarity of the generating random process, the non linearity of the manifold and the complex topology of the domain. In this paper, we propose to use a random domain decomposition approach to estimate an ensemble of local models and then to aggregate the predictions of the local models through Fréchet averaging. The algorithm is introduced in complete generality and is valid for data belonging to any smooth Riemannian manifold but it is then described in details for the case of the manifold of positive definite matrices, the hypersphere and the Cholesky manifold. The predictive performance of the method are explored via simulation studies for covariance matrices and correlation matrices, where the Cholesky manifold geometry is used. Finally, the method is illustrated on an environmental dataset observed over the Chesapeake Bay (USA).

stat.ME

A Spatial Modeling Approach for Linguistic Object Data: Analysing dialect sound variations across Great Britain

Dialect variation is of considerable interest in linguistics and other social sciences. However, traditionally it has been studied using proxies (transcriptions) rather than acoustic recordings directly. We introduce novel statistical techniques to analyse geolocalised speech recordings and to explore the spatial variation of pronunciations continuously over the region of interest, as opposed to traditional isoglosses, which provide a discrete partition of the region. Data of this type require an explicit modeling of the variation in the mean and the covariance. Usual Euclidean metrics are not appropriate, and we therefore introduce the concept of $d$-covariance, which allows consistent estimation both in space and at individual locations. We then propose spatial smoothing for these objects which accounts for the possibly non convex geometry of the domain of interest. We apply the proposed method to data from the spoken part of the British National Corpus, deposited at the British Library, London, and we produce maps of the dialect variation over Great Britain. In addition, the methods allow for acoustic reconstruction across the domain of interest, allowing researchers to listen to the statistical analysis.

stat.ME

The statistical analysis of acoustic phonetic data: exploring differences between spoken Romance languages

The historical and geographical spread from older to more modern languages has long been studied by examining textual changes and in terms of changes in phonetic transcriptions. However, it is more difficult to analyze language change from an acoustic point of view, although this is usually the dominant mode of transmission. We propose a novel analysis approach for acoustic phonetic data, where the aim will be to statistically model the acoustic properties of spoken words. We explore phonetic variation and change using a time-frequency representation, namely the log-spectrograms of speech recordings. We identify time and frequency covariance functions as a feature of the language; in contrast, mean spectrograms depend mostly on the particular word that has been uttered. We build models for the mean and covariances (taking into account the restrictions placed on the statistical analysis of such objects) and use these to define a phonetic transformation that models how an individual speaker would sound in a different language, allowing the exploration of phonetic differences between languages. Finally, we map back these transformations to the domain of sound recordings, allowing us to listen to the output of the statistical analysis. The proposed approach is demonstrated using recordings of the words corresponding to the numbers from "one" to "ten" as pronounced by speakers from five different Romance languages.

stat.AP

Permutation tests for the equality of covariance operators of functional data with applications to evolutionary biology

In this paper, we generalize the metric-based permutation test for the equality of covariance operators proposed by Pigoli et al. (2014) to the case of multiple samples of functional data. To this end, the non-parametric combination methodology of Pesarin and Salmaso (2010) is used to combine all the pairwise comparisons between samples into a global test. Different combining functions and permutation strategies are reviewed and analyzed in detail. The resulting test allows to make inference on the equality of the covariance operators of multiple groups and, if there is evidence to reject the null hypothesis, to identify the pairs of groups having different covariances. It is shown that, for some combining functions, step-down adjusting procedures are available to control for the multiple testing problem in this setting. The empirical power of this new test is then explored via simulations and compared with those of existing alternative approaches in different scenarios. Finally, the proposed methodology is applied to data from wheel running activity experiments, that used selective breeding to study the evolution of locomotor behavior in mice.

stat.ME

Tests for separability in nonparametric covariance operators of random surfaces

The assumption of separability of the covariance operator for a random image or hypersurface can be of substantial use in applications, especially in situations where the accurate estimation of the full covariance structure is unfeasible, either for computational reasons, or due to a small sample size. However, inferential tools to verify this assumption are somewhat lacking in high-dimensional or functional {data analysis} settings, where this assumption is most relevant. We propose here to test separability by focusing on $K$-dimensional projections of the difference between the covariance operator and a nonparametric separable approximation. The subspace we project onto is one generated by the eigenfunctions of the covariance operator estimated under the separability hypothesis, negating the need to ever estimate the full non-separable covariance. We show that the rescaled difference of the sample covariance operator with its separable approximation is asymptotically Gaussian. As a by-product of this result, we derive asymptotically pivotal tests under Gaussian assumptions, and propose bootstrap methods for approximating the distribution of the test statistics. We probe the finite sample performance through simulations studies, and present an application to log-spectrogram images from a phonetic linguistics dataset.

stat.ME