SearcharxivSearch

arXiv subjects

Shahin Tavakoli

Publications and source records attributed to Shahin Tavakoli.

11 recordsLinked to original sources

On the existence and non-existence of centres of mass on Hilbert spheres

Fréchet means and $L^p$ centres of mass provide notions of average location in metric spaces. On finite-dimensional spheres, existence follows from compactness. On infinite-dimensional spheres, it is not known whether a centre of mass always exists. We show that this is not always the case, and give a simple assumption under which a centre of mass exists. We then show that finding the sample centre of mass of data $x_1, \ldots, x_n$ on the sphere is always an optimisation problem on a subsphere of manifold dimension at most $n$, regardless of the potentially infinite dimension of the sphere. We conclude with some statistical implications.

math.ST

Modeling Time-course Gene Expression Data through Bayesian Partition Functional Principal Component Analysis

High-dimensional biomarkers such as gene expression levels are now routinely measured over time, allowing biological processes to be studied dynamically rather than through cross-sectional snapshots. However, existing methods do not adequately address the central applied challenges posed by such data: simultaneously reducing dimensionality, quantifying inter-individual variability and uncovering temporal structure shared across biomarkers. We introduce Partition Functional Principal Component Analysis (PFPCA), a Bayesian model that jointly learns shared temporal patterns and clusters variables according to their latent dynamics. PFPCA combines a mixture model with multivariate functional principal component analysis performed within each group. We develop a scalable mean-field variational algorithm for joint inference of functional principal component loadings, individual-level scores, group assignments and partition sizes. Simulations show clear gains from joint inference: PFPCA recovers both the partition and the latent functional structure more accurately than a two-step baseline. In the most challenging settings, PFPCA retrieves the true partition in 27% of replicates compared with 1% for the two-step baseline. Applied to longitudinal gene-expression data from individuals experimentally infected with H3N2 influenza virus, PFPCA identifies groups of genes with coordinated activation patterns and reveals temporal signatures associated with immune-response dynamics and symptom status.

stat.ME

Infinite-Dimensional Spherical Kernel ridge Regression

We introduce a novel regression framework designed to model non-linear responses situated on a sphere $\mathbb{S}$ of finite or infinite dimension. Unlike traditional tangent-space regressions, which lift responses to a tangent space $T_o \mathbb{S}$ and thereby violate intrinsic spherical distances, our proposed method employs an intrinsic approach. We model the conditional mean through an intercept $o \in \mathbb{S}$ and a linear predictor function $f: \mathfrak{X} \to T_o \mathbb{S}$. This formulation transforms the estimation problem into finding a linear predictor within a function space, but utilizing a metric defined by spherical geometry rather than standard Euclidean distance. Leveraging vector-valued reproducing kernel Hilbert space theory, our approach reduces the infinite-dimensional estimation challenge to a manageable finite-dimensional problem via the representer theorem, leading to an efficient BFGS-based estimation algorithm. We establish convergence rates and analyze the finite-sample behavior of our estimator, concluding with a practical application to density regression. The full implementation is available in R.

stat.ME

Normative brain mapping of 3-dimensional morphometry imaging data using skewed functional data analysis

Tensor-based morphometry (TBM) aims at showing local differences in brain volumes with respect to a common template. TBM images are smooth but they exhibit (especially in diseased groups) higher values in some brain regions called lateral ventricles. More specifically, our voxelwise analysis shows both a mean-variance relationship in these areas and evidence of spatially dependent skewness. We propose a model for 3-dimensional functional data where mean, variance, and skewness functions vary smoothly across brain locations. We model the voxelwise distributions as skew-normal. The smooth effects of age and sex are estimated on a reference population of cognitively normal subjects from the Alzheimer's Disease Neuroimaging Initiative (ADNI) dataset and mapped across the whole brain. The three parameter functions allow to transform each TBM image (in the reference population as well as in a test set) into a Gaussian process. These subject-specific normative maps are used to derive indices of deviation from a healthy condition to assess the individual risk of pathological degeneration.

stat.AP

Statistics in Phonetics

Phonetics is the scientific field concerned with the study of how speech is produced, heard and perceived. It abounds with data, such as acoustic speech recordings, neuroimaging data, or articulatory data. In this paper, we provide an introduction to different areas of phonetics (acoustic phonetics, sociophonetics, speech perception, articulatory phonetics, speech inversion, sound change, and speech technology), an overview of the statistical methods for analyzing their data, and an introduction to the signal processing methods commonly applied to speech recordings. A major transition in the statistical modeling of phonetic data has been the shift from fixed effects to random effects regression models, the modeling of curve data (for instance via GAMMs or FDA methods), and the use of Bayesian methods. This shift has been driven in part by the increased focus on large speech corpora in phonetics, which has been driven by machine learning methods such as forced alignment. We conclude by identifying opportunities for future research.

stat.AP

Exploring British Accents: Modelling the Trap-Bath Split with Functional Data Analysis

The sound of our speech is influenced by the places we come from. Great Britain contains a wide variety of distinctive accents which are of interest to linguistics. In particular, the "a" vowel in words like "class" is pronounced differently in the North and the South. Speech recordings of this vowel can be represented as formant curves or as Mel-frequency cepstral coefficient curves. Functional data analysis and generalized additive models offer techniques to model the variation in these curves. Our first aim is to model the difference between typical Northern and Southern vowels /ae/ and /a/, by training two classifiers on the North-South Class Vowels dataset collected for this paper (Koshy 2020). Our second aim is to visualize geographical variation of accents in Great Britain. For this we use speech recordings from a second dataset, the British National Corpus (BNC) audio edition (Coleman et al. 2012). The trained models are used to predict the accent of speakers in the BNC, and then we model the geographical patterns in these predictions using a soap film smoother. This work demonstrates a flexible and interpretable approach to modeling phonetic accent variation in speech recordings.

cs.SD

Functional regression clustering with multiple functional gene expressions

Gene expression data is often collected in time series experiments, under different experimental conditions. There may be genes that have very different gene expression profiles over time, but that adjust their gene expression patterns in the same way under experimental conditions. Our aim is to develop a method that finds clusters of genes in which the relationship between these temporal gene expression profiles are similar to one another, even if the individual temporal gene expression profiles differ. We propose a $K$-means type algorithm in which each cluster is defined by a function-on-function regression model, which, inter alia, allows for multiple functional explanatory variables. We validate this novel approach through extensive simulations and then apply it to identify groups of genes whose diurnal expression pattern is perturbed by the season in a similar way. Our clusters are enriched for genes with similar biological functions, including one cluster enriched in both photosynthesis-related functions and polysomal ribosomes, which shows that our method provides useful and novel biological insights.

stat.ME

Factor Models for High-Dimensional Functional Time Series

In this paper, we set up the theoretical foundations for a high-dimensional functional factor model approach in the analysis of large cross-sections (panels) of functional time series (FTS). We first establish a representation result stating that, under mild assumptions on the covariance operator of the cross-section, we can represent each FTS as the sum of a common component driven by scalar factors loaded via functional loadings, and a mildly cross-correlated idiosyncratic component. Our model and theory are developed in a general Hilbert space setting that allows for mixed panels of functional and scalar time series. We then turn to the identification of the number of factors, and the estimation of the factors, their loadings, and the common components. We provide a family of information criteria for identifying the number of factors, and prove their consistency. We provide average error bounds for the estimators of the factors, loadings, and common component; our results encompass the scalar case, for which they reproduce and extend, under weaker conditions, well-established similar results. Under slightly stronger assumptions, we also provide uniform bounds for the estimators of factors, loadings, and common component, thus extending existing scalar results. Our consistency results in the asymptotic regime where the number $N$ of series and the number $T$ of time observations diverge thus extend to the functional context the "blessing of dimensionality" that explains the success of factor models in the analysis of high-dimensional (scalar) time series. We provide numerical illustrations that corroborate the convergence rates predicted by the theory, and provide finer understanding of the interplay between $N$ and $T$ for estimation purposes. We conclude with an application to forecasting mortality curves, where we demonstrate that our approach outperforms existing methods.

math.ST

A Spatial Modeling Approach for Linguistic Object Data: Analysing dialect sound variations across Great Britain

Dialect variation is of considerable interest in linguistics and other social sciences. However, traditionally it has been studied using proxies (transcriptions) rather than acoustic recordings directly. We introduce novel statistical techniques to analyse geolocalised speech recordings and to explore the spatial variation of pronunciations continuously over the region of interest, as opposed to traditional isoglosses, which provide a discrete partition of the region. Data of this type require an explicit modeling of the variation in the mean and the covariance. Usual Euclidean metrics are not appropriate, and we therefore introduce the concept of $d$-covariance, which allows consistent estimation both in space and at individual locations. We then propose spatial smoothing for these objects which accounts for the possibly non convex geometry of the domain of interest. We apply the proposed method to data from the spoken part of the British National Corpus, deposited at the British Library, London, and we produce maps of the dialect variation over Great Britain. In addition, the methods allow for acoustic reconstruction across the domain of interest, allowing researchers to listen to the statistical analysis.

stat.ME

Tests for separability in nonparametric covariance operators of random surfaces

The assumption of separability of the covariance operator for a random image or hypersurface can be of substantial use in applications, especially in situations where the accurate estimation of the full covariance structure is unfeasible, either for computational reasons, or due to a small sample size. However, inferential tools to verify this assumption are somewhat lacking in high-dimensional or functional {data analysis} settings, where this assumption is most relevant. We propose here to test separability by focusing on $K$-dimensional projections of the difference between the covariance operator and a nonparametric separable approximation. The subspace we project onto is one generated by the eigenfunctions of the covariance operator estimated under the separability hypothesis, negating the need to ever estimate the full non-separable covariance. We show that the rescaled difference of the sample covariance operator with its separable approximation is asymptotically Gaussian. As a by-product of this result, we derive asymptotically pivotal tests under Gaussian assumptions, and propose bootstrap methods for approximating the distribution of the test statistics. We probe the finite sample performance through simulations studies, and present an application to log-spectrogram images from a phonetic linguistics dataset.

stat.ME

Fourier analysis of stationary time series in function space

We develop the basic building blocks of a frequency domain framework for drawing statistical inferences on the second-order structure of a stationary sequence of functional data. The key element in such a context is the spectral density operator, which generalises the notion of a spectral density matrix to the functional setting, and characterises the second-order dynamics of the process. Our main tool is the functional Discrete Fourier Transform (fDFT). We derive an asymptotic Gaussian representation of the fDFT, thus allowing the transformation of the original collection of dependent random functions into a collection of approximately independent complex-valued Gaussian random functions. Our results are then employed in order to construct estimators of the spectral density operator based on smoothed versions of the periodogram kernel, the functional generalisation of the periodogram matrix. The consistency and asymptotic law of these estimators are studied in detail. As immediate consequences, we obtain central limit theorems for the mean and the long-run covariance operator of a stationary functional time series. Our results do not depend on structural modelling assumptions, but only functional versions of classical cumulant mixing conditions, and are shown to be stable under discrete observation of the individual curves.

math.ST