Searcharxiv⌕ Search

arXiv subjects

Sophie Dabo-Niang

Publications and source records attributed to Sophie Dabo-Niang.

At least 19 recordsLinked to original sources

Variable-Projection Sparse Functional Principal Component Analysis: Interpretable Functional Dimensionality Reduction with Applications to Raman Spectral Data

Functional principal component analysis (FPCA) provides low-rank representations of functional data but generally produces dense components, making it difficult to identify the localised regions contributing to dominant modes of variation. This limitation is particularly relevant in Raman spectroscopy, where spectra are observed over an ordered domain and interpretation often focuses on chemically meaningful spectral regions. This study proposes variable-projection sparse FPCA (VP-SFPCA) for interpretable functional dimensionality reduction through sparse weight functions that promote localisation. The method formulates sparse FPCA as a regularised matrix-factorisation problem that incorporates the functional inner-product geometry and distinguishes sparse weight functions used to generate component scores from orthonormal loading functions used for reconstruction. Variable projection conditionally minimises over the loading functions, reducing the optimisation problem to the sparse weights. Performance was evaluated through simulation studies and empirical analyses of surface-enhanced Raman scattering (SERS) spectra, with conventional FPCA and SCAD-SFPCA serving as dense and sparse functional benchmarks, respectively. In the simulations, VP-SFPCA recovered localised functional structure while requiring substantially less computation than SCAD-SFPCA. In the empirical analysis, VP-SFPCA retained this computational advantage while yielding held-out reconstruction error close to that of conventional FPCA. Several prominent features of the estimated weight functions also coincided with established adenine SERS bands. Overall, VP-SFPCA provides a computationally practical approach to improving the interpretability of dominant functional modes through localisation.

stat.ME↗

Neighbourhood-Based Generalized Dynamic Principal Components for Spatial Functional Data

Regular-grid spatial functional datasets arise naturally in gridded environmental, oceanographic, climate, and remote-sensing applications, where each spatial location is associated with an entire curve. Existing spectral spatial functional principal component analysis methods provide an important frequency-domain approach for such data, but they do not directly target finite-neighbourhood least-squares reconstruction from an estimated latent spatial component field. To address this, we propose Spatial Functional Generalized Dynamic Principal Components (SFGDPC), a reconstruction-based dimension-reduction method for regular-grid spatial functional data. Each function is first represented by basis coefficients, and each coefficient vector is reconstructed from a scalar latent spatial field and its Chebyshev neighbourhood on the rectangular grid. The spatial neighbourhood radius is selected using a Bayesian information criterion (BIC)-type conditional reconstruction criterion. In the local-neighbourhood transfer simulation, SFGDPC reduced mean cumulative normalized mean squared error (NMSE) relative to spatial functional principal component analysis (SFPCA) by approximately 38-53% across the reported component counts and covariance conditions. In the Indian Ocean sea surface temperature (SST) application, one SFGDPC component produced lower whole-grid reconstruction error than both the boundary-safe and high-cap SFPCA benchmarks across all 33 annual fields. The results support SFGDPC as a local reconstruction-based complement to spectral SFPCA for regular-grid spatial functional data.

stat.ME↗

Spatial Principal Component Analysis and Moran Statistics for Multivariate Functional Areal Data

The paper introduces a multivariate functional areal spatial principal component analysis (mfasPCA) framework, together with multivariate functional Moran's I statistics, to enable the assessment of spatial autocorrelation and dimension reduction for multivariate functional data observed over areal units. The proposed framework is spatial-functional in scope: the functional argument may represent time, age, wavelength, or another ordered continuum, while spatial dependence is introduced across areal units through a spatial weight matrix. The principal component method is defined through a Moran-type spatially weighted criterion. We propose eigenvalue-based permutation tests to assess the significance of spatially structured components. The testing framework includes omnibus tests, componentwise tests with Holm adjustment, and sequential rank-wise tests based on tail sums of eigenvalues. Simulation studies show that mfasPCA captures positive and negative spatial-functional structures and concentrates them in the leading components under the respective autocorrelation regimes. A real-data application illustrates how mfasPCA identifies spatially structured modes of multivariate functional variation.

stat.ME↗

A spatial scan statistical for categorical, functional data

We have developed and tested a spatial scan statistic for categorical, functional data (CFSS) - a data structure within which current approaches cannot identify spatial clusters. Our methodology combines an encoding scheme for categorical, functional observations with a nonparametric scan statistic. In a simulation study with three distinct scenarios, the CFSS accurately recovered the simulated spatial clusters and gave very low false positive rates, high true positive rates, and high positive predictive values. We have also used the CFSS to identify and characterize spatial clusters in French air pollution data from the winter of 2024.

stat.ME↗

Generalized dynamic functional principal component analysis

In this paper, we explore dimension reduction for functional time series. We propose a generalized dynamic functional principal component analysis (GDFPCA) which does not rely on spectral density estimation and demonstrates strong empirical performance for both stationary and nonstationary functional time series. We define the generalized dynamic functional principal components (GDFPCs) as static factor time series in a functional dynamic factor model and obtain their multivariate representation from a truncation of the functional dynamic factor model. Estimation is based on a least-squares reconstruction criterion and implemented via a two-step procedure for the coefficient vectors of the loading curves under a basis expansion. We establish mean-square consistency of the reconstructed functional time series under weak stationarity. Simulation studies show that GDFPCA performs comparably to dynamic functional principal component analysis (DFPCA) for stationary data, while providing improved reconstruction accuracy in nonstationary settings, where both DFPCA and functional principal component analysis (FPCA) deteriorate. Applications to real datasets support the empirical advantages observed in the simulations.

stat.ME↗

Compositional data analysis for modelling and forecasting mortality using the α-transformation

Mortality forecasting is crucial for demographic planning and actuarial studies, especially for projecting population ageing and longevity risk. Classical approaches largely rely on extrapolative methods, such as the Lee-Carter (LC) model, which use mortality rates as the mortality measure. In recent years, compositional data analysis (CoDA), which respects summability and non-negativity constraints, has gained increasing attention for mortality forecasting. While the centred log-ratio (CLR) transformation is commonly used to map compositional data to real space, the α-transformation, a generalisation of log-ratio transformations, offers greater flexibility and adaptability. This study contributes to mortality forecasting by introducing the α-transformation as an alternative to the CLR transformation within a non-functional CoDA model that has not been previously investigated in existing literature. To fairly compare the impact of transformation choices on forecast accuracy, zero values in the data are imputed, although the α-transformation can inherently handle them. Using age-specific life table death counts for males and females in 31 selected European countries/regions from 1983 to 2018, the proposed method demonstrates comparable performance to the CLR transformation in most cases, with improved forecast accuracy in some instances. These findings highlight the potential of the α-transformation for enhancing mortality forecasting within the non-functional CoDA framework.

stat.AP↗

Reframing Three-Dimensional Morphometrics Through Functional Data Innovations

This study innovates geometric morphometrics by incorporating functional data analysis, the square-root velocity function (SRVF), and arc-length parameterisation for 3D morphometric data, leading to the development of seven new pipelines in addition to the standard geometric morphometrics (GM) approach.. This enables three-dimensional images to be examined from perspectives that do not neglect curvature, through the combined use of arc-length parameterisation, soft-alignment, and elastic-alignment. A simulation study was conducted to demonstrate the general effectiveness of eight pipelines: geometric morphometrics (GM, baseline), arc-GM, functional data morphometrics (FDM), arc-FDM, soft-SRV-FDM, arc-soft-SRV-FDM, elastic-SRV-FDM, and arc-elastic-SRV-FDM. These pipelines were also applied to distinguish dietary categories of kangaroos (omnivores, mixed feeders, browsers, and grazers) using cranial landmarks obtained from 41 extant species. Principal component analysis was conducted, followed by classification analysis using linear discriminant analysis, multinomial regression and support vector machines with a linear kernel. The results highlight the effectiveness of functional data analysis, together with arc-length and SRVF-based approaches, in opening the door to more robust perspectives for analysing three-dimensional morphometrics, while establishing geometric morphometrics as the baseline for comparison.

stat.AP↗

Functional regression with randomized signatures: An application to age-specific mortality rates

We propose a novel extension of the Hyndman-Ullah (HU) model to forecast mortality rates by integrating randomized signatures, referred to as the HU model with randomized signatures (HUrs). Unlike truncated signatures, which grow exponentially with order, randomized signatures, based on the Johnson-Lindenstrauss lemma, are able to approximate higher-order interactions in a computationally feasible way. Using mortality data from four countries, we evaluate the performance of the novel HUrs model compared to two alternative HU model versions. Our empirical results show that the proposed HUrs model performs well, particularly for Bulgarian and Japanese data.

stat.ME↗

Forecasting mortality rates with functional signatures

This study introduces an innovative methodology for mortality forecasting, which integrates signature-based methods within the functional data framework of the Hyndman-Ullah (HU) model. This new approach, termed the Hyndman-Ullah with truncated signatures (HUts) model, aims to enhance the accuracy and robustness of mortality predictions. By utilizing signature regression, the HUts model is able to capture complex, nonlinear dependencies in mortality data which enhances forecasting accuracy across various demographic conditions. The model is applied to mortality data from 12 countries, comparing its forecasting performance against variants of the HU models across multiple forecast horizons. Our findings indicate that overall the HUts model not only provides more precise point forecasts but also shows robustness against data irregularities, such as those observed in countries with historical outliers. The integration of signature-based methods enables the HUts model to capture complex patterns in mortality data, making it a powerful tool for actuaries and demographers. Prediction intervals are also constructed with bootstrapping methods

stat.ME↗

Fusion regression methods with repeated functional data

Linear regression and classification methods with repeated functional data are considered. For each statistical unit in the sample, a real-valued parameter is observed over time under different conditions related by some neighborhood structure (spatial, group, etc.). Two regression methods based on fusion penalties are proposed to consider the dependence induced by this structure. These methods aim to obtain parsimonious coefficient regression functions, by determining if close conditions are associated with common regression coefficient functions. The first method is a generalization to functional data of the variable fusion methodology based on the 1-nearest neighbor. The second one relies on the group fusion lasso penalty which assumes some grouping structure of conditions and allows for homogeneity among the regression coefficient functions within groups. Numerical simulations and an application of electroencephalography data are presented.

stat.ME↗

Classification of multivariate functional data on different domains with Partial Least Squares approaches

Classification (supervised-learning) of multivariate functional data is considered when the elements of the random functional vector of interest are defined on different domains. In this setting, PLS classification and tree PLS-based methods for multivariate functional data are presented. From a computational point of view, we show that the PLS components of the regression with multivariate functional data can be obtained using only the PLS methodology with univariate functional data. This offers an alternative way to present the PLS algorithm for multivariate functional data.

stat.ME↗

Uncovering Data Across Continua: An Introduction to Functional Data Analysis

In a world increasingly awash with data, the need to extract meaningful insights from data has never been more crucial. Functional Data Analysis (FDA) goes beyond traditional data points, treating data as dynamic, continuous functions, capturing ever-changing phenomena nuances. This article introduces FDA, merging statistics with real-world complexity, ideal for those with mathematical skills but no FDA background.

math.ST↗

A Markov-switching spatio-temporal ARCH model

Stock market indices are volatile by nature, and sudden shocks are known to affect volatility patterns. The autoregressive conditional heteroskedasticity (ARCH) and generalized ARCH (GARCH) models neglect structural breaks triggered by sudden shocks that may lead to an overestimation of persistence, causing an upward bias in the estimates. Different regime-switching models that have abrupt regime-switching governed by a Markov chain were developed to model volatility in financial time series data. Volatility modelling was also extended to spatially interconnected time series, resulting in spatial variants of ARCH models. This inspired us to propose a Markov switching framework of the spatio-temporal log-ARCH model. In this article, we discuss the Markov-switching extension of the model, the estimation procedure and the smooth inferences of the regimes. The Monte-Carlo simulation studies show that the maximum likelihood estimation method for our proposed model has good finite sample properties. The proposed model was applied to 28 stock indices data that were presumably affected by the 2015-2016 Chinese stock market crash. The results showed that our model is a better fit compared to that of the one-regime counterpart. Furthermore, the smoothed inference of the data indicated the approximate periods where structural breaks occurred. This model can capture structural breaks that simultaneously occur in nearby locations.

stat.ME↗

A unified approach for morphometrics and functional data analysis with machine learning for craniodental shape quantification in shrew species

This work proposes a functional data analysis approach for morphometrics with applications in classifying three shrew species (S. murinus, C. monticola and C. malayana) based on the images. The discrete landmark data of craniodental views (dorsal, jaw and lateral) are converted into continuous curves where the curves are represented as linear combinations of basis functions. A comparative study based on four machine learning algorithms such as naive Bayes, support vector machine, random forest, and generalized linear models was conducted on the predicted principal component scores obtained from the FDA approach and classical approach (combination of all three craniodental views and individual views). The FDA approach produced better results in separating the three clusters of shrew species compared to the classical method and the dorsal view gave the best representation in classifying the three shrew species. Overall, based on the FDA approach, GLM of the predicted PCA scores was the most accurate (95.4% accuracy) among the four classification models.

q-bio.QM↗

k-nearest neighbors prediction and classification for spatial data

This paper proposes a spatial k-nearest neighbor method for nonparametric prediction of real-valued spatial data and supervised classification for categorical spatial data. The proposed method is based on a double nearest neighbor rule which combines two kernels to control the distances between observations and locations. It uses a random bandwidth in order to more appropriately fit the distributions of the covariates. The almost complete convergence with rate of the proposed predictor is established and the almost sure convergence of the supervised classification rule was deduced. Finite sample properties are given for two applications of the k-nearest neighbor prediction and classification rule to the soil and the fisheries datasets

math.ST↗

A Bayesian shared-frailty spatial scan statistic model for time-to-event data

Spatial scan statistics are well known and widely used methods for the detection of spatial clusters of events. In the field of spatial analysis of time-to-event data, several models of scan statistics have been proposed. However, these models do not take into account the potential intra-unit spatial correlation of individuals nor a potential correlation between spatial units. To overcome this problem, we propose here a scan statistic based on a Cox model with shared frailty that takes into account the spatial correlation between spatial units. In simulation studies, we have shown that (i) classical models of spatial scan statistics for time-to-event data fail to maintain the type I error in the presence of intra-spatial unit correlation, and (ii) our model performs well in the presence of both intra-spatial unit correlation and inter-spatial unit correlation. Our method has been applied to epidemiological data and to the detection of spatial clusters of mortality in patients with end-stage renal disease in northern France.

stat.ME↗

Multivariate nonparametric regression by least squares Jacobi polynomials approximations

In this work, we study a random orthogonal projection based least squares estimator for the stable solution of a multivariate nonparametric regression (MNPR) problem. More precisely, given an integer $d\geq 1$ corresponding to the dimension of the MNPR problem, a positive integer $N\geq 1$ and a real parameter $α\geq -\frac{1}{2},$ we show that a fairly large class of $d-$variate regression functions are well and stably approximated by its random projection over the orthonormal set of tensor product $d-$variate Jacobi polynomials with parameters $(α,α).$ The associated uni-variate Jacobi polynomials have degree at most $N$ and their tensor products are orthonormal over $\mathcal U=[0,1]^d,$ with respect to the associated multivariate Jacobi weights. In particular, if we consider $n$ random sampling points $\mathbf X_i$ following the $d-$variate Beta distribution, with parameters $(α+1,α+1),$ then we give a relation involving $n, N, α$ to ensure that the resulting $(N+1)^d\times (N+1)^d$ random projection matrix is well conditioned. Moreover, we provide squared integrated as well as $L^2-$risk errors of this estimator. Precise estimates of these errors are given in the case where the regression function belongs to an isotropic Sobolev space $H^s(I^d),$ with $s> \frac{d}{2}.$ Also, to handle the general and practical case of an unknown distribution of the $\mathbf X_i,$ we use Shepard's scattered interpolation scheme in order to generate fairly precise approximations of the observed data at $n$ i.i.d. sampling points $\mathbf X_i$ following a $d-$variate Beta distribution. Finally, we illustrate the performance of our proposed multivariate nonparametric estimator by some numerical simulations with synthetic as well as real data.

math.ST↗

Investigating spatial scan statistics for multivariate functional data

This paper introduces new scan statistics for multivariate functional data indexed in space. The new methods are derivated from a MANOVA test statistic for functional data, an adaptation of the Hotelling T2-test statistic, and a multivariate extension of the Wilcoxon rank-sum test statistic. In a simulation study, the latter two methods present very good performances and the adaptation of the functional MANOVA also shows good performances for a normal distribution. Our methods detect more accurate spatial clusters than an existing nonparametric functional scan statistic. Lastly we applied the methods on multivariate functional data to search for spatial clusters of abnormal daily concentrations of air pollutants in the north of France in May and June 2020.

stat.ME↗