SearcharxivSearch

arXiv subjects

Jan Gertheiss

Publications and source records attributed to Jan Gertheiss.

At least 19 recordsLinked to original sources

Removal of Multivariate Environmental Influences in Structural Health Monitoring through Conditional Covariances and Supervised Learning

In structural health monitoring (SHM) systems, data is collected from a multitude of sensors measuring, for example, vibration or strain in the structure, along with additional features that capture environmental or operational information. It is well known that changes in the measured sensor outputs do not necessarily originate from structural damage but are often induced by environmental changes. One popular approach to account for these effects is regressing the system outputs on the confounding factors, also known as "response surface modeling". Afterward, the predicted values are subtracted from the observed ones to obtain corrected data with the environmental effects (supposedly) removed. However, the evaluation of real-world SHM data shows that environmental conditions may affect not only the expected output values but also higher-order statistical moments, particularly the variances of and the covariances and correlations between the output quantities, such as eigenfrequencies of different modes or strain sensors at different locations. By construction, the (supervised) machine learning techniques commonly used for response surface modeling cannot account for those higher-order effects. To address these issues, we present and discuss several approaches for identifying and quantifying multivariate confounding effects on output covariances and correlations: a nonparametric, kernel-based estimator, a random forest, a semiparametric additive model, and a deep learning approach. Furthermore, we show how the resulting conditional covariance matrices can be used in an SHM pipeline. We compare the competing methods on both artificial data and real-world load test data from the Vahrendorfer Stadtweg bridge in Hamburg, Germany, as well as eigenfrequency data from the railway bridge KW51 near Leuven, Belgium.

stat.ME

Feature Reconstruction and Monitoring of Load Test Data under Varying Environmental Conditions

System outputs in Structural Health Monitoring (SHM), such as sensor measurements or extracted features like eigenfrequencies, are influenced not only by (potential) damage but also by environmental and operational variables (EOV). Identifying these factors and removing their effects from the data is essential before proceeding with further analysis. Most existing methods for this task focus on the expected values of system outputs, e.g., using different types of response surface modeling. However, it has been shown that confounding variables can also affect the (co-)variance of and between system outputs. This is particularly important because the covariance matrix is an essential building block in many damage detection methods in SHM. Beyond standard response surface modeling, a nonparametric kernel approach can be used to estimate a conditional covariance matrix that can change depending on the identified confounding factor. This improves our understanding of how, e.g., temperature affects the system outputs. In this work, we present a new confounder-adjusted version of feature reconstruction. It uses the conditional covariance matrix as the basis for (conditional) principal component analysis. The resulting (conditional) principal component scores are then used to reconstruct system outputs with the confounding influences removed. In particular, the new approach eliminates the confounders effect on both the mean and the covariance. As will be shown on load test data from the Vahrendorfer Stadtweg bridge in Hamburg, Germany, the reconstructed features can then be employed for monitoring, e.g., using an appropriate control chart, resulting in fewer false alarms and a higher probability of detecting damage.

stat.AP

Higher-Order Multivariate Environmental Influences in Structural Health Monitoring

System outputs such as eigenfrequencies or strain data, often used in structural health monitoring (SHM), not only react to damage but also depend on environmental conditions. When trying to correct for these confounding effects, it is often (at least implicitly) assumed that only the expected, i.e., mean, output values are affected by environmental conditions. However, the evaluation of real-world SHM data indicates that environmental conditions may influence not only the mean output but also higher-order statistical moments, particularly the variances of and the covariances and correlations between the output quantities, such as eigenfrequencies of different modes or strain sensors at different locations. To address these issues, we discuss two approaches for identifying and quantifying multivariate confounding effects on output covariances and correlations: a random forest and a nonparametric, kernel-based approach. We compare the two competing methods on both artificial and real-world SHM data, finding that the kernel-based approach achieves higher accuracy, but the random forest produces estimates that are more robust and sometimes easier to interpret.

stat.AP

Covariate-Dependent Functional Principal Component Analysis for SHM

In Structural Health Monitoring (SHM), sensor measurements and derived features such as eigenfrequencies often exhibit systematic daily patterns and can therefore be naturally represented as functional data. Furthermore, these patterns are typically influenced by environmental factors, particularly temperature, which can substantially affect the observed system response. While most existing methods for removing environmental effects assume that confounding influences affect only the mean response, it has been shown that environmental and operational factors may also alter the covariance structure of the residual process. To address this limitation in a functional data monitoring framework, we incorporate so-called covariate-dependent functional principal component analysis (CD-FPCA), which allows eigenfunctions and eigenvalues of the residual process to vary smoothly with covariates such as temperature. The proposed methodology is illustrated using an extended version of the KW51 railway bridge eigenfrequency dataset. This case study suggests that accounting for covariate effects beyond the functional mean can improve the robustness of the monitoring procedure, in particular by reducing environmentally induced (false) alarms under challenging low-temperature conditions.

stat.ME

Confidence Intervals for Conditional Covariances of Natural Frequencies

In structural health monitoring (SHM), sensor measurements are collected, and damage-sensitive features such as natural frequencies are extracted for damage detection. However, these features depend not only on damage but are also influenced by various confounding factors, including environmental conditions and operational parameters. These factors must be identified, and their effects must be removed before further analysis. However, it has been shown that confounding variables may influence the mean and the covariance of the extracted features. This is particularly significant since the covariance is an essential building block in many damage detection tools. To account for the complex relationships resulting from the confounding factors, a nonparametric kernel approach can be used to estimate a conditional covariance matrix. By doing so, the covariance matrix is allowed to change depending on the identified confounding factor, thus providing a clearer understanding of how, for example, temperature influences the extracted features. This paper presents two bootstrap-based methods for obtaining confidence intervals for the conditional covariances, providing a way to quantify the uncertainty associated with the conditional covariance estimator. A proof-of-concept Monte Carlo study compares the two bootstrap versions proposed and evaluates their effectiveness. Finally, the methods are applied to the natural frequency data of the KW51 railway bridge near Leuven, Belgium. This real-world application highlights the practical implications of the findings. It underscores the importance of accurately accounting for confounding factors to generate more reliable diagnostic values with fewer false alarms.

stat.AP

Multivariate Long-term Profile Monitoring with Application to the KW51 Railway Bridge

Structural Health Monitoring (SHM) plays a pivotal role in modern civil engineering, providing critical insights into the health and integrity of infrastructure systems. This work presents a novel multivariate long-term profile monitoring approach to eliminate fluctuations in the measured response quantities, e.g., caused by environmental influences or measurement error. Our methodology addresses critical challenges in SHM and combines supervised methods with unsupervised, principal component analysis-based approaches in a single overarching framework, offering both flexibility and robustness in handling real-world large and/or sparse sensor data streams. We propose a function-on-function regression framework, which leverages functional data analysis for multivariate sensor data and integrates nonlinear modeling techniques, mitigating covariate-induced variations that can obscure structural changes.

stat.AP

Confounder-adjusted Covariances of System Outputs and Applications to Structural Health Monitoring

Automated damage detection is an integral component of each structural health monitoring (SHM) system. Typically, measurements from various sensors are collected and reduced to damage-sensitive features, and diagnostic values are generated by statistically evaluating the features. Since changes in data do not only result from damage, it is necessary to determine the confounding factors (environmental or operational variables) and to remove their effects from the measurements or features. Many existing methods for correcting confounding effects are based on different types of mean regression. This neglects potential changes in higher-order statistical moments, but in particular, the output covariances are essential for generating reliable diagnostics for damage detection. This article presents an approach to explicitly quantify the changes in the covariance, using conditional covariance matrices based on a non-parametric, kernel-based estimator. The method is applied to the Munich Test Bridge and the KW51 Railway Bridge in Leuven, covering both raw sensor measurements (acceleration, strain, inclination) and extracted damage-sensitive features (natural frequencies). The results show that covariances between different vibration or inclination sensors can significantly change due to temperature changes, and the same is true for natural frequencies. To highlight the advantages, it is explained how conditional covariances can be combined with standard approaches for damage detection, such as the Mahalanobis distance and principal component analysis. As a result, more reliable diagnostic values can be generated with fewer false alarms.

stat.AP

Covariate-Adjusted Functional Data Analysis for Structural Health Monitoring

Structural Health Monitoring (SHM) is increasingly applied in civil engineering. One of its primary purposes is detecting and assessing changes in structure conditions to increase safety and reduce potential maintenance downtime. Recent advancements, especially in sensor technology, facilitate data measurements, collection, and process automation, leading to large data streams. We propose a function-on-function regression framework for (nonlinear) modeling the sensor data and adjusting for covariate-induced variation. Our approach is particularly suited for long-term monitoring when several months or years of training data are available. It combines highly flexible yet interpretable semi-parametric modeling with functional principal component analysis and uses the corresponding out-of-sample Phase-II scores for monitoring. The method proposed can also be described as a combination of an ``input-output'' and an ``output-only'' method.

stat.AP

Regularization and Model Selection for Ordinal-on-Ordinal Regression with Applications to Food Products' Testing and Survey Data

Ordinal data are quite common in applied statistics. Although some model selection and regularization techniques for categorical predictors and ordinal response models have been developed over the past few years, less work has been done concerning ordinal-on-ordinal regression. Motivated by a consumer test and a survey on the willingness to pay for luxury food products consisting of Likert-type items, we propose a strategy for smoothing and selecting ordinally scaled predictors in the cumulative logit model. First, the group lasso is modified by the use of difference penalties on neighboring dummy coefficients, thus taking into account the predictors' ordinal structure. Second, a fused lasso-type penalty is presented for the fusion of predictor categories and factor selection. The performance of both approaches is evaluated in simulation studies and on real-world data.

stat.ME

Structural Health Monitoring with Functional Data: Two Case Studies

Structural Health Monitoring (SHM) is increasingly used in civil engineering. One of its main purposes is to detect and assess changes in infrastructure conditions to reduce possible maintenance downtime and increase safety. Ideally, this process should be automated and implemented in real-time. Recent advances in sensor technology facilitate data collection and process automation, resulting in massive data streams. Functional data analysis (FDA) can be used to model and aggregate the data obtained transparently and interpretably. In two real-world case studies of bridges in Germany and Belgium, this paper demonstrates how a function-on-function regression approach, combined with profile monitoring, can be applied to SHM data to adjust sensor/system outputs for environmental-induced variation and detect changes in construction. Specifically, we consider the R package \texttt{funcharts} and discuss some challenges when using this software on real-world SHM data. For instance, we show that pre-smoothing of the data can improve and extend its usability.

stat.AP

Functional Data Analysis: An Introduction and Recent Developments

Functional data analysis (FDA) is a statistical framework that allows for the analysis of curves, images, or functions on higher dimensional domains. The goals of FDA, such as descriptive analyses, classification, and regression, are generally the same as for statistical analyses of scalar-valued or multivariate data, but FDA brings additional challenges due to the high- and infinite dimensionality of observations and parameters, respectively. This paper provides an introduction to FDA, including a description of the most common statistical analysis techniques, their respective software implementations, and some recent developments in the field. The paper covers fundamental concepts such as descriptives and outliers, smoothing, amplitude and phase variation, and functional principal component analysis. It also discusses functional regression, statistical inference with functional data, functional classification and clustering, and machine learning approaches for functional data analysis. The methods discussed in this paper are widely applicable in fields such as medicine, biophysics, neuroscience, and chemistry, and are increasingly relevant due to the widespread use of technologies that allow for the collection of functional data. Sparse functional data methods are also relevant for longitudinal data analysis. All presented methods are demonstrated using available software in R by analyzing a data set on human motion and motor control. To facilitate the understanding of the methods, their implementation, and hands-on application, the code for these practical examples is made available on Github: https://github.com/davidruegamer/FDA_tutorial .

stat.ME

Nonparametric Estimation of the Underlying Distribution of Binned Continuous Data

The estimation of cumulative distribution functions (CDF) and probability density functions (PDF) is a fundamental practice in applied statistics. However, challenges often arise when dealing with data arranged in grouped intervals. In this paper, we discuss a suitable and highly flexible non-parametric density estimation approach for binned distributions, based on cubic monotonicity-preserving splines - known as cubic spline interpolation. Results from simulation studies demonstrate that this approach outperforms many widely used heuristic methods. Additionally, the application of this method to a dataset of train delays in Germany and micro census data on distance and travel time to work yields both meaningful but also some questionable results.

stat.ME

A Modification of McFadden's $R^2$ for Binary and Ordinal Response Models

A lot of studies on the summary measures of predictive strength of categorical response models consider the likelihood ratio index (LRI), also known as the McFadden-$R^2$, a better option than many other measures. We propose a simple modification of the LRI that adjusts for the effect of the number of response categories on the measure and that also rescales its values, mimicking an underlying latent measure. The modified measure is applicable to both binary and ordinal response models fitted by maximum likelihood. Results from simulation studies and a real data example on the olfactory perception of boar taint show that the proposed measure outperforms most of the widely used goodness-of-fit measures for binary and ordinal models. The proposed $R^2$ interestingly proves quite invariant to an increasing number of response categories of an ordinal model.

stat.ME

Penalized Optimal Scaling for Ordinal Variables with an Application to International Classification of Functioning Core Sets

Ordinal data occur frequently in the social sciences. When applying principal component analysis (PCA), however, those data are often treated as numeric implying linear relationships between the variables at hand, or non-linear PCA is applied where the obtained quantifications are sometimes hard to interpret. Non-linear PCA for categorical data, also called optimal scoring/scaling, constructs new variables by assigning numerical values to categories such that the proportion of variance in those new variables that is explained by a predefined number of principal components is maximized. We propose a penalized version of non-linear PCA for ordinal variables that is a smoothed intermediate between standard PCA on category labels and non-linear PCA as used so far. The new approach is by no means limited to monotonic effects and offers both better interpretability of the non-linear transformation of the category labels as well as better performance on validation data than unpenalized non-linear PCA and/or standard linear PCA. In particular, an application of penalized optimal scaling to ordinal data as given with the International Classification of Functioning, Disability and Health (ICF) is provided.

stat.AP

Nonparametric Regression and Classification with Functional, Categorical, and Mixed Covariates

We consider nonparametric prediction with multiple covariates, in particular categorical or functional predictors, or a mixture of both. The method proposed bases on an extension of the Nadaraya-Watson estimator where a kernel function is applied on a linear combination of distance measures each calculated on single covariates, with weights being estimated from the training data. The dependent variable can be categorical (binary or multi-class) or continuous, thus we consider both classification and regression problems. The methodology presented is illustrated and evaluated on artificial and real world data. Particularly it is observed that prediction accuracy can be increased, and irrelevant, noise variables can be identified/removed by `downgrading' the corresponding distance measures in a completely data-driven way.

stat.ME

Ride Sharing & Data Privacy: An Analysis of the State of Practice

Digital services like ride sharing rely heavily on personal data as individuals have to disclose personal information in order to gain access to the market and exchange their information with other participants; yet, the service provider usually gives little to no information regarding the privacy status of the disclosed information though privacy concerns are a decisive factor for individuals to (not) use these services. We analyzed how popular ride sharing services handle user privacy to assess the current state of practice. The results show that services include a varying set of personal data and offer limited privacy-related features.

cs.CY

Statistical Inference for Ordinal Predictors in Generalized Linear and Additive Models with Application to Bronchopulmonary Dysplasia

Discrete but ordered covariates are quite common in applied statistics, and some regularized fitting procedures have been proposed for proper handling of ordinal predictors in statistical modeling. In this study, we show how quadratic penalties on adjacent dummy coefficients of ordinal predictors proposed in the literature can be incorporated in the framework of generalized additive models, making tools for statistical inference developed there available for ordinal predictors as well. Motivated by an application from neonatal medicine, we discuss whether results obtained when constructing confidence intervals and testing significance of smooth terms in generalized additive models are useful with ordinal predictors/penalties as well.

stat.ME

Generalized Functional Additive Mixed Models

We propose a comprehensive framework for additive regression models for non-Gaussian functional responses, allowing for multiple (partially) nested or crossed functional random effects with flexible correlation structures for, e.g., spatial, temporal, or longitudinal functional data as well as linear and nonlinear effects of functional and scalar covariates that may vary smoothly over the index of the functional response. Our implementation handles functional responses from any exponential family distribution as well as many others like Beta- or scaled non-central $t$-distributions. Development is motivated by and evaluated on an application to large-scale longitudinal feeding records of pigs. Results in extensive simulation studies as well as replications of two previously published simulation studies for generalized functional mixed models demonstrate the good performance of our proposal. The approach is implemented in well-documented open source software in the "pffr()" function in R-package "refund".

stat.ME