SearcharxivSearch

arXiv subjects

Lassi Roininen

Publications and source records attributed to Lassi Roininen.

At least 19 recordsLinked to original sources

An unsupervised kernel norm monitoring for fault detection in a time series photovoltaic system

Grid-connected photovoltaic systems (GCPVS) are generally robust but remain susceptible to faults that can compromise energy conversion efficiency or raise safety concerns. Promptly and automatically detecting such anomalies is therefore essential for maintaining system reliability and performance. However, in practice, labeled fault data are rarely available in real-world deployments, which limits the applicability of supervised approaches. Conventional unsupervised baseline models, including a one-class support vector machine (OCSVM), isolation forest (iForest), and local outlier factor (LOF), are trained on normal operation data and assign anomaly scores reflecting how closely new observations resemble that baseline. Although these methods already accommodate non-linear behavior to varying degrees, kernel-based formulations offer further flexibility in shaping the decision boundary; however, tuning the kernel hyperparameters ordinarily requires some prior knowledge of the fault regime. We overcome this limitation by proposing kernel-based norm monitoring (KNM), a non-linear, unsupervised, window-based fault-detection method designed for continuous processes. Although the paper focuses on the GCPVS as a case study, KNM is a general-purpose monitoring framework applicable to a wide range of industrial processes. Using the Grid-connected PV System Faults (GPVS-Faults) dataset operating in intermediate power point tracking (IPPT) mode, KNM is evaluated in two fault scenarios, sensor faults and partial shading, against three benchmark techniques: OCSVM, iForest, and LOF. KNM achieves up to 99.1% and 98.3% accuracy on the two fault scenarios, respectively, using the Cauchy kernel, compared to 93.5% for the best-performing benchmark. The method is interpretable, and variable contribution plots are proposed to support fault identification.

stat.AP

Efficient Amortized Bayesian Inference for Markov Random Fields via Gradient-Informed Grid Selection

Bayesian inference for models with intractable likelihoods, such as Markov random fields, poses a fundamental computational challenge due to the tradeoff between inferential accuracy and computational cost. Various MCMC methods have been developed to address this challenge. The exchange algorithm targets the exact posterior, but requires an expensive perfect sampling step at each iteration, which is often infeasible in practice. In contrast, path sampling approximates the Metropolis acceptance ratio using a precomputed grid of likelihood values, but may introduce bias when the grid is poorly chosen. We introduce a novel amortized MCMC framework that retains the theoretical validity of exact methods while substantially reducing the computational burden. The proposed approach employs a gradient-informed grid selection procedure and constructs a surrogate likelihood via Hermite interpolation, yielding a smooth approximation with low error. A simulation study characterizes the rate at which inferential accuracy improves as the number of grid points increases. We further demonstrate the practical performance of the method through applications to a hidden Potts model for satellite imagery and an autologistic model for Arctic ice floes.

stat.ME

Identifiability and amortized inference limitations in Kuramoto models

Bayesian inference is a powerful tool for parameter estimation and uncertainty quantification in dynamical systems. However, for nonlinear oscillator networks such as Kuramoto models, widely used to study synchronization phenomena in physics, biology, and engineering, inference is often computationally prohibitive due to high-dimensional state spaces and intractable likelihood functions. We present an amortized Bayesian inference approach that learns a neural approximation of the posterior from simulated phase dynamics, enabling fast, scalable inference without repeated sampling or optimization. Applied to synthetic Kuramoto networks, the method shows promising results in approximating posterior distributions and capturing uncertainty, with computational savings compared to traditional Bayesian techniques. These findings suggest that amortized inference is a practical and flexible framework for uncertainty-aware analysis of oscillator networks.

stat.AP

Probabilistic multivariate statistical process control via kernel parameter uncertainty propagation

Kernel-based multivariate statistical process control (K-MSPC) extends classical monitoring to nonlinear industrial processes. Its performance depends critically on kernel parameters such as lengthscales and variance terms. In current practice these parameters are typically selected by heuristics or deterministic optimisation, and then treated as fixed, despite being inferred from finite and noisy data. This can lead to overconfident control limits and unstable alarm behaviour when the kernel choice is uncertain. This work proposes a probabilistic K-MSPC framework that quantifies and propagates kernel parameter uncertainty to the monitoring statistics. The approach follows a two-stage workflow: (i) deterministic kernel calibration using supervised or unsupervised models, and (ii) Bayesian inference of kernel parameters via Markov chain Monte Carlo. Posterior samples are propagated through kernel Principal Component Analysis to produce probabilistic $T^2$ and squarred prediction error control charts, together with uncertainty-aware contribution plots. The framework is evaluated on the Tennessee Eastman Process benchmark. Results show that posterior-mean monitoring often improves fault detection compared to deterministic prior-mean charts for the squared exponential kernel, while credible bands remain narrow in-control and widen under faults, reflecting amplified epistemic uncertainty in abnormal regimes. The automatic relevance determination kernel reduces posterior uncertainty and yields performance close to the deterministic baseline, whereas unsupervised calibration produces wider posterior bands but still robust fault detection.

stat.AP

Investigating the Electronic and Magnetic Properties of Na$_x$Fe$_{1/2}$Mn$_{1/2}$O$_2$ Cathode Materials with X-ray Compton Scattering

We discuss electronic and magnetic properties of Na$_x$Fe$_{1/2}$Mn$_{1/2}$O$_2$, a promising Na-ion battery cathode material. Using x-ray Compton scattering, SQUID magnetometry, and density-functional-theory based modeling, we probe how electrons and spins evolve during sodiation. By comparing Compton profiles of sodiated and desodiated samples, we show that oxygen 2$p$ orbitals drive the redox process, while transition-metal 3$d$ electrons become more delocalized, explaining the metallic phase at $x=2/3$. These profile differences define a quantitative descriptor for the sodiation range associated with improved conductivity. Electron holes on oxygen, reflected in oxygen magnetization, confirm the important role of oxygen in the electrochemical activity of the cathode.

cond-mat.mtrl-sci

Inhomogeneous Priors for Bayesian Inverse Problems

Many inverse problems arising in engineering and applied sciences involve unknown quantities with pronounced spatial inhomogeneity, such as localized defects or spatially varying material properties, making reliable uncertainty quantification particularly challenging. While Bayesian inverse problem methodologies provide a principled framework for assessing reconstruction reliability, commonly used Gaussian priors, such as Whittle-Matern models, impose globally homogeneous assumptions that limit their ability to capture such structure in large-scale settings. We introduce a new class of inhomogeneous priors defined via convolution with white noise, yielding nonstationary Whittle-Matern-type random fields with a rigorous mathematical construction. These priors fit naturally within existing Bayesian well-posedness theory and enable efficient sampling by reducing prior realizations to the solution of a pseudo-differential equation, for which we develop numerical schemes with quantified approximation error. Numerical experiments in one-dimensional denoising and two-dimensional limited-angle X-ray tomography demonstrate significant improvements in reconstruction quality and uncertainty quantification, particularly in data-limited scenarios.

math.NA

Spatiotemporal Dynamics of Conflict Occurrence and Fatalities in Ethiopia: A Bayesian Model and Predictive Insights Using Event-level Data (1997--2024)

This study presents a spatiotemporal dual Bayesian model that examines both the occurrence and number of conflict fatalities using event-level data from Ethiopia (1997-2024), sourced from the Armed Conflict Location and Event Data (ACLED) project. Fatalities are treated as two linked outcomes: the binary occurrence of deaths and the count of deaths when they occur. The model combines additive fixed effects for covariates with random effects capturing spatiotemporal influences, allowing for outcome-specific effects. Covariates include event type and season as categorical variables, proximity to cities and borders as nonlinear effects, and population as an offset term in the count model. A latent spatiotemporal process accounts for shared spatial and temporal dependence, with the spatial structure modeled using a Mat\'ern field prior and inference via Integrated Nested Laplace Approximation (INLA). Results show strong spatial clustering and temporal variation in fatality risk, emphasizing the importance of modeling both dimensions for better understanding and prediction. Airstrikes, shelling, and attacks show the highest fatality likelihood and counts, while communal and rebel actors cause the most deaths. Multiple fatalities are more likely in summer, and proximity to borders drives intense violence, whereas remoteness from urban centers is linked to lower-intensity events. These results provide insight for planning, policy, and resource allocation to protect vulnerable communities.

stat.AP

Computational design of personalized drugs via robust optimization under uncertainty

Effective disease treatment often requires precise control of the release of the active pharmaceutical ingredient (API). In this work, we present a computational inverse design approach to determine the optimal drug composition that yields a target release profile. We assume that the drug release is governed by the Noyes-Whitney model, meaning that dissolution occurs at the surface of the drug. Our inverse design method is based on topology optimization. The method optimizes the drug composition based on the target release profile, considering the drug material parameters and the shape of the final drug. Our method is non-parametric and applicable to arbitrary drug shapes. The inverse design method is complemented by robust topology optimization, which accounts for the random drug material parameters. We use the stochastic reduced-order method (SROM) to propagate the uncertainty in the dissolution model. Unlike Monte Carlo methods, SROM requires fewer samples and improves computational performance. We apply our method to designing drugs with several target release profiles. The numerical results indicate that the release profiles of the designed drugs closely resemble the target profiles. The SROM-based drug designs exhibit less uncertainty in their release profiles, suggesting that our method is a convincing approach for uncertainty-aware drug design.

cs.CE

Hierarchical Bayesian Modeling of Total Column Ozone: Unraveling Equatorial Variability over Ethiopia Using Satellite Data and Multisource Covariates

Understanding the spatiotemporal dynamics of total column ozone (TCO) is critical for monitoring ultraviolet (UV) exposure and ozone trends, particularly in equatorial regions where variability remains underexplored. This study investigates monthly TCO over Ethiopia (2012-2022) using a Bayesian hierarchical model implemented via Integrated Nested Laplace Approximation (INLA). The model incorporates nine environmental covariates, capturing meteorological, stratospheric, and topographic influences alongside spatiotemporal random effects. Spatial dependence is modeled using the Stochastic Partial Differential Equation (SPDE) approach, while temporal autocorrelation is handled through an autoregressive structure. The model shows strong predictive accuracy, with correlation coefficients of 0.94 (training) and 0.91 (validation), and RMSE values of 3.91 DU and 4.45 DU, respectively. Solar radiation, stratospheric temperature, and the Quasi-Biennial Oscillation are positively associated with TCO, whereas surface temperature, precipitation, humidity, water vapor, and altitude exhibit negative associations. Random effects highlight persistent regional clusters and seasonal peaks during summer. These findings provide new insights into regional ozone behavior over complex equatorial terrains, contributing to the understanding of the equatorial ozone paradox. The approach demonstrates the utility of combining satellite observations with environmental data in data-scarce regions, supporting improved UV risk monitoring and climate-informed policy planning.

stat.AP

Optimising Kernel-based Multivariate Statistical Process Control

Multivariate Statistical Process Control (MSPC) is a framework for monitoring and diagnosing complex processes by analysing the relationships between multiple process variables simultaneously. Kernel MSPC extends the methodology by leveraging kernel functions to capture non-linear relationships between the data, enhancing the process monitoring capabilities. However, optimising the kernel MSPC parameters, such as the kernel type and kernel parameters, is often done in literature in time-consuming and non-procedural manners such as cross-validation or grid search. In the present paper, we propose optimising the kernel MSPC parameters with Kernel Flows (KF), a recent kernel learning methodology introduced for Gaussian Process Regression (GPR). Apart from the optimisation technique, the novelty of the study resides also in the utilisation of kernel combinations for learning the optimal kernel type, and introduces individual kernel parameters for each variable. The proposed methodology is evaluated with multiple cases from the benchmark Tennessee Eastman Process. The faults are detected for all evaluated cases, including the ones not detected in the original study.

cs.CE

Partially stochastic deep learning with uncertainty quantification for model predictive heating control

Making the control of building heating systems more energy efficient is crucial for reducing global energy consumption and greenhouse gas emissions. Traditional rule-based control methods use a static, outdoor temperature-dependent heating curve to regulate heat input. This open-loop approach fails to account for both the current state of the system (indoor temperature) and free heat gains, such as solar radiation, often resulting in poor thermal comfort and overheating. Model Predictive Control (MPC) addresses these drawbacks by using predictive modeling to optimize heating based on a building's learned thermal behavior, current system state, and weather forecasts. However, current industrial MPC solutions often employ simplified physics-inspired indoor temperature models, sacrificing accuracy for robustness and interpretability. While purely data-driven models offer superior predictive performance and therefore more accurate control, they face challenges such as a lack of transparency. To bridge this gap, we propose a partially stochastic deep learning (DL) architecture, dubbed LSTM+BNN, for building-specific indoor temperature modeling. Unlike most studies that evaluate model performance through simulations or limited test buildings, our experiments across a comprehensive dataset of 100 real-world buildings, under various weather conditions, demonstrate that LSTM+BNN outperforms an industry-proven reference model, reducing the average prediction error measured as RMSE by more than 40% for the 48-hour prediction horizon of interest. Unlike deterministic DL approaches, LSTM+BNN offers a critical advantage by enabling pre-assessment of model competency for control optimization through uncertainty quantification. Thus, the proposed model shows significant potential to improve thermal comfort and energy efficiency achieved with heating MPC solutions.

stat.AP

Uncertainty quantification for electrical impedance tomography using quasi-Monte Carlo methods

The theoretical development of quasi-Monte Carlo (QMC) methods for uncertainty quantification of partial differential equations (PDEs) is typically centered around simplified model problems such as elliptic PDEs subject to homogeneous zero Dirichlet boundary conditions. In this paper, we present a theoretical treatment of the application of randomly shifted rank-1 lattice rules to electrical impedance tomography (EIT). EIT is an imaging modality, where the goal is to reconstruct the interior conductivity of an object based on electrode measurements of current and voltage taken at the boundary of the object. This is an inverse problem, which we tackle using the Bayesian statistical inversion paradigm. As the reconstruction, we consider QMC integration to approximate the unknown conductivity given current and voltage measurements. We prove under moderate assumptions placed on the parameterization of the unknown conductivity that the QMC approximation of the reconstructed estimate has a dimension-independent, faster-than-Monte Carlo cubature convergence rate. Finally, we present numerical results for examples computed using simulated measurement data.

math.NA

Log-Gaussian Cox Processes for Spatiotemporal Traffic Fatality Estimation in Addis Ababa

We investigate the spatiotemporal dynamics of traffic accidents in Addis Ababa, Ethiopia, using 2016--2019 data. We formulate the traffic accident intensity as a log-Gaussian Cox Process and model it as a spatiotemporal point process with and without fixed and random effect components that incorporate possible covariates and spatial correlation information. The covariate includes population density and distance of accident locations from schools, from markets, from bus stops and from worship places. We estimate the posterior of the state variables using integrated nested Laplace approximations with stochastic partial differential equations approach by considering Mat\`ern prior. Deviance and Watanabe - Akaike information criteria are used to check the performance of the models. We implement the methodology to map traffic accident intensity over Addis Ababa entirely and on its road networks and visualize the potential traffic accident hotspot areas. The comparison of the observation with the model output reveals that the covariates considered has significant effect for the accident intensity. Moreover, the information criteria results reveal the model with covariate performs well compared with the model without covariates. We obtained temporal correlation of the log-intensity as 0.78 indicating the existence of similar traffic fatality trend in space during the study period.

stat.AP

Statistical Batch-Based Bearing Fault Detection

In the domain of rotating machinery, bearings are vulnerable to different mechanical faults, including ball, inner, and outer race faults. Various techniques can be used in condition-based monitoring, from classical signal analysis to deep learning methods. Based on the complex working conditions of rotary machines, multivariate statistical process control charts such as Hotelling's $T^2$ and Squared Prediction Error are useful for providing early warnings. However, these methods are rarely applied to condition monitoring of rotating machinery due to the univariate nature of the datasets. In the present paper, we propose a multivariate statistical process control-based fault detection method that utilizes multivariate data composed of Fourier transform features extracted for fixed-time batches. Our approach makes use of the multidimensional nature of Fourier transform characteristics, which record more detailed information about the machine's status, in an effort to enhance early defect detection and diagnosis. Experiments with varying vibration measurement locations (Fan End, Drive End), fault types (ball, inner, and outer race faults), and motor loads (0-3 horsepower) are used to validate the suggested approach. The outcomes illustrate our method's effectiveness in fault detection and point to possible broader uses in industrial maintenance.

stat.ML

Bayesian inversion with Student's t priors based on Gaussian scale mixtures

Many inverse problems focus on recovering a quantity of interest that is a priori known to exhibit either discontinuous or smooth behavior. Within the Bayesian approach to inverse problems, such structural information can be encoded using Markov random field priors. We propose a class of priors that combine Markov random field structure with Student's t distribution. This approach offers flexibility in modeling diverse structural behaviors depending on available data. Flexibility is achieved by including the degrees of freedom parameter of Student's t distribution in the formulation of the Bayesian inverse problem. To facilitate posterior computations, we employ Gaussian scale mixture representation for the Student's t Markov random field prior, which allows expressing the prior as a conditionally Gaussian distribution depending on auxiliary hyperparameters. Adopting this representation, we can derive most of the posterior conditional distributions in a closed form and utilize the Gibbs sampler to explore the posterior. We illustrate the method with two numerical examples: signal deconvolution and image deblurring.

stat.CO

Geometry parameter estimation for sparse X-ray log imaging

We consider geometry parameter estimation in industrial sawmill fan-beam X-ray tomography. In such industrial settings, scanners do not always allow identification of the location of the source-detector pair, which creates the issue of unknown geometry. This work considers an approach for geometry estimation based on the calibration object. We parametrise the geometry using a set of 5 parameters. To estimate the geometry parameters, we calculate the maximum cross-correlation between a known-sized calibration object image and its filtered backprojection reconstruction and use differential evolution as an optimiser. The approach allows estimating geometry parameters from full-angle measurements as well as from sparse measurements. We show numerically that different sets of parameters can be used for artefact-free reconstruction. We deploy Bayesian inversion with first-order isotropic Cauchy difference priors for reconstruction of synthetic and real sawmill data with a very low number of measurements.

cs.CE

Reconstruction and segmentation from sparse sequential X-ray measurements of wood logs

In industrial applications, it is common to scan objects on a moving conveyor belt. If slice-wise 2D computed tomography (CT) measurements of the moving object are obtained we call it a sequential scanning geometry. In this case, each slice on its own does not carry sufficient information to reconstruct a useful tomographic image. Thus, here we propose the use of a Dimension reduced Kalman Filter to accumulate information between slices and allow for sufficiently accurate reconstructions for further assessment of the object. Additionally, we propose to use an unsupervised clustering approach known as Density Peak Advanced, to perform a segmentation and spot density anomalies in the internal structure of the reconstructed objects. We evaluate the method in a proof of concept study for the application of wood log scanning for the industrial sawing process, where the goal is to spot anomalies within the wood log to allow for optimal sawing patterns. Reconstruction and segmentation quality are evaluated from experimental measurement data for various scenarios of severely undersampled X-measurements. Results show clearly that an improvement in reconstruction quality can be obtained by employing the Dimension reduced Kalman Filter allowing to robustly obtain the segmented logs.

eess.SP

Log-Gaussian Gamma Processes for Training Bayesian Neural Networks in Raman and CARS Spectroscopies

We propose an approach utilizing gamma-distributed random variables, coupled with log-Gaussian modeling, to generate synthetic datasets suitable for training neural networks. This addresses the challenge of limited real observations in various applications. We apply this methodology to both Raman and coherent anti-Stokes Raman scattering (CARS) spectra, using experimental spectra to estimate gamma process parameters. Parameter estimation is performed using Markov chain Monte Carlo methods, yielding a full Bayesian posterior distribution for the model which can be sampled for synthetic data generation. Additionally, we model the additive and multiplicative background functions for Raman and CARS with Gaussian processes. We train two Bayesian neural networks to estimate parameters of the gamma process which can then be used to estimate the underlying Raman spectrum and simultaneously provide uncertainty through the estimation of parameters of a probability distribution. We apply the trained Bayesian neural networks to experimental Raman spectra of phthalocyanine blue, aniline black, naphthol red, and red 264 pigments and also to experimental CARS spectra of adenosine phosphate, fructose, glucose, and sucrose. The results agree with deterministic point estimates for the underlying Raman and CARS spectral signatures.

stat.AP