SearcharxivSearch

arXiv subjects

Jörg Martin

Publications and source records attributed to Jörg Martin.

13 recordsLinked to original sources

Low Rank Based Subspace Inference for the Laplace Approximation of Bayesian Neural Networks

Subspace inference for neural networks assumes that a subspace of their parameter space suffices to produce a reliable uncertainty quantification. In this work, we underpin the validity of this assumption by using low rank techniques. We derive an expression for a subspace model to a Bayesian inference scenario based on the Laplace approximation that is, in a certain sense, optimal given a specific dataset. We empirically show that a Laplace approximation constructed with a dimensionally reduced covariance matrix closely matches the full Laplace approximation obtained using the exact covariance matrix. Where feasible, this subspace model can serve as a baseline for benchmarking the performance of subspace models. In addition, we provide a scalable approximation of this subspace construction that is usable in practice and compare it to existing subspace models from the literature. In general, our approximation scheme outperforms previous work. Furthermore, we present a metric to qualitatively compare the approximation quality of different subspace models even if the exact Laplace approximation is unknown.

cs.LG

Explainable AI needs formalization

The field of "explainable artificial intelligence" (XAI) seemingly addresses the desire that decisions of machine learning systems should be human-understandable. However, in its current state, XAI itself needs scrutiny. Popular methods cannot reliably answer relevant questions about ML models, their training data, or test inputs, because they systematically attribute importance to input features that are independent of the prediction target. This limits the utility of XAI for diagnosing and correcting data and models, for scientific discovery, and for identifying intervention targets. The fundamental reason for this is that current XAI methods do not address well-defined problems and are not evaluated against targeted criteria of explanation correctness. Researchers should formally define the problems they intend to solve and design methods accordingly. This will lead to diverse use-case-dependent notions of explanation correctness and objective metrics of explanation performance that can be used to validate XAI algorithms.

cs.LG

cc-Shapley: Measuring Multivariate Feature Importance Needs Causal Context

Explainable artificial intelligence promises to yield insights into relevant features, thereby enabling humans to examine and scrutinize machine learning models or even facilitating scientific discovery. Considering the widespread technique of Shapley values, we find that purely data-driven operationalization of multivariate feature importance is unsuitable for such purposes. Even for simple problems with two features, spurious associations due to collider bias and suppression arise from considering one feature only in the observational context of the other, which can lead to misinterpretations. Causal knowledge about the data-generating process is required to identify and correct such misleading feature attributions. We propose cc-Shapley (causal context Shapley), an interventional modification of conventional observational Shapley values leveraging knowledge of the data's causal structure, thereby analyzing the relevance of a feature in the causal context of the remaining features. We show theoretically that this eradicates spurious association induced by collider bias. We compare the behavior of Shapley and cc-Shapley values on various, synthetic, and real-world datasets. We observe nullification or reversal of associations compared to univariate feature importance when moving from observational to cc-Shapley.

cs.LG

Use of synthetic data for training dose estimation neural networks in CT dosimetry

Personalized computed tomography (CT) dosimetry has great potential in assessing patient-specific radiation exposure, supporting risk assessment, and optimizing clinical protocols. The aim of this study is to evaluate the potential of synthetic anatomical data for improving machine learning-based personalized computed tomography (CT) dosimetry. It is investigated whether the combination of synthetic human body geometries with real patient data can improve model accuracy and generalization for CT organ dose estimation while maintaining the uncertainty requirements outlined in IAEA TRS-457. Deep learning models for organ dose prediction are trained using datasets with varying proportions of real and synthetic data. Synthetic datasets are generated from computational human phantoms with controlled distributions of organ volumes and body. A dedicated model uncertainty evaluation method is implemented to quantify prediction reliability and verify compliance with TRS-457 accuracy limits. Model performance and uncertainty are compared across different training data compositions, including a model trained solely on real patient data. As baseline validated Monte Carlo simulation is used. Models trained solely on synthetic data show limited predictive accuracy, particularly for small or peripheral organs. Incorporating as little as 10 % real patient data significantly improves both statistical accuracy and uncertainty estimates, achieving a performance comparable to that of real-only models. The hybrid training approach improves robustness across different anatomies while maintaining TRS-457-compliant uncertainty levels (k=2 uncertainty < 20% for adults). The results indicate that the combination of real and synthetic data in combination with a systematic uncertainty assessment supports the development of CT dosimetry models and at the same time reduces the amount of real data required.

physics.med-ph

Aleatoric uncertainty for Errors-in-Variables models in deep regression

A Bayesian treatment of deep learning allows for the computation of uncertainties associated with the predictions of deep neural networks. We show how the concept of Errors-in-Variables can be used in Bayesian deep regression to also account for the uncertainty associated with the input of the employed neural network. The presented approach thereby exploits a relevant, but generally overlooked, source of uncertainty and yields a decomposition of the predictive uncertainty into an aleatoric and epistemic part that is more complete and, in many cases, more consistent from a statistical perspective. We discuss the approach along various simulated and real examples and observe that using an Errors-in-Variables model leads to an increase in the uncertainty while preserving the prediction performance of models without Errors-in-Variables. For examples with known regression function we observe that this ground truth is substantially better covered by the Errors-in-Variables model, indicating that the presented approach leads to a more reliable uncertainty estimation.

cs.LG

A framework for benchmarking uncertainty in deep regression

We propose a framework for the assessment of uncertainty quantification in deep regression. The framework is based on regression problems where the regression function is a linear combination of nonlinear functions. Basically, any level of complexity can be realized through the choice of the nonlinear functions and the dimensionality of their domain. Results of an uncertainty quantification for deep regression are compared against those obtained by a statistical reference method. The reference method utilizes knowledge of the underlying nonlinear functions and is based on a Bayesian linear regression using a reference prior. Reliability of uncertainty quantification is assessed in terms of coverage probabilities, and accuracy through the size of calculated uncertainties. We illustrate the proposed framework by applying it to current approaches for uncertainty quantification in deep regression. The flexibility, together with the availability of a reference solution, makes the framework suitable for defining benchmark sets for uncertainty quantification.

cs.LG

About exchanging expectation and supremum for conditional Wasserstein GANs

In cases where a Wasserstein GAN depends on a condition the latter is usually handled via an expectation within the loss function. Depending on the way this is motivated, the discriminator is either required to be Lipschitz-1 in both or in only one of its arguments. For the weaker requirement to become usable one needs to exchange a supremum and an expectation. This is a mathematically perilous operation, which is, so far, only partially justified in the literature. This short mathematical note intends to fill this gap and provides the mathematical rationale for discriminators that are only partially Lipschitz-1 for cases where this approach is more appropriate or successful.

cs.LG

Detecting unusual input to neural networks

Evaluating a neural network on an input that differs markedly from the training data might cause erratic and flawed predictions. We study a method that judges the unusualness of an input by evaluating its informative content compared to the learned parameters. This technique can be used to judge whether a network is suitable for processing a certain input and to raise a red flag that unexpected behavior might lie ahead. We compare our approach to various methods for uncertainty evaluation from the literature for various datasets and scenarios. Specifically, we introduce a simple, effective method that allows to directly compare the output of such metrics for single input points even if these metrics live on different scales.

cs.LG

The variation of the posterior variance and Bayesian sample size determination

We consider Bayesian sample size determination using a criterion that utilizes the first two moments of the expected posterior variance. We study the resulting sample size in dependence on the chosen prior and explore the success rate for bounding the posterior variance below a prescribed limit under the true sampling distribution. Compared with sample size determination based on the expected average of the posterior variance the proposed criterion leads to an increase in sample size and significantly improved success rates. Generic asymptotic properties are proven, such as an asymptotic expression for the sample size and a sort of phase transition. Our study is illustrated using two real world datasets with Poisson and normally distributed data. Based on our results some recommendations are given.

math.ST

Inspecting adversarial examples using the Fisher information

Adversarial examples are slight perturbations that are designed to fool artificial neural networks when fed as an input. In this work the usability of the Fisher information for the detection of such adversarial attacks is studied. We discuss various quantities whose computation scales well with the network size, study their behavior on adversarial examples and show how they can highlight the importance of single input neurons, thereby providing a visual tool for further analyzing (un-)reasonable behavior of a neural network. The potential of our methods is demonstrated by applications to the MNIST, CIFAR10 and Fruits-360 datasets.

cs.LG

Paracontrolled distributions on Bravais lattices and weak universality of the 2d parabolic Anderson model

We develop a discrete version of paracontrolled distributions as a tool for deriving scaling limits of lattice systems, and we provide a formulation of paracontrolled distribution in weighted Besov spaces. Moreover, we develop a systematic martingale approach to control the moments of polynomials of i.i.d. random variables and to derive their scaling limits. As an application, we prove a weak universality result for the parabolic Anderson model: We study a nonlinear population model in a small random potential and show that under weak assumptions it scales to the linear parabolic Anderson model.

math.PR

A Littlewood-Paley description of modelled distributions

We exhibit a fundamental link between Hairer's theory of regularity structures and the paracontrolled calculus of Gubinelli, Imkeller and Perkowski. By using paraproducts we provide a Littlewood-Paley description of the spaces of modelled distributions in regularity structures that is similar to the Besov description of classical Hölder spaces.

math.PR

Solution to the stochastic Schrödinger equation on the full space

We here show how the methods recently applied by [DW16] to solve the stochastic nonlinear Schrödinger equation on $\mathbb{T}^2$ can be enhanced to yield solutions on $\mathbb{R}^2$ if the non-linearity is weak enough. We prove that the solutions remains localized on compact time intervals which allows us to apply energy methods on the full space.

math.PR