SearcharxivSearch

arXiv subjects

Sonja Greven

Publications and source records attributed to Sonja Greven.

At least 19 recordsLinked to original sources

ICON Decomposition: Auditing deep neural networks for shortcuts by decomposing layer-wise representations using concepts

Deep neural networks often exploit spurious associations, a failure known as shortcut learning. Before deployment, models should be audited for reliance on a set of concepts, such as acquisition artifacts or demographics. Current methods, such as linear probes and concept activation vectors, measure reliance by asking whether each concept, in isolation, is decodable from a layer. Their scores therefore reflect not only reliance but also correlations in the audit dataset. We introduce Independent Canonical cONcept (ICON) decomposition, which quantifies the share of a layer's variance each concept explains, conditional on all other concepts and the outcome. ICON scores are variance shares, comparable across layers and between continuous and categorical concepts. ICON also reports the share the set leaves unexplained. On simulated data, ICON recovers the true importance more accurately than seven baselines. On skin-cancer and neuroimaging models, ICON distinguishes learned shortcuts from correlated concepts, confirmed by retraining and out-of-distribution tests.

cs.LG

Random mixtures in Bayes Hilbert spaces

We present a framework for the analysis and unmixing of random density mixtures in the Bayes Hilbert space. General identifiability results for mixtures in Hilbert spaces are established and applied to the Bayes Hilbert space setting. Building on these results, we propose a penalised maximum likelihood approach for the unmixing of Bayes Hilbert mixtures aimed at recovering the statistically space-efficient representation, together with a computationally efficient coordinate-wise maximisation algorithm for its implementation. The methodology is illustrated through a hyperspectral data application, where observations can be naturally embedded in the Bayes Hilbert space and analyzed in terms of distributional shape rather than amplitude. A complementary simulation study demonstrates the interpretability and practical performance of the proposed approach.

stat.ME

Embedded Conditional Independence Tests for Large Language Model Generated Text with an Application to German Parliament Speeches

Conditional independence tests (CITs) test for conditional dependence between two random objects $X$ and $Y$ given a third random object $Z$. Existing CITs have limited applicability to high-dimensional data, especially multimodal data like text. However, we show that such tests are of interest for large language model (LLM) outputs, where we test whether an output $X$ generated from a source text $Z$ carries information about an attribute $Y$ beyond $Z$ itself. For this purpose, we propose embedded CITs (eCITs), which embed $X$ and $Z$ and apply an existing CIT to the resulting representations and to $Y$. We show that, provided the embedding of $Z$ is sufficient, i.e. retains the information $Z$ carries about either $Y$ or the representation of $X$, the null hypothesis transfers from $X$ and $Z$ to their representations, so that a CIT valid for the embedded hypothesis is valid for the original one. We further give conditions for equivalence of the two hypotheses, and show that sufficiency weakens to mean sufficiency when the embedded test targets conditional mean independence. We propose a semi-synthetic simulation design to assess type I error (T1E) control and power of the eCITs for given embedding maps on a specific dataset and task, and use it to evaluate them on our application. Applying the eCITs to German Parliament speeches, we find for all combinations of embedding maps considered that the summaries of two LLMs contain information about the speaker's faction and gender beyond the speech they were generated from.

stat.ML

Controlling for Omitted Variable Bias in Deep Neural Networks

Control variables are widely used in statistical modelling to account for omitted variable bias of known confounders. However, they have largely been underexplored in deep learning. This is surprising, given that deep learning models encode image-inferable covariates, such as demographic variables, into their predictions when these covariates are correlated with the outcome---a form of omitted variable bias referred to as 'shortcut learning'. While many existing confound-control or fairness methods try to restrict the correlation of such covariates with model predictions, we show that this fails to correct for omitted variable bias. We therefore propose a control variable approach for deep learning models, based on generalised additive modelling of the effects of model inputs and covariates. As flexible additive models can suffer from concurvity, we introduce an estimation procedure that refits the final layer of a pre-trained network to include covariate effects, using cross-fitting with ridge penalisation. We show how these effects can be orthogonalised with respect to covariates to exclude their mediated effects and that model predictions can be marginalised over the covariate distribution to control for their effect. This yields unbiased, interpretable predictions and offers flexibility to model the desired effects depending on the scientific or fairness objective. We verify our approach using simulated images, and demonstrate consistent estimation of true effects. Existing methods either require more data or fail to recover the true effects. We apply our method to real neuroimaging data with experimentally induced confounding, where it recovers prediction performance to near the level of a model trained on unconfounded data. Code is available at https://github.com/mpff/cocodeel.

stat.ME

Deep Shape Regression for Planar Curves with Multimodal Covariates

The shape of a planar curve is the geometric information that remains once translation, rotation, scale and reparametrisation are removed and is of interest in many health applications, e.g. in neuroimaging. We propose a deep shape regression model for open planar curves that admits multimodal and high-dimensional covariates. Representing curves as complex-valued functions, we show that the conditional full Procrustes mean is the leading eigenfunction of the conditional covariance. To estimate this covariance surface, we propose a novel deep conditional covariance smoother with modality-specific encoders - e.g. splines for scalar covariates and convolutional networks for images, which classical spline smoothers cannot accommodate. Our model is by construction invariant to the translation, rotation and scaling of the input curves and handles sparsely and irregularly sampled curves. We further provide an algorithm for elastic mean estimation that also removes parametrisation by iterating covariance smoothing, rotational alignment and parametrisation alignment. We illustrate the method on simulated outlines with known conditional mean and multimodal covariates, and give a first application to hippocampal outlines from the ADNI cohort, recovering covariate effects consistent with the literature. Code is available at https://github.com/mpff/dnn-shapes.

stat.ME

Dimension reduction of multivariate densities in Bayes spaces

The Bayes space provides a Hilbert space structure for analysing probability density functions (PDFs), equipping them with a geometry that reflects their relative and constrained nature. A key tool in this framework is the centred logratio (clr) transformation, which establishes an isometric isomorphism between the Bayes space and (a subspace of) the classical $L^2$ space. This makes it possible to apply functional data analysis (FDA) techniques, particularly functional principal component analysis (FPCA), to both univariate and multivariate density data in the context of dimension reduction. For multivariate PDFs, embedding them in the Bayes space enables an orthogonal decomposition into independent and interactive components. Furthermore, the independent part can be decomposed into mutually orthogonal geometric marginals. This structure provides more profound insights into the sources of variation in multivariate densities. We show that this decomposition of the total variance is optimal in a PCA sense, impacting the interpretation of the eigenfunctions and scores resulting from FPCA. We demonstrate that applying FPCA directly to multivariate densities is equivalent in a certain sense to applying multivariate FPCA to their decomposed form, with the resulting eigenfunctions and scores decomposing accordingly. The unique decomposition based on these theoretical results is applied to housing and geological empirical data respectively, demonstrating the interpretability and practical value of this approach.

stat.ME

Variable Domain Multivariate Functional Principal Component Analysis

Multivariate functional principal component analysis (MFPCA) is a powerful dimension reduction technique for analyzing multiple functional variables simultaneously. However, existing MFPCA methods assume that all functional observations are recorded over a common, fixed domain. This assumption is often violated in practical applications where the observation period varies across subjects, leading to what is known as variable domain functional data. We propose a novel approach for MFPCA that explicitly accommodates variable domains by extending existing multivariate functional principal component analysis to the variable domain setting. Our methodology involves performing univariate variable domain FPCA for each functional variable separately, stacking the resulting univariate scores, and then smoothing the empirical covariance matrix of these stacked scores over the domain length. This allows us to estimate multivariate eigenfunctions and scores that properly account for varying observation periods. We demonstrate through extensive simulation studies that our proposed method outperforms approaches that ignore the variable domain structure and rely on binning strategies. The practical utility of our method is illustrated through an application analyzing body temperature and capillary oxygen saturation (SpO$_2$) trajectories in COVID-19 hospital admitted patients, where patients experienced varying lengths of stay and monitoring periods.

stat.ME

Counterfactual Density Effects and the German East--West Income Gap

We propose a novel framework for conducting causal inference based on counterfactual densities. While the current paradigm of causal inference is mostly focused on estimating average treatment effects (ATEs), which restricts the analysis to the first moment of the outcome variable, our density-based approach is able to detect causal effects based on general distributional characteristics. Following the Oaxaca-Blinder decomposition approach, we consider two types of counterfactual density effects that together explain observed discrepancies between the densities of the treated and control group. First, the distribution effect is the counterfactual effect of changing the conditional density of the control group to that of the treatment group, while keeping the covariates fixed at the treatment group distribution. Second, the covariate effect represents the effect of a hypothetical change in the covariate distribution. Both effects have a causal interpretation under the classical unconfoundedness and overlap assumptions. Methodologically, our approach is based on analyzing the conditional densities as elements of a Bayes Hilbert space, which preserves the non-negativity and integration-to-one constraints. We specify a flexible functional additive regression model estimating the conditional densities. We apply our method to analyze the German East--West income gap, i.e., the observed differences in wages between East Germans and West Germans. While most of the existing studies focus on the average differences and neglect other distributional characteristics, our density-based approach is suited to detect all nuances of the counterfactual distributions, including differences in probability masses at zero.

econ.EM

Better than Average: Spatially-Aware Aggregation of Segmentation Uncertainty Improves Downstream Performance

Uncertainty Quantification (UQ) is crucial for ensuring the reliability of automated image segmentations in safety-critical domains like biomedical image analysis or autonomous driving. In segmentation, UQ generates pixel-wise uncertainty scores that must be aggregated into image-level scores for downstream tasks like Out-of-Distribution (OoD) or failure detection. Despite routine use of aggregation strategies, their properties and impact on downstream task performance have not yet been comprehensively studied. Global Average is the default choice, yet it does not account for spatial and structural features of segmentation uncertainty. Alternatives like patch-, class- and threshold-based strategies exist, but lack systematic comparison, leading to inconsistent reporting and unclear best practices. We address this gap by (1) formally analyzing properties, limitations, and pitfalls of common strategies; (2) proposing novel strategies that incorporate spatial uncertainty structure and (3) benchmarking their performance on OoD and failure detection across ten datasets that vary in image geometry and structure. We find that aggregators leveraging spatial structure yield stronger performance in both downstream tasks studied. However, the performance of individual aggregators depends heavily on dataset characteristics, so we (4) propose a meta-aggregator that integrates multiple aggregators and performs robustly across datasets.

cs.CV

Physiological and Behavioral Modeling of Stress and Cognitive Load in Web-Based Question Answering

Time pressure and question difficulty can trigger stress and cognitive overload in web-based surveys, compromising data quality and user experience. Most stress detection methods are based on low-resolution self-reports, which are poorly suited for capturing fast, moment-to-moment changes during short online tasks. Addressing this gap, we conducted a 2x2 within-subjects study (N = 29), manipulating question difficulty and time pressure in a web-based multiple-choice task. Participants completed general knowledge and cognitive questions while we collected multimodal data: mouse dynamics, eye tracking, electrocardiogram, and electrodermal activity. Using condition-based and self-reported labels, we used statistical and machine learning models to model stress and question difficulty. Our results show distinct physiological and behavioral patterns within very short timeframes. This work demonstrates the feasibility of rapidly detecting cognitive-affective states in digital environments, paving the way for more adaptive, ethical, and user-aware survey interfaces.

cs.HC

Additive Density Regression

We present a structured additive regression approach to model conditional densities given scalar covariates, where only samples of the conditional distributions are observed. This links our approach to distributional regression models for scalar data. The model is formulated in a Bayes Hilbert space -- preserving nonnegativity and integration to one under summation and scalar multiplication -- with respect to an arbitrary finite measure. This allows to consider, amongst others, continuous, discrete and mixed densities. Our theoretical results include asymptotic existence, uniqueness, consistency, and asymptotic normality of the penalized maximum likelihood estimator, as well as confidence regions and inference for the (effect) densities. For estimation, we propose to maximize the penalized log-likelihood corresponding to an appropriate multinomial, or equivalently, Poisson regression model, which we show to approximate the original penalized maximum likelihood problem. We apply our framework to a motivating gender economic data set from the German Socio-Economic Panel Study (SOEP), analyzing the distribution of the woman's share in a couple's total labor income given covariate effects for year, place of residence and age of the youngest child. As the income share is a continuous variable having discrete point masses at zero and one for single-earner couples, the corresponding densities are of mixed type.

stat.ME

On spatial point processes with composition-valued marks

Methods for marked spatial point processes with scalar marks have seen extensive development in recent years. While the impressive progress in data collection and storage capacities has yielded an immense increase in spatial point process data with highly challenging non-scalar marks, methods for their analysis are not equally well developed. In particular, there are no methods for composition-valued marks, i.e. vector-valued marks with a sum-to-constant constrain (typically 1 or 100). Prompted by the need for a suitable methodological framework, we extend existing methods to spatial point processes with composition-valued marks and adapt common mark characteristics to this context. The proposed methods are applied to analyse spatial correlations in data on tree crown-to-base and business sector compositions.

stat.ME

Deep Nonparametric Conditional Independence Tests for Images

Conditional independence tests (CITs) test for conditional dependence between random variables. As existing CITs are limited in their applicability to complex, high-dimensional variables such as images, we introduce deep nonparametric CITs (DNCITs). The DNCITs combine embedding maps, which extract feature representations of high-dimensional variables, with nonparametric CITs applicable to these feature representations. For the embedding maps, we derive general properties on their parameter estimators to obtain valid DNCITs and show that these properties include embedding maps learned through (conditional) unsupervised or transfer learning. For the nonparametric CITs, appropriate tests are selected and adapted to be applicable to feature representations. Through simulations, we investigate the performance of the DNCITs for different embedding maps and nonparametric CITs under varying confounder dimensions and confounder relationships. We apply the DNCITs to brain MRI scans and behavioral traits, given confounders, of healthy individuals from the UK Biobank (UKB), confirming null results from a number of ambiguous personality neuroscience studies with a larger data set and with our more powerful tests. In addition, in a confounder control study, we apply the DNCITs to brain MRI scans and a confounder set to test for sufficient confounder control, leading to a potential reduction in the confounder dimension under improved confounder control compared to existing state-of-the-art confounder control studies for the UKB. Finally, we provide an R package implementing the DNCITs.

stat.ML

Generalized Multivariate Functional Additive Mixed Models for Location, Scale, and Shape

We propose a flexible regression framework to model the conditional distribution of multilevel generalized multivariate functional data of potentially mixed type, e.g. binary and continuous data. We make pointwise parametric distributional assumptions for each dimension of the multivariate functional data and model each distributional parameter as an additive function of covariates. The dependency between the different outcomes and, for multilevel functional data, also between different functions within a level is modelled by shared latent multivariate Gaussian processes. For a parsimonious representation of the latent processes, (generalized) multivariate functional principal components are estimated from the data and used as an empirical basis for these latent processes in the regression framework. Our modular two-step approach is very general and can easily incorporate new developments in the estimation of functional principal components for all types of (generalized) functional data. Flexible additive covariate effects for scalar or even functional covariates are available and are estimated in a Bayesian framework. We provide an easy-to-use implementation in the accompanying R package 'gmfamm' on CRAN and conduct a simulation study to confirm the validity of our regression framework and estimation strategy. The proposed multivariate functional model is applied to four dimensional traffic data in Berlin, which consists of the hourly numbers and mean speed of cars and trucks at different locations.

stat.ME

Additive Density-on-Scalar Regression in Bayes Hilbert Spaces with an Application to Gender Economics

Motivated by research on gender identity norms and the distribution of the woman's share in a couple's total labor income, we consider functional additive regression models for probability density functions as responses with scalar covariates. To preserve nonnegativity and integration to one under vector space operations, we formulate the model for densities in a Bayes Hilbert space, which allows to not only consider continuous densities, but also, e.g., discrete or mixed densities. Mixed ones occur in our application, as the woman's income share is a continuous variable having discrete point masses at zero and one for single-earner couples. Estimation is based on a gradient boosting algorithm, allowing for potentially numerous flexible covariate effects and model selection. We develop properties of Bayes Hilbert spaces related to subcompositional coherence, yielding (odds-ratio) interpretation of effect functions and simplified estimation for mixed densities via an orthogonal decomposition. Applying our approach to data from the German Socio-Economic Panel Study (SOEP) shows a more symmetric distribution in East German than in West German couples after reunification and a smaller child penalty comparing couples with and without minor children. These West-East differences become smaller, but are persistent over time.

stat.ME

Approximation of bivariate densities with compositional splines

Reliable estimation and approximation of probability density functions is fundamental for their further processing. However, their specific properties, i.e. scale invariance and relative scale, prevent the use of standard methods of spline approximation and have to be considered when building a suitable spline basis. Bayes Hilbert space methodology allows to account for these properties of densities and enables their conversion to a standard Lebesgue space of square integrable functions using the centered log-ratio transformation. As the transformed densities fulfill a zero integral constraint, the constraint should likewise be respected by any spline basis used. Bayes Hilbert space methodology also allows to decompose bivariate densities into their interactive and independent parts with univariate marginals. As this yields a useful framework for studying the dependence structure between random variables, a spline basis ideally should admit a corresponding decomposition. This paper proposes a new spline basis for (transformed) bivariate densities respecting the desired zero integral property. We show that there is a one-to-one correspondence of this basis to a corresponding basis in the Bayes Hilbert space of bivariate densities using tools of this methodology. Furthermore, the spline representation and the resulting decomposition into interactive and independent parts are derived. Finally, this novel spline representation is evaluated in a simulation study and applied to empirical geochemical data.

stat.ME

Flexible joint models for multivariate longitudinal and time-to-event data using multivariate functional principal components

The joint modeling of multiple longitudinal biomarkers together with a time-to-event outcome is a challenging modeling task of continued scientific interest. In particular, the computational complexity of high dimensional (generalized) mixed effects models often restricts the flexibility of shared parameter joint models, even when the subject-specific marker trajectories follow highly nonlinear courses. We propose a parsimonious multivariate functional principal components representation of the shared random effects. This allows better scalability, as the dimension of the random effects does not directly increase with the number of markers, only with the chosen number of principal component basis functions used in the approximation of the random effects. The functional principal component representation additionally allows to estimate highly flexible subject-specific random trajectories without parametric assumptions. The modeled trajectories can thus be distinctly different for each biomarker. We build on the framework of flexible Bayesian additive joint models implemented in the R-package 'bamlss', which also supports estimation of nonlinear covariate effects via Bayesian P-splines. The flexible yet parsimonious functional principal components basis used in the estimation of the joint model is first estimated in a preliminary step. We validate our approach in a simulation study and illustrate its advantages by analyzing a study on primary biliary cholangitis.

stat.ME

Functional Data Analysis: An Introduction and Recent Developments

Functional data analysis (FDA) is a statistical framework that allows for the analysis of curves, images, or functions on higher dimensional domains. The goals of FDA, such as descriptive analyses, classification, and regression, are generally the same as for statistical analyses of scalar-valued or multivariate data, but FDA brings additional challenges due to the high- and infinite dimensionality of observations and parameters, respectively. This paper provides an introduction to FDA, including a description of the most common statistical analysis techniques, their respective software implementations, and some recent developments in the field. The paper covers fundamental concepts such as descriptives and outliers, smoothing, amplitude and phase variation, and functional principal component analysis. It also discusses functional regression, statistical inference with functional data, functional classification and clustering, and machine learning approaches for functional data analysis. The methods discussed in this paper are widely applicable in fields such as medicine, biophysics, neuroscience, and chemistry, and are increasingly relevant due to the widespread use of technologies that allow for the collection of functional data. Sparse functional data methods are also relevant for longitudinal data analysis. All presented methods are demonstrated using available software in R by analyzing a data set on human motion and motor control. To facilitate the understanding of the methods, their implementation, and hands-on application, the code for these practical examples is made available on Github: https://github.com/davidruegamer/FDA_tutorial .

stat.ME