SearcharxivSearch

arXiv subjects

Elmar Plischke

Publications and source records attributed to Elmar Plischke.

6 recordsLinked to original sources

Trustworthy Feature Importance Avoids Unrestricted Permutations

Feature importance methods using unrestricted permutations are flawed due to extrapolation errors; such errors appear in all non-trivial variable importance approaches. We propose three new approaches: conditional model reliance and Knockoffs with Gaussian transformation, and restricted ALE plot designs. Theoretical and numerical results show our strategies reduce/eliminate extrapolation.

stat.ML

No for Some, Yes for Others: Persona Prompts and Other Sources of False Refusal in Language Models

Large language models (LLMs) are increasingly integrated into our daily lives and personalized. However, LLM personalization might also increase unintended side effects. Recent work suggests that persona prompting can lead models to falsely refuse user requests. However, no work has fully quantified the extent of this issue. To address this gap, we measure the impact of 15 sociodemographic personas (based on gender, race, religion, and disability) on false refusal. To control for other factors, we also test 16 different models, 3 tasks (Natural Language Inference, politeness, and offensiveness classification), and nine prompt paraphrases. We propose a Monte Carlo-based method to quantify this issue in a sample-efficient manner. Our results show that as models become more capable, personas impact the refusal rate less and less. Certain sociodemographic personas increase false refusal in some models, which suggests underlying biases in the alignment strategies or safety mechanisms. However, we find that the model choice and task significantly influence false refusals, especially in sensitive content tasks. Our findings suggest that persona effects have been overestimated, and might be due to other factors.

cs.CL

gsaot: an R package for Optimal Transport-based sensitivity analysis

gsaot is an R package for Optimal Transport-based global sensitivity analysis. It provides a simple interface for indices estimation using a variety of state-of-the-art Optimal Transport solvers such as the network simplex and Sinkhorn-Knopp. The package is model-agnostic, allowing analysts to perform the sensitivity analysis as a post-processing step. Moreover, gsaot provides functions for indices and statistics visualization. In this work, we provide an overview of the theoretical grounds, of the implemented algorithms, and show how to use the package in different examples.

stat.CO

X-matrices

We evidence a family $\mathcal{X}$ of square matrices over a field $\mathbb{K}$, whose elements will be called X-matrices. We show that this family is shape invariant under multiplication as well as transposition. We show that $\mathcal{X}$ is a (in general non-commutative) subring of $GL(n,\mathbb{K})$. Moreover, we analyse the condition for a matrix $A \in \mathcal{X}$ to be invertible in $\mathcal{X}$. We also show that, if one adds a symmetry condition called here bi-symmetry, then the set $\mathcal{X}^b$ of bi-symmetric X-matrices is a commutative subring of $\mathcal{X}$. We propose results for eigenvalue inclusion, showing that for X-matrices eigenvalues lie exactly on the boundary of Cassini ovals. It is shown that any monic polynomial on $ \mathbb{R} $ can be associated with a companion matrix in $ \mathcal{X} $.

math.RA

Computing Shapley Effects for Sensitivity Analysis

Shapley effects are attracting increasing attention as sensitivity measures. When the value function is the conditional variance, they account for the individual and higher order effects of a model input. They are also well defined under model input dependence. However, one of the issues associated with their use is computational cost. We present a new algorithm that offers major improvements for the computation of Shapley effects, reducing computational burden by several orders of magnitude (from $k!\cdot k$ to $2^k$, where $k$ is the number of inputs) with respect to currently available implementations. The algorithm works in the presence of input dependencies. The algorithm also makes it possible to estimate all generalized (Shapley-Owen) effects for interactions.

stat.CO

Functional ANOVA with Multiple Distributions: Implications for the Sensitivity Analysis of Computer Experiments

The functional ANOVA expansion of a multivariate mapping plays a fundamental role in statistics. The expansion is unique once a unique distribution is assigned to the covariates. Recent investigations in the environmental and climate sciences show that analysts may not be in a position to assign a unique distribution in realistic applications. We offer a systematic investigation of existence, uniqueness, orthogonality, monotonicity and ultramodularity of the functional ANOVA expansion of a multivariate mapping when a multiplicity of distributions is assigned to the covariates. In particular, we show that a multivariate mapping can be associated with a core of probability measures that guarantee uniqueness. We obtain new results for variance decomposition and dimension distribution under mixtures. Implications for the global sensitivity analysis of computer experiments are also discussed.

stat.CO