SearcharxivSearch

arXiv subjects

Marta Oliveira

Publications and source records attributed to Marta Oliveira.

7 recordsLinked to original sources

Feature salience - not task-informativeness - drives machine learning model explanations

Explainable AI (XAI) promises to provide insight into machine learning models' decision processes, where one goal is to identify failures such as shortcut learning. This promise relies on the field's assumption that input features marked as important by an XAI must contain information about the target variable. However, it is unclear whether informativeness is indeed the main driver of importance attribution in practice, or if other data properties such as statistical suppression, novelty at test-time, or high feature salience substantially contribute. To clarify this, we trained deep learning models on three variants of a binary image classification task, in which translucent watermarks are either absent, act as class-dependent confounds, or represent class-independent noise. Results for five popular attribution methods show substantially elevated relative importance in watermarked areas (RIW) for all models regardless of the training setting ($R^2 \geq .45$). By contrast, whether the presence of watermarks is class-dependent or not only has a marginal effect on RIW ($R^2 \leq .03$), despite a clear impact impact on model performance and generalisation ability. XAI methods show similar behaviour to model-agnostic edge detection filters and attribute substantially less importance to watermarks when bright image intensities are encoded by smaller instead of larger feature values. These results indicate that importance attribution is most strongly driven by the salience of image structures at test time rather than statistical associations learned by machine learning models. Previous studies demonstrating successful XAI application should be reevaluated with respect to a possibly spurious concurrency of feature salience and informativeness, and workflows using feature attribution methods as building blocks should be scrutinised.

cs.LG

Benchmarking the Influence of Pre-training on Explanation Performance in MR Image Classification

Convolutional Neural Networks (CNNs) are frequently and successfully used in medical prediction tasks. They are often used in combination with transfer learning, leading to improved performance when training data for the task are scarce. The resulting models are highly complex and typically do not provide any insight into their predictive mechanisms, motivating the field of "explainable" artificial intelligence (XAI). However, previous studies have rarely quantitatively evaluated the "explanation performance" of XAI methods against ground-truth data, and transfer learning and its influence on objective measures of explanation performance has not been investigated. Here, we propose a benchmark dataset that allows for quantifying explanation performance in a realistic magnetic resonance imaging (MRI) classification task. We employ this benchmark to understand the influence of transfer learning on the quality of explanations. Experimental results show that popular XAI methods applied to the same underlying model differ vastly in performance, even when considering only correctly classified examples. We further observe that explanation performance strongly depends on the task used for pre-training and the number of CNN layers pre-trained. These results hold after correcting for a substantial correlation between explanation and classification performance.

cs.CV

GECOBench: A Gender-Controlled Text Dataset and Benchmark for Quantifying Biases in Explanations

Large pre-trained language models have become a crucial backbone for many downstream tasks in natural language processing (NLP), and while they are trained on a plethora of data containing a variety of biases, such as gender biases, it has been shown that they can also inherit such biases in their weights, potentially affecting their prediction behavior. However, it is unclear to what extent these biases also affect feature attributions generated by applying "explainable artificial intelligence" (XAI) techniques, possibly in unfavorable ways. To systematically study this question, we create a gender-controlled text dataset, GECO, in which the alteration of grammatical gender forms induces class-specific words and provides ground truth feature attributions for gender classification tasks. This enables an objective evaluation of the correctness of XAI methods. We apply this dataset to the pre-trained BERT model, which we fine-tune to different degrees, to quantitatively measure how pre-training induces undesirable bias in feature attributions and to what extent fine-tuning can mitigate such explanation bias. To this extent, we provide GECOBench, a rigorous quantitative evaluation framework for benchmarking popular XAI methods. We show a clear dependency between explanation performance and the number of fine-tuned layers, where XAI methods are observed to benefit particularly from fine-tuning or complete retraining of embedding layers.

cs.LG

EXACT: Towards a platform for empirically benchmarking Machine Learning model explanation methods

The evolving landscape of explainable artificial intelligence (XAI) aims to improve the interpretability of intricate machine learning (ML) models, yet faces challenges in formalisation and empirical validation, being an inherently unsupervised process. In this paper, we bring together various benchmark datasets and novel performance metrics in an initial benchmarking platform, the Explainable AI Comparison Toolkit (EXACT), providing a standardised foundation for evaluating XAI methods. Our datasets incorporate ground truth explanations for class-conditional features, and leveraging novel quantitative metrics, this platform assesses the performance of post-hoc XAI methods in the quality of the explanations they produce. Our recent findings have highlighted the limitations of popular XAI methods, as they often struggle to surpass random baselines, attributing significance to irrelevant features. Moreover, we show the variability in explanations derived from different equally performing model architectures. This initial benchmarking platform therefore aims to allow XAI researchers to test and assure the high quality of their newly developed methods.

cs.LG

Hesperos: A geophysical mission to Venus

The Hesperos mission proposed in this paper is a mission to Venus to investigate the interior structure and the current level of activity. The main questions to be answered with this mission are whether Venus has an internal structure and composition similar to Earth and if Venus is still tectonically active. To do so the mission will consist of two elements: an orbiter to investigate the interior and changes over longer periods of time and a balloon floating at an altitude between 40 and 60km to investigate the composition of the atmosphere. The mission will start with the deployment of the balloon which will operate for about 25 days. During this time the orbiter acts as a relay station for data communication with Earth. Once the balloon phase is finished the orbiter will perform surface and gravity gradient mapping over the course of 7 Venus days. This mission proposal is the result of the Alpbach Summer School and the post-Alpbach week.

astro-ph.EP

Detailed experimental and numerical analysis of a cylindrical cup deep drawing: pros and cons of using solid-shell elements

The Swift test was originally proposed as a formability test to reproduce the conditions observed in deep drawing operations. This test consists on forming a cylindrical cup from a circular blank, using a flat bottom cylindrical punch and has been extensively studied using both analytical and numerical methods. This test can also be combined with the Demeri test, which consists in cutting a ring from the wall of a cylindrical cup, in order to open it afterwards to measure the springback. This combination allows their use as benchmark test, in order to improve the knowledge concerning the numerical simulation models, through the comparison between experimental and numerical results. The focus of this study is the experimental and numerical analyses of the Swift cup test, followed by the Demeri test, performed with an AA5754-O alloy at room temperature. In this context, a detailed analysis of the punch force evolution, the thickness evolution along the cup wall, the earing profile, the strain paths and their evolution and the ring opening is performed. The numerical simulation is performed using the finite element code ABAQUS, with solid and solid-shell elements, in order to compare the computational efficiency of these type of elements. The results show that the solid-shell element is more cost-effective than the solid, presenting global accurate predictions, excepted for the thinning zones. Both the von Mises and the Hill48 yield criteria predict the strain distributions in the final cup quite accurately. However, improved knowledge concerning the stress states is still required, because the Hill48 criterion showed difficulties in the correct prediction of the springback, whatever the type of finite element adopted.

physics.class-ph

Thermo-mechanical finite element analysis of the AA5086 alloy under warm forming conditions

Warm forming processes have been successfully applied at laboratory level to overcome some important drawbacks of the Al-Mg alloys, such as poor formability and large springback. However, the numerical simulation of these processes requires the adoption of coupled thermo-mechanical finite element analysis, using temperature-dependent material models. The numerical description of the thermo-mechanical behaviour can require a large set of experimental tests. These experimental tests should be performed under conditions identical to the ones observed in the forming process. In this study, the warm deep drawing of a cylindrical cup is analysed, including the split-ring test to assess the temperature effect on the springback. Based on the analysis of the forming process conditions, the thermo-mechanical behaviour of the AA5086 aluminium alloy is described by a rate-independent thermo-elasto-plastic material model. The hardening law adopted is temperature-dependent while the yield function is temperature-independent. Nevertheless, the yield criterion parameters are selected based on the temperature of the heated tools. In fact, the model assumes that the temperature of the tools is uniform and constant, adopting a variable interfacial heat transfer coefficient. The accuracy of the proposed finite element model is assessed by comparing numerical and experimental results. The predicted punch force, thickness distribution and earing profile are in very good agreement with the experimental measurements, when the anisotropic behaviour of the blank is accurately described. However, this does not guarantee a correct springback prediction, which is strongly influenced by the elastic properties, namely the Young's modulus.

cond-mat.mtrl-sci