SearcharxivSearch

arXiv subjects

Maximilian Schmidt

Publications and source records attributed to Maximilian Schmidt.

At least 19 recordsLinked to original sources

Flowcean - Model Learning for Cyber-Physical Systems

Effective models of Cyber-Physical Systems (CPS) are crucial for their design and operation. Constructing such models is difficult and time-consuming due to the inherent complexity of CPS. As a result, data-driven model generation using machine learning methods is gaining popularity. In this paper, we present Flowcean, a novel framework designed to automate the generation of models through data-driven learning that focuses on modularity and usability. By offering various learning strategies, data processing methods, and evaluation metrics, our framework provides a comprehensive solution, tailored to CPS scenarios. Flowcean facilitates the integration of diverse learning libraries and tools within a modular and flexible architecture, ensuring adaptability to a wide range of modeling tasks. This streamlines the process of model generation and evaluation, making it more efficient and accessible.

cs.LG

Carbon Footprint Evaluation of Code Generation through LLM as a Service

Due to increased computing use, data centers consume and emit a lot of energy and carbon. These contributions are expected to rise as big data analytics, digitization, and large AI models grow and become major components of daily working routines. To reduce the environmental impact of software development, green (sustainable) coding and claims that AI models can improve energy efficiency have grown in popularity. Furthermore, in the automotive industry, where software increasingly governs vehicle performance, safety, and user experience, the principles of green coding and AI-driven efficiency could significantly contribute to reducing the sector's environmental footprint. We present an overview of green coding and metrics to measure AI model sustainability awareness. This study introduces LLM as a service and uses a generative commercial AI language model, GitHub Copilot, to auto-generate code. Using sustainability metrics to quantify these AI models' sustainability awareness, we define the code's embodied and operational carbon.

cs.CY

Smooth Deep Saliency

In this work, we investigate methods to reduce the noise in deep saliency maps coming from convolutional downsampling. Those methods make the investigated models more interpretable for gradient-based saliency maps, computed in hidden layers. We evaluate the faithfulness of those methods using insertion and deletion metrics, finding that saliency maps computed in hidden layers perform better compared to both the input layer and GradCAM. We test our approach on different models trained for image classification on ImageNet1K, and models trained for tumor detection on Camelyon16 and in-house real-world digital pathology scans of stained tissue samples. Our results show that the checkerboard noise in the gradient gets reduced, resulting in smoother and therefore easier to interpret saliency maps.

cs.CV

Prompting-based Synthetic Data Generation for Few-Shot Question Answering

Although language models (LMs) have boosted the performance of Question Answering, they still need plenty of data. Data annotation, in contrast, is a time-consuming process. This especially applies to Question Answering, where possibly large documents have to be parsed and annotated with questions and their corresponding answers. Furthermore, Question Answering models often only work well for the domain they were trained on. Since annotation is costly, we argue that domain-agnostic knowledge from LMs, such as linguistic understanding, is sufficient to create a well-curated dataset. With this motivation, we show that using large language models can improve Question Answering performance on various datasets in the few-shot setting compared to state-of-the-art approaches. For this, we perform data generation leveraging the Prompting framework, suggesting that language models contain valuable task-agnostic knowledge that can be used beyond the common pre-training/fine-tuning scheme. As a result, we consistently outperform previous approaches on few-shot Question Answering.

cs.CL

Learn to Code Sustainably: An Empirical Study on LLM-based Green Code Generation

The increasing use of information technology has led to a significant share of energy consumption and carbon emissions from data centers. These contributions are expected to rise with the growing demand for big data analytics, increasing digitization, and the development of large artificial intelligence (AI) models. The need to address the environmental impact of software development has led to increased interest in green (sustainable) coding and claims that the use of AI models can lead to energy efficiency gains. Here, we provide an empirical study on green code and an overview of green coding practices, as well as metrics used to quantify the sustainability awareness of AI models. In this framework, we evaluate the sustainability of auto-generated code. The auto-generate codes considered in this study are produced by generative commercial AI language models, GitHub Copilot, OpenAI ChatGPT-3, and Amazon CodeWhisperer. Within our methodology, in order to quantify the sustainability awareness of these AI models, we propose a definition of the code's "green capacity", based on certain sustainability metrics. We compare the performance and green capacity of human-generated code and code generated by the three AI language models in response to easy-to-hard problem statements. Our findings shed light on the current capacity of AI models to contribute to sustainable software development.

cs.SE

Model Stitching and Visualization How GAN Generators can Invert Networks in Real-Time

In this work, we propose a fast and accurate method to reconstruct activations of classification and semantic segmentation networks by stitching them with a GAN generator utilizing a 1x1 convolution. We test our approach on images of animals from the AFHQ wild dataset, ImageNet1K, and real-world digital pathology scans of stained tissue samples. Our results show comparable performance to established gradient descent methods but with a processing time that is two orders of magnitude faster, making this approach promising for practical applications.

cs.CV

Learning-based approaches for reconstructions with inexact operators in nanoCT applications

Imaging problems such as the one in nanoCT require the solution of an inverse problem, where it is often taken for granted that the forward operator, i.e., the underlying physical model, is properly known. In the present work we address the problem where the forward model is inexact due to stochastic or deterministic deviations during the measurement process. We particularly investigate the performance of non-learned iterative reconstruction methods dealing with inexactness and learned reconstruction schemes, which are based on U-Nets and conditional invertible neural networks. The latter also provide the opportunity for uncertainty quantification. A synthetic large data set in line with a typical nanoCT setting is provided and extensive numerical experiments are conducted evaluating the proposed methods.

math.NA

An Educated Warm Start For Deep Image Prior-Based Micro CT Reconstruction

Deep image prior (DIP) was recently introduced as an effective unsupervised approach for image restoration tasks. DIP represents the image to be recovered as the output of a deep convolutional neural network, and learns the network's parameters such that the output matches the corrupted observation. Despite its impressive reconstructive properties, the approach is slow when compared to supervisedly learned, or traditional reconstruction techniques. To address the computational challenge, we bestow DIP with a two-stage learning paradigm: (i) perform a supervised pretraining of the network on a simulated dataset; (ii) fine-tune the network's parameters to adapt to the target reconstruction task. We provide a thorough empirical analysis to shed insights into the impacts of pretraining in the context of image reconstruction. We showcase that pretraining considerably speeds up and stabilizes the subsequent reconstruction task from real-measured 2D and 3D micro computed tomography data of biological specimens. The code and additional experimental materials are available at https://educateddip.github.io/docs.educated_deep_image_prior/.

eess.IV

Multi-Tracer Groundwater Dating in Southern Oman using Bayesian Modelling

In the scope of assessing aquifer systems in areas where freshwater is scarce, estimation of transit times is a vital step to quantify the effect of groundwater abstraction. Transit time distributions of different shapes, mean residence times, and contributions are used to represent the hydrogeological conditions in aquifer systems and are typically inferred from measured tracer concentrations by inverse modeling. In this study, a multi-tracer sampling campaign was conducted in the Salalah Plain in Southern Oman including CFCs, SF6, 39Ar, 14C, and 4He. Based on the data of three tracers, a two-component Dispersion Model (DMmix) and a nonparametric model with six age bins were assumed and evaluated using Bayesian statistics. In a Markov Chain Monte Carlo approach, the maximum likelihood parameter estimates and their uncertainties were determined. Model performance was assessed using Bayes factor and leave-one-out cross-validation. Both models suggest that the groundwater in the Salalah Plain is composed of a very young component below 30 yr and a very old component beyond 1,000 yr, with the nonparametric model performing slightly better than the DMmix model. All wells except one exhibit reasonable goodness of fit. Our results support the relevance of Bayesian modeling in hydrology and the potential of nonparametric models for an adequate representation of aquifer dynamics.

physics.geo-ph

Seshadri constants on abelian surfaces

So far, Seshadri constants on abelian surfaces are completely understood only in the cases of Picard number one and on principally polarized abelian surfaces with real multiplication. Beyond that, there are partial results for products of elliptic curves. In this paper, we show how to compute the Seshadri constant of any nef line bundle on any abelian surface over the complex numbers. We develop an effective algorithm depending only on the basis of the Néron-Severi group to compute not only the Seshadri constants but also the numerical data of their Seshadri curves. Access to the Seshadri curves allows us to plot Seshadri functions and better understand their structure. We show that already in the case of Picard number two the complexity of Seshadri functions can vary to a great degree. Our results indicate that aside from finitely many cases the complexity of the Seshadri function is at least as high as in the Cantor function.

math.AG

Conditional Invertible Neural Networks for Medical Imaging

Over the last years, deep learning methods have become an increasingly popular choice to solve tasks from the field of inverse problems. Many of these new data-driven methods have produced impressive results, although most only give point estimates for the reconstruction. However, especially in the analysis of ill-posed inverse problems, the study of uncertainties is essential. In our work, we apply generative flow-based models based on invertible neural networks to two challenging medical imaging tasks, i.e. low-dose computed tomography and accelerated medical resonance imaging. We test different architectures of invertible neural networks and provide extensive ablation studies. In most applications, a standard Gaussian is used as the base distribution for a flow-based model. Our results show that the choice of a radial distribution can improve the quality of reconstructions.

eess.IV

Evolving Neuronal Plasticity Rules using Cartesian Genetic Programming

We formulate the search for phenomenological models of synaptic plasticity as an optimization problem. We employ Cartesian genetic programming to evolve biologically plausible human-interpretable plasticity rules that allow a given network to successfully solve tasks from specific task families. While our evolving-to-learn approach can be applied to various learning paradigms, here we illustrate its power by evolving plasticity rules that allow a network to efficiently determine the first principal component of its input distribution. We demonstrate that the evolved rules perform competitively with known hand-designed solutions. We explore how the statistical properties of the datasets used during the evolutionary search influences the form of the plasticity rules and discover new rules which are adapted to the structure of the corresponding datasets.

cs.NE

Evolving to learn: discovering interpretable plasticity rules for spiking networks

Continuous adaptation allows survival in an ever-changing world. Adjustments in the synaptic coupling strength between neurons are essential for this capability, setting us apart from simpler, hard-wired organisms. How these changes can be mathematically described at the phenomenological level, as so called "plasticity rules", is essential both for understanding biological information processing and for developing cognitively performant artificial systems. We suggest an automated approach for discovering biophysically plausible plasticity rules based on the definition of task families, associated performance measures and biophysical constraints. By evolving compact symbolic expressions we ensure the discovered plasticity rules are amenable to intuitive understanding, fundamental for successful communication and human-guided generalization. We successfully apply our approach to typical learning scenarios and discover previously unknown mechanisms for learning efficiently from rewards, recover efficient gradient-descent methods for learning from target signals, and uncover various functionally equivalent STDP-like rules with tuned homeostatic mechanisms.

q-bio.NC

Seshadri constants on principally polarized abelian surfaces with real multiplication

Seshadri constants on abelian surfaces are fully understood in the case of Picard number one. Little is known so far for simple abelian surfaces of higher Picard number. In this paper we investigate principally polarized abelian surfaces with real multiplication. They are of Picard number two and might be considered the next natural case to be studied. The challenge is to not only determine the Seshadri constants of individual line bundles, but to understand the whole \emph{Seshadri function} on these surfaces. Our results show on the one hand that this function is surprisingly complex: On surfaces with real multiplication in $\mathbb Z[\sqrt e]$ it consists of linear segments that are never adjacent to each other -- it behaves like the Cantor function. On the other hand, we prove that the Seshadri function it is invariant under an infinite group of automorphisms, which shows that it does have interesting regular behavior globally.

math.AG

Conditional Normalizing Flows for Low-Dose Computed Tomography Image Reconstruction

Image reconstruction from computed tomography (CT) measurement is a challenging statistical inverse problem since a high-dimensional conditional distribution needs to be estimated. Based on training data obtained from high-quality reconstructions, we aim to learn a conditional density of images from noisy low-dose CT measurements. To tackle this problem, we propose a hybrid conditional normalizing flow, which integrates the physical model by using the filtered back-projection as conditioner. We evaluate our approach on a low-dose CT benchmark and demonstrate superior performance in terms of structural similarity of our flow-based method compared to other deep learning based approaches.

eess.IV

ADVISER: A Toolkit for Developing Multi-modal, Multi-domain and Socially-engaged Conversational Agents

We present ADVISER - an open-source, multi-domain dialog system toolkit that enables the development of multi-modal (incorporating speech, text and vision), socially-engaged (e.g. emotion recognition, engagement level prediction and backchanneling) conversational agents. The final Python-based implementation of our toolkit is flexible, easy to use, and easy to extend not only for technically experienced users, such as machine learning researchers, but also for less technically experienced users, such as linguists or cognitive scientists, thereby providing a flexible platform for collaborative research. Link to open-source code: https://github.com/DigitalPhonetics/adviser

cs.CL

The LoDoPaB-CT Dataset: A Benchmark Dataset for Low-Dose CT Reconstruction Methods

Deep Learning approaches for solving Inverse Problems in imaging have become very effective and are demonstrated to be quite competitive in the field. Comparing these approaches is a challenging task since they highly rely on the data and the setup that is used for training. We provide a public dataset of computed tomography images and simulated low-dose measurements suitable for training this kind of methods. With the LoDoPaB-CT Dataset we aim to create a benchmark that allows for a fair comparison. It contains over 40,000 scan slices from around 800 patients selected from the LIDC/IDRI Database. In this paper we describe how we processed the original slices and how we simulated the measurements. We also include first baseline results.

eess.IV

Computed Tomography Reconstruction Using Deep Image Prior and Learned Reconstruction Methods

In this work, we investigate the application of deep learning methods for computed tomography in the context of having a low-data regime. As motivation, we review some of the existing approaches and obtain quantitative results after training them with different amounts of data. We find that the learned primal-dual has an outstanding performance in terms of reconstruction quality and data efficiency. However, in general, end-to-end learned methods have two issues: a) lack of classical guarantees in inverse problems and b) lack of generalization when not trained with enough data. To overcome these issues, we bring in the deep image prior approach in combination with classical regularization. The proposed methods improve the state-of-the-art results in the low data-regime.

eess.IV