SearcharxivSearch

arXiv subjects

Jan Ewald

Publications and source records attributed to Jan Ewald.

3 recordsLinked to original sources

BBOmix: A Tabular Benchmark for Hyperparameter Optimization of Unsupervised Biological Representation Learning

The rapid advancement of high-throughput sequencing has led to large, high-dimensional omics datasets. Deep unsupervised learning architectures, particularly Autoencoders (AEs), are increasingly used for dimensionality reduction and representation learning in this domain. However, AEs are highly sensitive to architectural choices and hyperparameters, and unsupervised optimization typically relies on reconstruction loss, which may be a poor proxy for downstream utility. Exhaustive hyperparameter optimization (HPO) is computationally expensive, leading researchers to frequently rely on suboptimal default configurations. To democratize access to large-scale unsupervised HPO research, we introduce $\textbf{BBOmix}$, the first open-source tabular benchmark for unsupervised representation learning on real-world biological data. Our benchmark includes 105,000 evaluations across four AE architectures and seven multi-omics modalities from the TCGA and SCHC datasets. We quantify the correlation between reconstruction loss and downstream task performance and provide an extensive evaluation of state-of-the-art single-fidelity, multi-fidelity, and transfer learning HPO methods, establishing a rigorous baseline for future research in unsupervised biological representation learning.

cs.LG

Federated Learning in Genetics: Extended Analysis of Accuracy, Performance and Privacy Trade-offs

Machine learning on large-scale genomic or transcriptomic data is important for many novel health applications. For example, precision medicine tailors medical treatments to patients on the basis of individual biomarkers, cellular and molecular states, etc. However, the data required is sensitive, voluminous, heterogeneous, and typically distributed across locations where dedicated machine learning hardware is not available. Due to privacy and regulatory reasons, it is also problematic to aggregate all data at a trusted third party. Federated learning is a promising solution to this dilemma, because it enables decentralized, collaborative machine learning without exchanging raw data. In this paper, we perform comparative experiments with the federated learning frameworks TensorFlow Federated and Flower. Our test case is the training of disease prognosis and cell type classification models. We train the models with distributed transcriptomic data, considering both data heterogeneity and architectural heterogeneity. We measure model quality, robustness against privacy-enhancing noise and computational performance. We evaluate the resource overhead of a federated system from both client and global perspectives and assess benefits and limitations. Each of the federated learning frameworks has different strengths. However, our experiments confirm that both frameworks can readily build models on transcriptomic data, without transferring personal raw data to a third party with abundant computational resources. This paper is the extended version of https://link.springer.com/chapter/10.1007/978-3-031-63772-8_26.

cs.LG

Optimizing defence, counter-defence and counter-counter defence in parasitic and trophic interactions -- A modelling study

In host-pathogen interactions, often the host (attacked organism) defends itself by some toxic compound and the parasite, in turn, responds by producing an enzyme that inactivates that compound. In some cases, the host can respond by producing an inhibitor of that enzyme, which can be considered as a counter-counter defence. An example is provided by cephalosporins, beta-lactamases and clavulanic acid (an inhibitor of beta-lactamases). Here, we tackle the question under which conditions it pays, during evolution, to establish a counter-counter defence rather than to intensify or widen the defence mechanisms. We establish a mathematical model describing this phenomenon, based on enzyme kinetics for competitive inhibition. We use an objective function based on Haber's rule, which says that the toxic effect is proportional to the time integral of toxin concentration. The optimal allocation of defence and counter-counter defence can be calculated in an analytical way despite the nonlinearity in the underlying differential equation. The calculation provides a threshold value for the dissociation constant of the inhibitor. Only if the inhibition constant is below that threshold, that is, in the case of strong binding of the inhibitor, it pays to have a counter-counter defence. This theoretical prediction accounts for the observation that not for all defence mechanisms, a counter-counter defence exists. Our results should be of interest for computing optimal mixtures of beta-lactam antibiotics and beta-lactamase inhibitors such as sulbactam, as well as for plant-herbivore and other molecular-ecological interactions and to fight antibiotic resistance in general.

q-bio.SC