SearcharxivSearch

arXiv subjects

Jonas Landsgesell

Publications and source records attributed to Jonas Landsgesell.

7 recordsLinked to original sources

Distributional Regression with Tabular Foundation Models: Evaluating Probabilistic Predictions via Proper Scoring Rules

Modern tabular foundation models such as TabPFN and TabICL naturally produce full predictive distributions, while the benchmarks used to evaluate them (TabArena, TALENT, and others) still rely almost exclusively on point-estimate metrics (RMSE, $R^2$). This mismatch implicitly rewards machine learning models or pipelines that elicit a good conditional mean while ignoring the quality of the predictive distribution. We make the case for using proper scoring rules for training, fine-tuning, and benchmarking (ranking) of tabular foundation models. Although all strictly proper scoring rules are theoretically equivalent at the population level, they may differ on finite data: We demonstrate analytically and empirically that different scoring rules can induce different inductive biases during finite-sample optimization, leading to different model performance. We validate this finding by running fine-tuning experiments with TabPFN and TabICL using different scoring rules for various data sets, revealing non-trivial interactions between training objectives and evaluation metrics. Our results show that practitioners can adapt tabular foundation models to task-specific scoring objectives, and that the choice of scoring rule can influence model behavior in practice.

cs.LG

ScoringBench: A Benchmark for Evaluating Tabular Foundation Models with Proper Scoring Rules

Tabular foundation models such as TabPFN and TabICL already produce full predictive distributions, yet prevailing regression benchmarks evaluate them almost exclusively via point-estimate metrics (RMSE, $R^2$). This discards precisely the distributional information these models are designed to provide - a critical gap for high-stakes domains where not all kinds of errors are equally costly. We introduce ScoringBench, an open and extensible benchmark that evaluates tabular regression models under a comprehensive suite of proper scoring rules - including CRPS, CRLS, interval score, energy score, and weighted CRPS - alongside standard point metrics. ScoringBench covers 97 regression datasets from diverse domains, supports transparent community contributions via a git-based leaderboard, and provides two complementary ranking protocols: an ordinal Demsar/autorank approach and a magnitude-preserving z-score ranking approach. Evaluating several models - spanning in-context learners, fine-tuned foundation models, gradient-boosted trees, and MLPs - we find that model rankings shift substantially depending on the scoring rule: models that excel on point-estimate metrics can rank poorly on probabilistic ones, and the top-performing model under one proper scoring rule may rank noticeably lower under another. These results demonstrate that the choice of evaluation metric is not a technicality but a modelling decision - and, for applications where e.g. tail errors are disproportionately costly, a domain-specific requirement with direct consequences for model deployment.

cs.AI

Joint Optimization of Neural Autoregressors via Scoring rules

Non-parametric distributional regression has achieved significant milestones in recent years. Among these, the Tabular Prior-Data Fitted Network (TabPFN) has demonstrated state-of-the-art performance on various benchmarks. However, a challenge remains in extending these grid-based approaches to a truly multivariate setting. In a naive non-parametric discretization with $N$ bins per dimension, the complexity of an explicit joint grid scales exponentially and the paramer count of the neural networks rise sharply. This scaling is particularly detrimental in low-data regimes, as the final projection layer would require many parameters, leading to severe overfitting and intractability.

cond-mat.soft

Cell Model Approaches for Predicting the Swelling and Mechanical Properties of Polyelectrolyte Gels

We present two successive mean-field approximations for describing the mechanical properties and the swelling equilibrium of polyelectrolyte gels in contact with a salt solution. The first mean-field approximation reduces the many-chain problem of a gel to a corresponding single chain problem. The second mean-field step integrates out the degrees of freedom of the flexible chain and the ions. It replaces the particle-based description of the polyelectrolyte with suitable charge distributions and an effective elasticity term. These simplifications result in a computationally very efficient Poisson-Boltzmann cell-gel description. Despite their simplicity, the single chain cell-gel model shows excellent and the PB model very good agreement with explicit molecular dynamics simulations of the reference periodic monodisperse network model for varying chain length, polymer charge fraction, and external reservoir salt concentrations. Comparisons of our models to the Katchalsky model reveal that our approach is superior for strongly charged chains and can also predict the bulk moduli more accurately. We further discuss chain length polydispersity effects, investigate changes in the solvent permittivity, and demonstrate the robustness of our approach to parameter variations coming from several modeling assumptions.

cond-mat.soft

Modeling Gel Swelling Equilibrium in Mean-Field: From explicit Models to Poisson-Boltzmann

We develop a double mean-field theory for charged macrogels immersed in electrolyte solutions in the spirit of the cell model approach. We first demonstrate that the equilibrium sampling of a single explicit coarse-grained charged polymer in a cell yields accurate predictions of the swelling equilibrium if the geometry is suitably chosen and all pressure contributions have been incorporated accurately. We then replace the explicit flexible chain by a suitably modeled penetrable charged rod that allows to compute all pressure terms within the Poisson-Boltzmann approximation. This model, albeit computationally cheap, yields excellent predictions of swelling equilibria under varying chain length, polymer charge fraction, and external reservoir salt concentrations when compared to coarse-grained molecular dynamics simulations of charged macrogels. We present an extension of the model to the experimentally relevant cases of pH-sensitive gels.

cond-mat.soft

ESPResSo 4.0 -- An Extensible Software Package for Simulating Soft Matter Systems

ESPResSo 4.0 is an extensible simulation package for research on soft matter. This versatile molecular dynamics program was originally developed for coarse-grained simulations of charged systems Limbach et al., Comput. Phys. Commun. 174, 704 (2006). The scope of the software has since broadened considerably: ESPResSo can now be used to simulate systems with length scales spanning from the molecular to the colloidal. Examples include, self-propelled particles in active matter, membranes in biological systems, and the aggregation of soot particles in process engineering. ESPResSo also includes solvers for hydrodynamic and electrokinetic problems, both on the continuum and on the explicit particle level. Since our last description of version 3.1 Arnold et al., Meshfree Methods for Partial Differential Equations VI, Lect. Notes Comput. Sci. Eng. 89, 1 (2013), the software has undergone considerable restructuring. The biggest change is the replacement of the Tcl scripting interface with a much more powerful Python interface. In addition, many new simulation methods have been implemented. In this article, we highlight the changes and improvements made to the interface and code, as well as the new simulation techniques that enable a user of ESPResSo 4.0 to simulate physics that is at the forefront of soft matter research.

cond-mat.soft

Simulation of weak polyelectrolytes: A comparison between the constant pH and the reaction ensemble method

The reaction ensemble and the constant pH method are well-known chemical equilibrium approaches to simulate protonation and deprotonation reactions in classical molecular dynamics and Monte Carlo simulations. In this article, we show similarity between both methods {under certain conditions}. We perform molecular dynamics simulations of a weak polyelectrolyte in order to compare the titration curves obtained by both approaches. Our findings reveal a good agreement between the methods when the reaction ensemble is used to sweep the reaction constant. Pronounced differences between the reaction ensemble and the constant pH method can be observed for stronger acids and bases in terms of adaptive pH values. These deviations are due to the presence of explicit protons in the reaction ensemble method which induce a screening of electrostatic interactions between the charged titrable groups of the polyelectrolyte. The outcomes of our simulation hint to a better applicability of the reaction ensemble method for systems in confined geometries and titrable groups in polyelectrolytes with different pK$_\text{a}$ values.

cond-mat.soft