SearcharxivSearch

arXiv subjects

Kathrin Gruber

Publications and source records attributed to Kathrin Gruber.

3 recordsLinked to original sources

Cascaded Flow Matching for Heterogeneous Tabular Data with Mixed-Type Features

Advances in generative modeling have recently been adapted to tabular data containing discrete and continuous features. However, generating mixed-type features that combine discrete states with an otherwise continuous distribution in a single feature remains challenging. We advance the state-of-the-art in diffusion models for tabular data with a cascaded approach. We first generate a low-resolution version of a tabular data row, that is, the collection of the purely categorical features and a coarse categorical representation of numerical features. Next, this information is leveraged in the high-resolution flow matching model via a novel guided conditional probability path and data-dependent coupling. The low-resolution representation of numerical features explicitly accounts for discrete outcomes, such as missing or inflated values, and therewith enables a more faithful generation of mixed-type features. We formally prove that this cascade tightens the transport cost bound. The results indicate that our model generates significantly more realistic samples and captures distributional details more accurately, for example, the detection score improves by 51.9\%. Code is available at https://github.com/muellermarkus/tabcascade.

cs.LG

Continuous Diffusion for Mixed-Type Tabular Data

Score-based generative models, commonly referred to as diffusion models, have proven to be successful at generating text and image data. However, their adaptation to mixed-type tabular data remains underexplored. In this work, we propose CDTD, a Continuous Diffusion model for mixed-type Tabular Data. CDTD is based on a novel combination of score matching and score interpolation to enforce a unified continuous noise distribution for both continuous and categorical features. We explicitly acknowledge the necessity of homogenizing distinct data types by relying on model-specific loss calibration and initialization schemes. To further address the high heterogeneity in mixed-type tabular data, we introduce adaptive feature- or type-specific noise schedules. These ensure balanced generative performance across features and optimize the allocation of model capacity across features and diffusion time. Our experimental results show that CDTD consistently outperforms state-of-the-art benchmark models, captures feature correlations exceptionally well, and that heterogeneity in the noise schedule design boosts sample quality. Replication code is available at https://github.com/muellermarkus/cdtd.

cs.LG

Predicting patterns for molecular self-organization on surfaces using interaction-site models

Molecular building blocks interacting at the nanoscale organize spontaneously into stable mono- layers that display intriguing long-range ordering motifs on the surface of atomic substrates. The patterning process, if appropriately controlled, represents a viable route to manufacture practical nanodevices. With this goal in mind, we seek to capture the salient features of the self-assembly process by means of an interaction-site model. The geometry of the building blocks, the symmetry of the underlying substrate, and the strength and range of interactions encode the self-assembly pro- cess. By means of Monte Carlo simulations, we have predicted an ample variety of ordering motifs which nicely reproduce the experimental results. Here, we explore in detail the phase behavior of the system in terms of the temperature and the lattice constant of the underlying substrate. Our method is suitable to investigate the stability of the emergent patterns as well as to identify the nature of the melting transition monitoring appropriate order parameters.

cond-mat.mes-hall