SearcharxivSearch

arXiv subjects

Tassilo Schwarz

Publications and source records attributed to Tassilo Schwarz.

6 recordsLinked to original sources

Milestoning Markov-jump dynamics: Stationary properties, thermodynamic consistency, kinetic hysteresis, and fluctuation symmetries

We derive an exact coarse graining of generic Markov-jump processes into observable semi-Markov dynamics. Exact results for waiting-time distributions for jumps between observable states are derived and proved that these decompose into conditionally independent dwell and transition times. Dwell times are proved to be a local property of mesostates - they depend on the initial but not final state. Conversely, transition-path times depend on both states, trigger kinetic hysteresis, and, under suitable conditions on the hidden sub-network, are shown to obey a reflection symmetry. We characterize the stationary properties of the milestoned dynamics, prove its thermodynamic consistency, and demonstrate robustness to milestone positioning. Surprisingly, even in the limit of a time-scale separation rendering the observed dynamics approximately Markovian, the effect of kinetic hysteresis on the dissipation persists. A minimal example shows how the results lay the foundation for inferring affinities of hidden dissipative cycles from observations of transition-path times.

cond-mat.stat-mech

Optimal Initialization in Depth: Lyapunov Initialization and Limit Theorems for Deep Leaky ReLU Networks

Effective initialization in deep networks requires an understanding of random neural networks. In this work, a rigorous probabilistic analysis of deep bias-free random Leaky ReLU networks is provided. We prove a Law of Large Numbers and a Central Limit Theorem for the logarithm of the norm of network activations, establishing that, as the number of layers increases, their growth is governed by a parameter called the Lyapunov exponent. This parameter characterizes a sharp phase transition between vanishing and exploding activations, and we calculate the Lyapunov exponent explicitly for Gaussian or orthogonal weight matrices. Our results reveal that standard methods, such as He initialization or orthogonal initialization, do not guarantee activation stability for deep networks of low width. Based on these theoretical insights, we propose a novel initialization method, referred to as Lyapunov initialization, which sets the Lyapunov exponent to zero and thereby ensures that the neural network is as stable as possible, leading empirically to improved learning.

stat.ML

Permutation-Invariant Spectral Learning via Dyson Diffusion

Diffusion models are central to generative modeling and have been adapted to graphs by diffusing adjacency matrix representations. The challenge of having up to $n!$ such representations for graphs with $n$ nodes is only partially mitigated by using permutation-equivariant learning architectures. Despite their computational efficiency, existing graph diffusion models struggle to distinguish certain graph families and their spectra, unless graph data are augmented with ad hoc features. This shortcoming stems from enforcing the inductive bias within the learning architecture. In this work, we leverage random matrix theory to analytically extract the spectral properties of the diffusion process, allowing us to push most of the inductive bias from the architecture into the dynamics. Building on this, we introduce the Dyson Diffusion Model, which employs Dyson's Brownian motion to capture the spectral dynamics of an Ornstein-Uhlenbeck process on the adjacency matrix. Furthermore, conditioned on the spectral dynamics, we formulate a Lie group diffusion, appropriately modeling the remaining degrees of freedom. Strikingly, the resulting learning problem becomes permutation invariant at the Lie algebra level. We demonstrate that the Dyson Diffusion Model learns graph spectra accurately and outperforms existing graph diffusion models.

stat.ML

Consistent time reversal and reliable and accurate inference in the presence of memory

Thermodynamic inference from coarse observations remains a key challenge. Memory, in particular correlations between consecutively observed mesostates, blur signatures of irreversibility and must be accounted for in defining physical time-reversal, which remains an open problem. We derive an experimentally accessible k-th order estimator for the entropy production rate. Using novel measure-theoretic techniques we prove necessary and sufficient conditions for guaranteed lower bounds on the dissipation even in the strongly non-Markovian setting. The proof reveals that estimators saturated in the order unravel the duration of memory which needs to be considered in defining physically consistent time-reversal. We show that Markovian estimators in absence of a time-scale separation lead to artifacts, which convey no physical meaning. Similarly, estimators not saturated in the order may overestimate the dissipation. The necessity of correctly accounting for memory in thermodynamic inference from strongly non-Markovian observations underscores the still underappreciated challenges and intricacies in defining and understanding irreversibility in presence of memory. Our results will hopefully stimulate experiments systematically considering thermodynamic inference on multiple scales consistently accounting for memory.

cond-mat.stat-mech

Complex-Weighted Convolutional Networks: Provable Expressiveness via Complex Diffusion

Graph Neural Networks (GNNs) have achieved remarkable success across diverse applications, yet they remain limited by oversmoothing and poor performance on heterophilic graphs. To address these challenges, we introduce a novel framework that equips graphs with a complex-weighted structure, assigning each edge a complex number to drive a diffusion process that extends random walks into the complex domain. We prove that this diffusion is highly expressive: with appropriately chosen complex weights, any node-classification task can be solved in the steady state of a complex random walk. Building on this insight, we propose the Complex-Weighted Convolutional Network (CWCN), which learns suitable complex-weighted structures directly from data while enriching diffusion with learnable matrices and nonlinear activations. CWCN is simple to implement, requires no additional hyperparameters beyond those of standard GNNs, and achieves competitive performance on benchmark datasets. Our results demonstrate that complex-weighted diffusion provides a principled and general mechanism for enhancing GNN expressiveness, opening new avenues for models that are both theoretically grounded and practically effective.

cs.LG

Randomized Controlled Trials Under Influence: Covariate Factors and Graph-Based Network Interference

Randomized controlled trials are not only the golden standard in medicine and vaccine trials but have spread to many other disciplines like behavioral economics, making it an important interdisciplinary tool for scientists. When designing randomized controlled trials, how to assign participants to treatments becomes a key issue. In particular in the presence of covariate factors, the assignment can significantly influence statistical properties and thereby the quality of the trial. Another key issue is the widely popular assumption among experimenters that participants do not influence each other -- which is far from reality in a field study and can, if unaccounted for, deteriorate the quality of the trial. We address both issues in our work. After introducing randomized controlled trials bridging terms from different disciplines, we first address the issue of participant-treatment assignment in the presence of known covariate factors. Thereby, we review a recent assignment algorithm that achieves good worst-case variance bounds. Second, we address social spillover effects. Therefore, we build a comprehensive graph-based model of influence between participants, for which we design our own average treatment effect estimator $\hat τ_{net}$. We discuss its bias and variance and reduce the problem of variance minimization to a certain instance of minimizing the norm of a matrix-vector product, which has been considered in literature before. Further, we discuss the role of disconnected components in the model's underlying graph.

stat.ME