SearcharxivSearch

arXiv subjects

Ard A. Louis

Publications and source records attributed to Ard A. Louis.

At least 19 recordsLinked to original sources

Counting Connected and Disconnected Ways to Assemble a Jigsaw Puzzle

A jigsaw puzzle may be assembled in many different ways. Some assembly sequences remain connected throughout, while others temporarily build separate parts of the puzzle before joining them together. By representing the puzzle as a graph, these assembly sequences become vertex orderings that can be counted using recent results from graph theory. We apply this framework to enumerate three natural assembly strategies: connected assembly, assembly beginning from several disconnected pieces, and assembly in which new disconnected sections may be started during the process. The resulting counts reveal that, even for modest puzzle sizes, connectivity-preserving assembly sequences are greatly outnumbered by those that pass through disconnected intermediate stages.

math.GM

Successive vertex orderings of graphs

A successive vertex ordering of a graph is a linear ordering of its vertices in which every vertex except the first has at least one neighbour appearing earlier. Such orderings arise naturally in incremental growth and connectivity-preserving constructions, where vertices are added sequentially and must attach to the existing structure. We derive an exact formula for the number of successive vertex orderings of any finite connected graph. The formula is obtained via an inclusion--exclusion argument over independent sets and depends on two explicit combinatorial parameters, one of which is defined recursively. The result applies to all finite connected graphs without requiring regularity or symmetry assumptions. We also express the enumeration as a weighted generating polynomial over independent sets; its value at $x = -1$ recovers the total count of successive orderings, and the $k$-th derivative at this point encodes the number of orderings in which exactly $k$ non-first vertices appear before all of their neighbours.

math.CO

Decoupling Dynamical Richness from Representation Learning: Towards Practical Measurement

Dynamic feature transformation (the rich regime) does not always align with predictive performance (better representation), yet accuracy is often used as a proxy for richness, limiting analysis of their relationship. We propose a computationally efficient, performance-independent metric of richness grounded in the low-rank bias of rich dynamics, which recovers neural collapse as a special case. The metric is empirically more stable than existing alternatives and captures known lazy-torich transitions (e.g., grokking) without relying on accuracy. We further use it to examine how training factors (e.g., learning rate) relate to richness, confirming recognized assumptions and highlighting new observations (e.g., batch normalization promotes rich dynamics). An eigendecomposition-based visualization is also introduced to support interpretability, together providing a diagnostic tool for studying the relationship between training factors, dynamics, and representations.

stat.ML

Sufficient Conditions for Stability of Minimum-Norm Interpolating Deep ReLU Networks

Algorithmic stability is a classical framework for analyzing the generalization error of learning algorithms. It predicts that an algorithm has small generalization error if it is insensitive to small perturbations in the training set such as the removal or replacement of a training point. While stability has been demonstrated for numerous well-known algorithms, this framework has had limited success in analyses of deep neural networks. In this paper we study the algorithmic stability of deep ReLU homogeneous neural networks that achieve zero training error using parameters with the smallest $L_2$ norm, also known as the minimum-norm interpolation, a phenomenon that can be observed in overparameterized models trained by gradient-based algorithms. We investigate sufficient conditions for such networks to be stable. We find that 1) such networks are stable when they contain a (possibly small) stable sub-network, followed by a layer with a low-rank weight matrix, and 2) such networks are not guaranteed to be stable even when they contain a stable sub-network, if the following layer is not low-rank. The low-rank assumption is inspired by recent empirical and theoretical results which demonstrate that training deep neural networks is biased towards low-rank weight matrices, for minimum-norm interpolation and weight-decay regularization.

cs.LG

Feature learning is decoupled from generalization in high capacity neural networks

Neural networks outperform kernel methods, sometimes by orders of magnitude, e.g. on staircase functions. This advantage stems from the ability of neural networks to learn features, adapting their hidden representations to better capture the data. We introduce a concept we call feature quality to measure this performance improvement. We examine existing theories of feature learning and demonstrate empirically that they primarily assess the strength of feature learning, rather than the quality of the learned features themselves. Consequently, current theories of feature learning do not provide a sufficient foundation for developing theories of neural network generalization.

cs.LG

Deep neural networks have an inbuilt Occam's razor

The remarkable performance of overparameterized deep neural networks (DNNs) must arise from an interplay between network architecture, training algorithms, and structure in the data. To disentangle these three components, we apply a Bayesian picture, based on the functions expressed by a DNN, to supervised learning. The prior over functions is determined by the network, and is varied by exploiting a transition between ordered and chaotic regimes. For Boolean function classification, we approximate the likelihood using the error spectrum of functions on data. When combined with the prior, this accurately predicts the posterior, measured for DNNs trained with stochastic gradient descent. This analysis reveals that structured data, combined with an intrinsic Occam's razor-like inductive bias towards (Kolmogorov) simple functions that is strong enough to counteract the exponential growth of the number of functions with complexity, is a key to the success of DNNs.

cs.LG

An exactly solvable model for emergence and scaling laws in the multitask sparse parity problem

Deep learning models can exhibit what appears to be a sudden ability to solve a new problem as training time, training data, or model size increases, a phenomenon known as emergence. In this paper, we present a framework where each new ability (a skill) is represented as a basis function. We solve a simple multi-linear model in this skill-basis, finding analytic expressions for the emergence of new skills, as well as for scaling laws of the loss with training time, data size, model size, and optimal compute. We compare our detailed calculations to direct simulations of a two-layer neural network trained on multitask sparse parity, where the tasks in the dataset are distributed according to a power-law. Our simple model captures, using a single fit parameter, the sigmoidal emergence of multiple new skills as training time, data size or model size increases in the neural network.

cs.LG

Exploiting the equivalence between quantum neural networks and perceptrons

Quantum machine learning models based on parametrized quantum circuits, also called quantum neural networks (QNNs), are considered to be among the most promising candidates for applications on near-term quantum devices. Here we explore the expressivity and inductive bias of QNNs by exploiting an exact mapping from QNNs with inputs $x$ to classical perceptrons acting on $x \otimes x$ (generalised to complex inputs). The simplicity of the perceptron architecture allows us to provide clear examples of the shortcomings of current QNN models, and the many barriers they face to becoming useful general-purpose learning algorithms. For example, a QNN with amplitude encoding cannot express the Boolean parity function for $n\geq 3$, which is but one of an exponential number of data structures that such a QNN is unable to express. Mapping a QNN to a classical perceptron simplifies training, allowing us to systematically study the inductive biases of other, more expressive embeddings on Boolean data. Several popular embeddings primarily produce an inductive bias towards functions with low class balance, reducing their generalisation performance compared to deep neural network architectures which exhibit much richer inductive biases. We explore two alternate strategies that move beyond standard QNNs. In the first, we use a QNN to help generate a classical DNN-inspired kernel. In the second we draw an analogy to the hierarchical structure of deep neural networks and construct a layered non-linear QNN that is provably fully expressive on Boolean data, while also exhibiting a richer inductive bias than simple QNNs. Finally, we discuss characteristics of the QNN literature that may obscure how hard it is to achieve quantum advantage over deep learning algorithms on classical data.

quant-ph

Exploring simplicity bias in 1D dynamical systems

Arguments inspired by algorithmic information theory predict an inverse relation between the probability and complexity of output patterns in a wide range of input-output maps. This phenomenon is known as \emph{simplicity bias}. By viewing the parameters of dynamical systems as inputs, and resulting (digitised) trajectories as outputs, we study simplicity bias in the logistic map, Gauss map, sine map, Bernoulli map, and tent map. We find that the logistic map, Gauss map, and sine map all exhibit simplicity bias upon sampling of map initial values and parameter values, but the Bernoulli map and tent map do not. The simplicity bias upper bound on output pattern probability is used to make \emph{a priori} predictions for the probability of output patterns. In some cases, the predictions are surprisingly accurate, given that almost no details of the underlying dynamical systems are assumed. More generally, we argue that studying probability-complexity relationships may be a useful tool in studying patterns in dynamical systems.

math.DS

Coarse-grained modelling of DNA-RNA hybrids

We introduce oxNA, a new model for the simulation of DNA-RNA hybrids which is based on two previously developed coarse-grained models$\unicode{x2014}$oxDNA and oxRNA. The model naturally reproduces the physical properties of hybrid duplexes including their structure, persistence length and force-extension characteristics. By parameterising the DNA-RNA hydrogen bonding interaction we fit the model's thermodynamic properties to experimental data using both average-sequence and sequence-dependent parameters. To demonstrate the model's applicability we provide three examples of its use$\unicode{x2014}$calculating the free energy profiles of hybrid strand displacement reactions, studying the resolution of a short R-loop and simulating RNA-scaffolded wireframe origami.

cond-mat.soft

Double-descent curves in neural networks: a new perspective using Gaussian processes

Double-descent curves in neural networks describe the phenomenon that the generalisation error initially descends with increasing parameters, then grows after reaching an optimal number of parameters which is less than the number of data points, but then descends again in the overparameterized regime. In this paper, we use techniques from random matrix theory to characterize the spectral distribution of the empirical feature covariance matrix as a width-dependent perturbation of the spectrum of the neural network Gaussian process (NNGP) kernel, thus establishing a novel connection between the NNGP literature and the random matrix theory literature in the context of neural networks. Our analytical expression allows us to study the generalisation behavior of the corresponding kernel and GP regression, and provides a new interpretation of the double-descent phenomenon, namely as governed by the discrepancy between the width-dependent empirical kernel and the width-independent NNGP kernel.

stat.ML

Robustness and Stability of Spin Glass Ground States to Perturbed Interactions

Across many scientific and engineering disciplines, it is important to consider how much the output of a given system changes due to perturbations of the input. Here, we investigate the glassy phase of $\pm J$ spin glasses at zero temperature by calculating the robustness of the ground states to flips in the sign of single interactions. For random graphs and the Sherrington-Kirkpatrick model, we find relatively large sets of bond configurations that generate the same ground state. These sets can themselves be analyzed as subgraphs of the interaction domain, and we compute many of their topological properties. In particular, we find that the robustness, equivalent to the average degree, of these subgraphs is much higher than one would expect from a random model. Most notably, it scales in the same logarithmic way with the size of the subgraph as has been found in genotype-phenotype maps for RNA secondary structure folding, protein quaternary structure, gene regulatory networks, as well as for models for genetic programming. The similarity between these disparate systems suggests that this scaling may have a more universal origin.

cond-mat.dis-nn

Designing the self-assembly of arbitrary shapes using minimal complexity building blocks

The design space for a self-assembled multicomponent objects ranges from a solution in which every building block is unique to one with the minimum number of distinct building blocks that unambiguously define the target structure. Using a novel pipeline, we explore the design spaces for a set of structures of various sizes and complexities. To understand the implications of the different solutions, we analyse their assembly dynamics using patchy particle simulations and study the influence of the number of distinct building blocks and the angular and spatial tolerances on their interactions on the kinetics and yield of the target assembly. We show that the resource-saving solution with minimum number of distinct blocks can often assemble just as well (or faster) than designs where each building block is unique. We further use our methods to design multifarious structures, where building blocks are shared between different target structures. Finally, we use coarse-grained DNA simulations to investigate the realisation of multicomponent shapes using DNA nanostructures as building blocks.

cond-mat.soft

Free-energy landscapes of DNA and its assemblies: Perspectives from coarse-grained modelling

This chapter will provide an overview of how characterizing free-energy landscapes can provide insights into the biophysical properties of DNA, as well as into the behaviour of the DNA assemblies used in the field of DNA nanotechnology. The landscapes for these complex systems are accessible through the use of accurate coarse-grained descriptions of DNA. Particular foci will be the landscapes associated with DNA self-assembly and mechanical deformation, where the latter can arise from either externally imposed forces or internal stresses.

cond-mat.soft

From genotypes to organisms: State-of-the-art and perspectives of a cornerstone in evolutionary dynamics

Understanding how genotypes map onto phenotypes, fitness, and eventually organisms is arguably the next major missing piece in a fully predictive theory of evolution. We refer to this generally as the problem of the genotype-phenotype map. Though we are still far from achieving a complete picture of these relationships, our current understanding of simpler questions, such as the structure induced in the space of genotypes by sequences mapped to molecular structures, has revealed important facts that deeply affect the dynamical description of evolutionary processes. Empirical evidence supporting the fundamental relevance of features such as phenotypic bias is mounting as well, while the synthesis of conceptual and experimental progress leads to questioning current assumptions on the nature of evolutionary dynamics-cancer progression models or synthetic biology approaches being notable examples. This work delves into a critical and constructive attitude in our current knowledge of how genotypes map onto molecular phenotypes and organismal functions, and discusses theoretical and empirical avenues to broaden and improve this comprehension. As a final goal, this community should aim at deriving an updated picture of evolutionary processes soundly relying on the structural properties of genotype spaces, as revealed by modern techniques of molecular and functional analysis.

q-bio.PE

Generalization bounds for deep learning

Generalization in deep learning has been the topic of much recent theoretical and empirical research. Here we introduce desiderata for techniques that predict generalization errors for deep learning models in supervised learning. Such predictions should 1) scale correctly with data complexity; 2) scale correctly with training set size; 3) capture differences between architectures; 4) capture differences between optimization algorithms; 5) be quantitatively not too far from the true error (in particular, be non-vacuous); 6) be efficiently computable; and 7) be rigorous. We focus on generalization error upper bounds, and introduce a categorisation of bounds depending on assumptions on the algorithm and data. We review a wide range of existing approaches, from classical VC dimension to recent PAC-Bayesian bounds, commenting on how well they perform against the desiderata. We next use a function-based picture to derive a marginal-likelihood PAC-Bayesian bound. This bound is, by one definition, optimal up to a multiplicative constant in the asymptotic limit of large training sets, as long as the learning curve follows a power law, which is typically found in practice for deep learning problems. Extensive empirical analysis demonstrates that our marginal-likelihood PAC-Bayes bound fulfills desiderata 1-3 and 5. The results for 6 and 7 are promising, but not yet fully conclusive, while only desideratum 4 is currently beyond the scope of our bound. Finally, we comment on why this function-based bound performs significantly better than current parameter-based PAC-Bayes bounds.

stat.ML

Is SGD a Bayesian sampler? Well, almost

Overparameterised deep neural networks (DNNs) are highly expressive and so can, in principle, generate almost any function that fits a training dataset with zero error. The vast majority of these functions will perform poorly on unseen data, and yet in practice DNNs often generalise remarkably well. This success suggests that a trained DNN must have a strong inductive bias towards functions with low generalisation error. Here we empirically investigate this inductive bias by calculating, for a range of architectures and datasets, the probability $P_{SGD}(f\mid S)$ that an overparameterised DNN, trained with stochastic gradient descent (SGD) or one of its variants, converges on a function $f$ consistent with a training set $S$. We also use Gaussian processes to estimate the Bayesian posterior probability $P_B(f\mid S)$ that the DNN expresses $f$ upon random sampling of its parameters, conditioned on $S$. Our main findings are that $P_{SGD}(f\mid S)$ correlates remarkably well with $P_B(f\mid S)$ and that $P_B(f\mid S)$ is strongly biased towards low-error and low complexity functions. These results imply that strong inductive bias in the parameter-function map (which determines $P_B(f\mid S)$), rather than a special property of SGD, is the primary explanation for why DNNs generalise so well in the overparameterised regime. While our results suggest that the Bayesian posterior $P_B(f\mid S)$ is the first order determinant of $P_{SGD}(f\mid S)$, there remain second order differences that are sensitive to hyperparameter tuning. A function probability picture, based on $P_{SGD}(f\mid S)$ and/or $P_B(f\mid S)$, can shed new light on the way that variations in architecture or hyperparameter settings such as batch size, learning rate, and optimiser choice, affect DNN performance.

cs.LG

Measuring internal forces in single-stranded DNA: Application to a DNA force clamp

We present a new method for calculating internal forces in DNA structures using coarse-grained models and demonstrate its utility with the oxDNA model. The instantaneous forces on individual nucleotides are explored and related to model potentials, and using our framework, internal forces are calculated for two simple DNA systems and for a recently-published nanoscopic force clamp. Our results highlight some pitfalls associated with conventional methods for estimating internal forces, which are based on elastic polymer models, and emphasise the importance of carefully considering secondary structure and ionic conditions when modelling the elastic behaviour of single-stranded DNA. Beyond its relevance to the DNA nanotechnological community, we expect our approach to be broadly applicable to calculations of internal force in a variety of structures -- from DNA to protein -- and across other coarse-grained simulation models.

cond-mat.soft