SearcharxivSearch

arXiv subjects

Daoyuan Qian

Publications and source records attributed to Daoyuan Qian.

5 recordsLinked to original sources

Modular connectivity in neural networks emerges from Poisson noise-motivated regularisation, and promotes robustness and compositional generalisation

Circuits in the brain commonly exhibit modular architectures that factorise complex tasks, resulting in the ability to compositionally generalise and reduce catastrophic forgetting. In contrast, artificial neural networks (ANNs) appear to mix all processing, because modular solutions are difficult to find as they are vanishing subspaces in the space of possible solutions. Here, we draw inspiration from fault-tolerant computation and the Poisson-like firing of real neurons to show that activity-dependent neural noise, combined with nonlinear neural responses, drives the emergence of solutions that reflect an accurate understanding of modular tasks, corresponding to acquisition of a correct world model. We find that noise-driven modularisation can be recapitulated by a deterministic regulariser that multiplicatively combines weights and activations, revealing rich phenomenology not captured in linear networks or by standard regularisation methods. Though the emergence of modular structure requires sufficiently many training samples (exponential in the number of modular task dimensions), we show that pre-modularised ANNs exhibit superior noise-robustness and the ability to generalise and extrapolate well beyond training data, compared to ANNs without such inductive biases. Together, our work demonstrates a regulariser and architectures that could encourage modularity emergence to yield functional benefits.

physics.bio-ph

Compositional Generalization via Forced Rendering of Disentangled Latents

Composition-the ability to generate myriad variations from finite means-is believed to underlie powerful generalization. However, compositional generalization remains a key challenge for deep learning. A widely held assumption is that learning disentangled (factorized) representations naturally supports this kind of extrapolation. Yet, empirical results are mixed, with many generative models failing to recognize and compose factors to generate out-of-distribution (OOD) samples. In this work, we investigate a controlled 2D Gaussian "bump" generation task with fully disentangled (x,y) inputs, demonstrating that standard generative architectures still fail in OOD regions when training with partial data, by re-entangling latent representations in subsequent layers. By examining the model's learned kernels and manifold geometry, we show that this failure reflects a "memorization" strategy for generation via data superposition rather than via composition of the true factorized features. We show that when models are forced-through architectural modifications with regularization or curated training data-to render the disentangled latents into the full-dimensional representational (pixel) space, they can be highly data-efficient and effective at composing in OOD regions. These findings underscore that disentangled latents in an abstract representation are insufficient and show that if models can represent disentangled factors directly in the output representational space, it can achieve robust compositional generalization.

cs.LG

Fundamental performance bounds on time-series generation using reservoir computing

Reservoir computing (RC) harnesses the intrinsic dynamics of a chaotic system, called the reservoir, to perform various time-varying functions. An important use-case of RC is the generation of target temporal sequences via a trainable output-to-reservoir feedback loop. Despite the promise of RC in various domains, we lack a theory of performance bounds on RC systems. Here, we formulate an existence condition for a feedback loop that produces the target sequence. We next demonstrate that, given a sufficiently chaotic neural network reservoir, two separate factors are needed for successful training: global network stability of the target orbit, and the ability of the training algorithm to drive the system close enough to the target, which we term `reach'. By computing the training phase diagram over a range of target output amplitudes and periods, we verify that reach-limited failures depend on the training algorithm while stability-limited failures are invariant across different algorithms. We leverage dynamical mean field theory (DMFT) to provide an analytical amplitude-period bound on achievable outputs by RC networks and propose a way of enhancing algorithm reach via forgetting. The resulting mechanistic understanding of RC performance can guide the future design and deployment of reservoir networks.

nlin.CD

Phase transitions in rolling of irregular cylinders and spheres

When placed on an inclined plane, a perfect 2D disk or 3D sphere simply rolls down in a straight line under gravity. But how is the rolling affected if these shapes are irregular or random? Treating the terminal rolling speed as an order parameter, we show that phase transitions arise as a function of the dimension of the state space and inertia. We calculate the scaling exponents and the macroscopic lag time associated with the presence of first and second order transitions, and describe the regimes of co-existence of stable states and the accompanying hysteresis. Experiments with rolling cylinders corroborate our theoretical results on the scaling of the lag time. Experiments with spheres reveal closed orbits and their period-doubling in the overdamped and inertial limits respectively, providing visible manifestations of the hairy ball theorem and the doubly-connected nature of SO(3), the space of 3-dimensional rotations. Going beyond simple curiosity, our study might be relevant in a number of natural and artificial systems that involve the rolling of irregular objects, in systems ranging from nanoscale cellular transport to robotics.

physics.class-ph

Modelling Mullins Effect Induced by Chain Delamination and Reattachment

We propose a continuum theory to model the Mullins effect, which is ubiquitously observed in polymer composites. In the theory, the softening of the materials during the stretching process is accounted for by considering the delamination of polymer chains from nano-/micro-sized fillers, and the recovery effect during the de-stretching process is due to the reattachment of the polymer chains to nano-/micro-sized fillers. By incorporating the chain entanglements, Log-Normal distribution of the mesh size in the network, etc., we can obtain a good agreement between our numerical calculation results and existing experimental data. This physical theory can be easily adapted to meet more practical needs and utilised in analysing mechanic properties of polymer composites.

cond-mat.soft