SearcharxivSearch

arXiv subjects

Claudius Gros

Publications and source records attributed to Claudius Gros.

At least 19 recordsLinked to original sources

Distance generalization in transformers: why bother with positional encoding?

Out-of-distribution length generalization, namely to extrapolate a task from short to longer context, has been studied intensively for transformers. Here we focus on distance generalization, which probes performance when inter-token distances are changed between training and inference, while keeping a fixed context length. We construct two synthetic delay copy tasks, both involving finite distances between source and recall, where tokens are copied either fully or selectively, and test models on delays unseen during training. We address three questions: (A) Do positional encoding schemes such as RoPE and ALiBi improve distance resolution relative to no positional encoding (NoPE)? (B) How does data diversity, the number of inter-token distances seen in training, affect performance? (C) When is distance transfer learning positive or negative? We present a thorough investigation, finding that it is paramount to improve our understanding of the underlying mechanisms.

cs.CL

Maximally chaotic competition for attention in the cultural domain

Memory is a key determinant when cultural items compete for attention and, consequently, for success, as in the case of songs on a music chart. For modeling, one adds memory to Lotka-Volterra models, the reference for Markovian competitive processes. Here we treat memory in terms of an exponential moving average, finding that it leads to an extended region of winnerless chaos characterized by log-normal popularity statistics in trailing top-k charts. Importantly, the observed log-normal behavior collapses to a power-law distribution when the feedback dynamics is fast on the scale of the charting period. This result is in agreement with the observed statistics of real-world music charts (e.g., Billboard and Spotify). In a chaotic state, the size of the largest Lyapunov exponent is a measure of how unpredictable the system is. We find that the largest Lyapunov exponent varies strongly as a function of parameters in the phase where winnerless chaos is stable. Interestingly, the sets of parameters obtained by comparing simulations with real-world cultural-item dynamics extracted from Google Books and Google Trends, movies, Reddit, Wikipedia, Twitter, and scientific publications, are located close to the points in parameter space where the largest Lyapunov exponent reaches its local maximum, namely, close to the point of maximal unpredictability. This result suggests that cultural competitive processes are maximally chaotic when memory is a key determinant.

nlin.CD

From generative AI to the brain: five takeaways

The big strides seen in generative AI are not based on somewhat obscure algorithms, but due to clearly defined generative principles. The resulting concrete implementations have proven themselves in large numbers of applications. We suggest that it is imperative to thoroughly investigate which of these generative principles may be operative also in the brain, and hence relevant for cognitive neuroscience. In addition, ML research led to a range of interesting characterizations of neural information processing systems. We discuss five examples, the shortcomings of world modelling, the generation of thought processes, attention, neural scaling laws, and quantization, that illustrate how much neuroscience could potentially learn from ML research.

cs.AI

Small transformer architectures for task switching

The rapid progress seen in terms of large-scale generative AI is largely based on the attention mechanism. It is conversely non-trivial to conceive small-scale applications for which attention-based architectures outperform traditional approaches, such as multi-layer perceptrons or recurrent networks. We examine this problem in the context of 'task switching'. In this framework models work on ongoing token sequences with the current task being determined by stochastically interspersed control tokens. We show that standard transformers cannot solve a basic task switching reference model based on finite domain arithmetics which contains subtasks dedicated to increment / addition / reverse copy / context (IARC). We show that transformers, long short-term memory recurrent networks (LSTM), and plain multi-layer perceptrons (MLPs) achieve similar, but only modest prediction accuracies. We enlarge our comparative study by including an extension of the standard transformer architecture to its non-translational invariant counterpart, the cisformer, and an alternative attention mechanism, extensive attention. A combination of the latter is found to be the only model able to achieve considerable performance levels, of around 95%. Our results indicate that the workings of attention can be understood better, and even improved, when comparing qualitatively different formulations in task-switching settings.

cs.LG

AlphaZero Neural Scaling and Zipf's Law: a Tale of Board Games and Power Laws

Neural scaling laws are observed in a range of domains, to date with no universal understanding of why they occur. Recent theories suggest that loss power laws arise from Zipf's law, a power law observed in domains like natural language. One theory suggests that language scaling laws emerge when Zipf-distributed task quanta are learned in descending order of frequency. In this paper we examine power-law scaling in AlphaZero, a reinforcement learning algorithm, using a model of language-model scaling. We find that game states in training and inference data scale with Zipf's law, which is known to arise from the tree structure of the environment, and examine the correlation between scaling-law and Zipf's-law exponents. In agreement with the quanta scaling model, we find that agents optimize state loss in descending order of frequency, even though this order scales inversely with modelling complexity. We also find that inverse scaling, the failure of models to improve with size, is correlated with unusual Zipf curves where end-game states are among the most frequent states. We show evidence that larger models shift their focus to these less-important states, sacrificing their understanding of important early-game states.

cs.LG

How oscillations in SIRS epidemic models are affected by the distribution of immunity times

Models for resident infectious diseases, like the SIRS model, may settle into an endemic state with constant numbers of susceptible ($S$), infected ($I$) and recovered ($R$) individuals, where recovered individuals attain a temporary immunity to reinfection. For many infectious pathogens, infection dynamics may also show periodic outbreaks corresponding to a limit cycle in phase space. One way to reproduce oscillations in SIRS models is to include a non-exponential dwell-time distribution in the recovered state. Here, we study a SIRS model with a step-function-like kernel for the immunity time, mapping out the model's full phase diagram. Using the kernel series framework, we are able to identify the onset of periodic outbreaks when successively broadening the step-width. We further investigate the shape of the outbreaks, finding that broader steps cause more sinusoidal oscillations while more uniform immunity time distributions are related to sharper outbreaks occurring after extended periods of low infection activity. Our main results concern recovery distributions characterized by a single dominant timescale. We also consider recovery distributions with two timescales, which may be observed when two or more distinct recovery processes co-exist. Surprisingly, two qualitatively different limit cycles are found to be stable in this case, with only one of the two limit cycles emerging via a standard supercritical Hopf bifurcation.

q-bio.PE

Self-organized attractoring in locomoting animals and robots: an emerging field

Locomotion may be induced on three levels. On a classical level, actuators and limbs follow the sequence of open-loop top-down control signals they receive. Limbs may move alternatively on their own, which implies that interlimb coordination must be mediated either by the body or via decentralized inter-limb signaling. In this case, when embodiment is present, two types of controllers are conceivable for the actuators of the limbs, local pacemaker circuits and control principles based on self-organized embodiment. The latter, self-organized control, is based on limit cycles and chaotic attractors that emerge within the feedback loop composed of controller, body, and environment. For this to happen, the sensorimotor loop must be locally closed, e.g. via propriosensation. Here we review the progress made within the framework of self-organized embodiment, with a particular focus on the concept of attractoring. This concept characterizes situations when sets of attractors combining discrete and continuous spectra are available as motor primitives for higher-order control schemes, such as kick control. In particular, we show that a simple generative principle allows for the robust formulation of self-organized embodiment. Based on the recurrent alternation between measuring the actual status of an actuator and providing a target for the actuator to achieve in the next step, we find that the mechanism leads to compliant locomotion for a range of simulated and real-world robots, which include barrel- and sphere-shaped agents, as well as wheeled and legged robots.

nlin.AO

Reorganizing attention-space geometry with expressive attention

Attention regulates information transfer between tokens. For this, query and key vectors are compared, typically in terms of a scalar product, $\mathbf{Q}^T\mathbf{K}$, together with a subsequent softmax normalization. In geometric terms, the standard dot-product attention (DPA) leads to large/small attention weights for parallel/antiparallel queries and keys. Here we study expressive attention (EA), which is based on $(\mathbf{Q}^T\mathbf{K})^2$, the squared dot product. In this case, attention is enhanced when query and key are either parallel or antiparallel, and suppressed for orthogonal configurations. EA can be introduced into any attention-based code without additional compute costs or memory requirements. For a series of autoregressive prediction tasks, we find that expressive attention performs at least as well as vanilla DPA. Increasing task complexity, EA is observed to outperform DPA with increasing margins, which also holds for multi-task settings. For a given model size, EA manages to achieve 100% performance for a range of complexity levels not accessible to DPA. Our results show that it is possible to reorganize the geometry of the matching condition in the space of attention heads without loss of performance.

cs.LG

A game of life with dormancy

The factors contributing to the persistence and stability of life are fundamental for understanding complex living systems. Organisms are commonly challenged by harsh and fluctuating environments that are suboptimal for growth and reproduction, which can lead to extinction. Species often contend with unfavorable and noisy conditions by entering a reversible state of reduced metabolic activity, a phenomenon known as dormancy. Here, we develop Spore Life, a model to investigate the effects of dormancy on population dynamics. It is based on Conway's Game of Life, a deterministic cellular automaton where simple rules govern the metabolic state of an individual based on the metabolic state of its neighbors. For individuals that would otherwise die, Spore Life provides a refuge in the form of an inactive state. These dormant individuals (spores) can resuscitate when local conditions improve. The model includes a parameter alpha that controls the survival probability of spores, interpolating between Game of Life (alpha = 0) and Spore Life (alpha = 1), while capturing stochastic dynamics in the intermediate regime (0 < alpha < 1). In addition to identifying the emergence of unique periodic configurations, we find that spore survival increases the average number of active individuals and buffers populations from extinction. Contrary to expectations, the stabilization of the population is not the result of a large and long-lived seed bank. Instead, the demographic patterns in Spore Life only require a small number of resuscitation events. Our approach yields novel insight into what is minimally required for the emergence of complex behaviors associated with dormancy and the seed banks that they generate.

q-bio.PE

Neural self-organization for muscle-driven robots

We present self-organizing control principles for simulated robots actuated by synthetic muscles. Muscles correspond to linear motors exerting force only when contracting, but not when expanding, with joints being actuated by pairs of antagonistic muscles. Individually, muscles are connected to a controller composed of a single neuron with a dynamical threshold that generates target positions for the respective muscle. A stable limit cycle is generated when the embodied feedback loop is closed, giving rise to regular locomotive patterns. In the absence of direct couplings between neurons, we show that force-mediated intra- and inter-leg couplings between muscles suffice to generate stable gaits.

nlin.AO

Mapping dynamical systems with distributed time delays to sets of ordinary differential equations

Real-world dynamical systems with retardation effects are described in general not by a single, precisely defined time delay, but by a range of delay times. An exact mapping onto a set of $N+1$ ordinary differential equations exists when the respective delay distribution is given in terms of a gamma distribution with discrete exponents. The number of auxiliary variables one needs to introduce, $N$, is inversely proportional to the variance of the delay distribution. The case of a single delay is therefore recovered when $N\to\infty$. Using this approach, denoted here the `kernel series framework', we examine systematically how the bifurcation phase diagram of the Mackey-Glass system changes under the influence of distributed delays. We find that local properties, f.i.\ the locus of a Hopf bifurcation, are robust against the introduction of broadened memory kernels. Period-doubling transitions and the onset of chaos, which involve non-local properties of the flow, are found in contrast to be more sensitive to distributed delays. In general, the observed effects are found to scale as $1/N$. Furthermore, we consider time-delayed systems exhibiting chaotic diffusion, which is present in particular for sinusoidal flows. We find that chaotic diffusion is substantially more pronounced for distributed delays. Our results indicate in consequence that modeling approaches of real-world processes should take the effects of distributed delay times into account.

math.DS

Scaling Laws for a Multi-Agent Reinforcement Learning Model

The recent observation of neural power-law scaling relations has made a significant impact in the field of deep learning. A substantial amount of attention has been dedicated as a consequence to the description of scaling laws, although mostly for supervised learning and only to a reduced extent for reinforcement learning frameworks. In this paper we present an extensive study of performance scaling for a cornerstone reinforcement learning algorithm, AlphaZero. On the basis of a relationship between Elo rating, playing strength and power-law scaling, we train AlphaZero agents on the games Connect Four and Pentago and analyze their performance. We find that player strength scales as a power law in neural network parameter count when not bottlenecked by available compute, and as a power of compute when training optimally sized agents. We observe nearly identical scaling exponents for both games. Combining the two observed scaling laws we obtain a power law relating optimal size to compute similar to the ones observed for language models. We find that the predicted scaling of optimal neural network size fits our data for both games. This scaling law implies that previously published state-of-the-art game-playing models are significantly smaller than their optimal size, given the respective compute budgets. We also show that large AlphaZero models are more sample efficient, performing better than smaller models with the same amount of training data.

cs.LG

Generic catastrophic poverty when selfish investors exploit a degradable common resource

The productivity of a common pool of resources may degrade when overly exploited by a number of selfish investors, a situation known as the tragedy of the commons (TOC). Without regulations, agents optimize the size of their individual investments into the commons by balancing incurring costs with the returns received. The resulting Nash equilibrium involves a self-consistency loop between individual investment decisions and the state of the commons. As a consequence, several non-trivial properties emerge. For $N$ investing actors we prove rigorously that typical payoffs do not scale as $1/N$, the expected result for cooperating agents, but as $(1/N)^2$. Payoffs are hence reduced with regard to the functional dependence on $N$, a situation denoted catastrophic poverty. We show that catastrophic poverty results from a fine-tuned balance between returns and costs. Additionally, a finite number of oligarchs may be present. Oligarchs are characterized by payoffs that are finite and not decreasing when $N$ increases. Our results hold for generic classes of models, including convex and moderately concave cost functions. For strongly concave cost functions the Nash equilibrium undergoes a collective reorganization, being characterized instead by entry barriers and sudden death forced market exits.

econ.TH

Collective strategy condensation towards class-separated societies

In physics, the wavefunctions of bosonic particles collapse when the system undergoes a Bose-Einstein condensation. In game theory, the strategy of an agent describes the probability to engage in a certain course of action. Strategies are expected to differ in competitive situations, namely when there is a penalty to do the same as somebody else. We study what happens when agents are interested how they fare not only in absolute terms, but also relative to others. This preference, denoted envy, is shown to induce the emergence of distinct social classes via a collective strategy condensation transition. Members of the lower class pursue identical strategies, in analogy to the Bose-Einstein condensation, with the upper class remaining individualistic.

econ.TH

When to end a lock down? How fast must vaccination campaigns proceed in order to keep health costs in check?

We propose a simple rule of thumb for countries which have embarked on a vaccination campaign while still facing the need to keep non-pharmaceutical interventions (NPI) in place because of the ongoing spread of SARS-CoV-2. If the aim is to keep the death rate from increasing, NPIs can be loosened when it is possible to vaccinate more than twice the growth rate of new cases. If the aim is to keep the pressure on hospitals under control, the vaccination rate has to be about four times higher. These simple rules can be derived from the observation that the risk of death or a severe course requiring hospitalization from a COVID-19 infection increases exponentially with age and that the sizes of age cohorts decrease linearly at the top of the population pyramid. Protecting the over 60-year-olds, which constitute approximately one-quarter of the population in Europe (and most OECD countries), reduces the potential loss of life by 95 percent.

q-bio.PE

Emotions as abstract evaluation criteria in biological and artificial intelligences

Biological as well as advanced artificial intelligences (AIs) need to decide which goals to pursue. We review nature's solution to the time allocation problem, which is based on a continuously readjusted categorical weighting mechanism we experience introspectively as emotions. One observes phylogenetically that the available number of emotional states increases hand in hand with the cognitive capabilities of animals and that raising levels of intelligence entail ever larger sets of behavioral options. Our ability to experience a multitude of potentially conflicting feelings is in this view not a leftover of a more primitive heritage, but a generic mechanism for attributing values to behavioral options that can not be specified at birth. In this view, emotions are essential for understanding the mind. For concreteness, we propose and discuss a framework which mimics emotions on a functional level. Based on time allocation via emotional stationarity (TAES), emotions are implemented as abstract criteria, such as satisfaction, challenge and boredom, which serve to evaluate activities that have been carried out. The resulting timeline of experienced emotions is compared with the `character' of the agent, which is defined in terms of a preferred distribution of emotional states. The long-term goal of the agent, to align experience with character, is achieved by optimizing the frequency for selecting individual tasks. Upon optimization, the statistics of emotion experience becomes stationary.

q-bio.NC

The economics of stop-and-go epidemic control

We analyse 'stop-and-go' containment policies that produce infection cycles as periods of tight lockdowns are followed by periods of falling infection rates. The subsequent relaxation of containment measures allows cases to increase again until another lockdown is imposed and the cycle repeats. The policies followed by several European countries during the Covid-19 pandemic seem to fit this pattern. We show that 'stop-and-go' should lead to lower medical costs than keeping infections at the midpoint between the highs and lows produced by 'stop-and-go'. Increasing the upper and reducing the lower limits of a stop-and-go policy by the same amount would lower the average medical load. But increasing the upper and lowering the lower limit while keeping the geometric average constant would have the opposite effect. We also show that with economic costs proportional to containment, any path that brings infections back to the original level (technically a closed cycle) has the same overall economic cost.

econ.TH

Charting closed-loop collective cultural decisions: From book best sellers and music downloads to Twitter hashtags and Reddit comments

Charts are used to measure relative success for a large variety of cultural items. Traditional music charts have been shown to follow self-organizing principles with regard to the distribution of item lifetimes, the on-chart residence times. Here we examine if this observation holds also for (a) music streaming charts (b) book best-seller lists and (c) for social network activity charts, such as Twitter hashtags and the number of comments Reddit postings receive. We find that charts based on the active production of items, like commenting, are more likely to be influenced by external factors, in particular by the 24 hour day-night cycle. External factors are less important for consumption-based charts (sales, downloads), which can be explained by a generic theory of decision-making. In this view, humans aim to optimize the information content of the internal representation of the outside world, which is logarithmically compressed. Further support for information maximization is argued to arise from the comparison of hourly, daily and weekly charts, which allow to gauge the importance of decision times with respect to the chart compilation period.

cs.SI