SearcharxivSearch

arXiv subjects

Daniele Tantari

Publications and source records attributed to Daniele Tantari.

At least 19 recordsLinked to original sources

A solvable model for unsupervised federated learning

We introduce a theoretical framework for analyzing federated learning in a generative setting through a teacher-multiple interacting students scenario, in which each student receives a distinct realization of the data, either through a different noise corruption or by accessing a different subset, possibly of varying size. Using theoretical tools in equilibrium disordered system, we analytically show that interactions among students systematically enhance learning performance: highly noisy students require fewer samples to recover the underlying pattern, while low-noise students achieve a larger overlap with the ground-truth signal. We derive the optimal Bayesian conditions for teacher recovery as functions of the sample complexity, noise level, and interaction strength, and validate these predictions through numerical simulations. The resulting dynamics can be mapped onto equilibrium sampling in a Restricted Boltzmann Machine with a structured hidden layer, providing a principled theoretical understanding of how interactions improve distributed generative modeling.

cond-mat.dis-nn

The effect of priors on Learning with Restricted Boltzmann Machines

Restricted Boltzmann Machines (RBMs) are generative models designed to learn from data with a rich underlying structure. In this work, we explore a teacher-student setting where a student RBM learns from examples generated by a teacher RBM, with a focus on the effect of the unit priors on learning efficiency. We consider a parametric class of priors that interpolate between continuous (Gaussian) and binary variables. This approach models various possible choices of visible units, hidden units, and weights for both the teacher and student RBMs. By analyzing the phase diagram of the posterior distribution in both the Bayes optimal and mismatched regimes, we demonstrate the existence of a triple point that defines the critical dataset size necessary for learning through generalization. The critical size is strongly influenced by the properties of the teacher, and thus the data, but is unaffected by the properties of the student RBM. Nevertheless, a prudent choice of student priors can facilitate training by expanding the so-called signal retrieval region, where the machine generalizes effectively.

cond-mat.dis-nn

Saddle Hierarchy in Dense Associative Memory

Dense Associative Memory (DAM) models have been attracting renewed attention since they were shown to be robust to adversarial examples and closely related to cutting edge machine learning paradigms, such as the attention mechanism and generative diffusion. We study a DAM built upon a three-layer Boltzmann machine with Potts hidden units, which represent data clusters and classes. Through a statistical mechanics analysis, we derive saddle-point equations that characterize both the stationary points of DAMs trained on real data and the fixed points of DAMs trained on synthetic data within a teacher-student framework. Based on these results, we propose a novel regularization scheme that makes training significantly more stable. Moreover, we show empirically that our DAM learns interpretable solutions to both supervised and unsupervised classification problems. Pushing our theoretical analysis further, we find that the weights learned by relatively small DAMs correspond to unstable saddle points in larger DAMs. We implement a network-growing algorithm that leverages this saddle-point hierarchy to drastically reduce the computational cost of training dense associative memory.

cs.LG

Network Security under Heterogeneous Cyber-Risk Profiles and Contagion

Cyber risk has become a critical financial threat in today's interconnected digital economy. This paper introduces a cyber-risk management framework for networked digital systems that combines the strategic behavior of players with contagion dynamics within a security game. We address the problem of optimally allocating cybersecurity resources across a network, focusing on the heterogeneous valuations of nodes by attackers and defenders, some areas may be of high interest to the attacker, while others are prioritized by the defender. We explore how this asymmetry drives attack and defense strategies and shapes the system's overall resilience. We extend a method to determine optimal resource allocation based on simple network metrics weighted by the defender's and attacker's risk profiles. We further propose risk measures based on contagion paths and analyze how propagation dynamics influence optimal defense strategies. Numerical experiments explore risk versus cost efficient frontiers varying network topologies and risk profiles, revealing patterns of resource allocation and cyber deception effects. These findings provide actionable insights for designing resilient digital infrastructures and mitigating systemic cyber risk.

cs.CR

On the phase diagram of the multiscale mean-field spin-glass

In this paper we study the phase diagram of a Sherrington-Kirkpatrick (SK) model where the couplings are forced to thermalize at different time scales. Besides being a challenging generalization of the SK model, such settings may arise naturally in physics whenever part of the many degrees of freedom of a system relaxes to equilibrium considerably faster than the others. For this model we compute the asymptotic value of the second moment of the overlap distribution. Furthermore, we provide a rigorous sufficient condition for an annealed solution to hold, identifying a high temperature, or weak coupling, region. In addition, we also prove that for sufficiently strong couplings the solution must present a number of replica symmetry breaking levels at least equal to the number of time scales already present in the multiscale model. Finally, we give a sufficient condition for the existence of gaps in the support of the functional order parameters.

math-ph

On the supra-linear storage in dense networks of grid and place cells

Place-cell networks, typically forced to pairwise synaptic interactions, are widely studied as models of cognitive maps: such models, however, share a severely limited storage capacity, scaling linearly with network size and with a very small critical storage. This limitation is a challenge for navigation in 3-dimensional space because, oversimplifying, if encoding motion along a one-dimensional trajectory embedded in 2-dimensions requires $O(K)$ patterns (interpreted as bins), extending this to a 2-dimensional manifold embedded in a 3-dimensional space -- yet preserving the same resolution -- requires roughly $O(K^2)$ patterns, namely a supra-linear amount of patterns. In these regards, dense Hebbian architectures, where higher-order neural assemblies mediate memory retrieval, display much larger capacities and are increasingly recognized as biologically plausible, but have never linked to place cells so far. Here we propose a minimal two-layer model, with place cells building a layer and leaving the other layer populated by neural units that account for the internal representations (so to qualitatively resemble grid cells in the MEC of mammals): crucially, by assuming that each place cell interacts with pairs of grid cells, we show how such a model is formally equivalent to a dense Battaglia-Treves-like Hebbian network of grid cells only endowed with four-body interactions. By studying its emergent computational properties by means of statistical mechanics of disordered systems, we prove -- analytically -- that such effective higher-order assemblies (constructed under the guise of biological plausibility) can support supra-linear storage of continuous attractors; furthermore, we prove -- numerically -- that the present neural network is capable of recognition and navigation on general surfaces embedded in a 3-dimensional space.

cond-mat.dis-nn

Modeling Structured Data Learning with Restricted Boltzmann Machines in the Teacher-Student Setting

Restricted Boltzmann machines (RBM) are generative models capable to learn data with a rich underlying structure. We study the teacher-student setting where a student RBM learns structured data generated by a teacher RBM. The amount of structure in the data is controlled by adjusting the number of hidden units of the teacher and the correlations in the rows of the weights, a.k.a. patterns. In the absence of correlations, we validate the conjecture that the performance is independent of the number of teacher patters and hidden units of the student RBMs, and we argue that the teacher-student setting can be used as a toy model for studying the lottery ticket hypothesis. Beyond this regime, we find that the critical amount of data required to learn the teacher patterns decreases with both their number and correlations. In both regimes, we find that, even with a relatively large dataset, it becomes impossible to learn the teacher patterns if the inference temperature used for regularization is kept too low. In our framework, the student can learn teacher patterns one-to-one or many-to-one, generalizing previous findings about the teacher-student setting with two hidden units to any arbitrary finite number of hidden units.

cs.LG

Dense Hopfield Networks in the Teacher-Student Setting

Dense Hopfield networks are known for their feature to prototype transition and adversarial robustness. However, previous theoretical studies have been mostly concerned with their storage capacity. We bridge this gap by studying the phase diagram of p-body Hopfield networks in the teacher-student setting of an unsupervised learning problem, uncovering ferromagnetic phases reminiscent of the prototype and feature learning regimes. On the Nishimori line, we find the critical size of the training set necessary for efficient pattern retrieval. Interestingly, we find that that the paramagnetic to ferromagnetic transition of the teacher-student setting coincides with the paramagnetic to spin-glass transition of the direct model, i.e. with random patterns. Outside of the Nishimori line, we investigate the learning performance in relation to the inference temperature and dataset noise. Moreover, we show that using a larger p for the student than the teacher gives the student an extensive tolerance to noise. We then derive a closed-form expression measuring the adversarial robustness of such a student at zero temperature, corroborating the positive correlation between number of parameters and robustness observed in large neural networks. We also use our model to clarify why the prototype phase of modern Hopfield networks is adversarially robust.

cond-mat.dis-nn

Hopfield model with planted patterns: a teacher-student self-supervised learning model

While Hopfield networks are known as paradigmatic models for memory storage and retrieval, modern artificial intelligence systems mainly stand on the machine learning paradigm. We show that it is possible to formulate a teacher-student self-supervised learning problem with Boltzmann machines in terms of a suitable generalization of the Hopfield model with structured patterns, where the spin variables are the machine weights and patterns correspond to the training set's examples. We analyze the learning performance by studying the phase diagram in terms of the training set size, the dataset noise and the inference temperature (i.e. the weight regularization). With a small but informative dataset the machine can learn by memorization. With a noisy dataset, an extensive number of examples above a critical threshold is needed. In this regime the memory storage limits of the system becomes an opportunity for the occurrence of a learning regime in which the system can generalize.

cond-mat.dis-nn

Contingent Convertible Bonds in Financial Networks

We study the role of contingent convertible bonds (CoCos) in a complex network of interconnected banks. By studying the system's phase transitions, we reveal that the structure of the interbank network is of fundamental importance for the effectiveness of CoCos as a financial stability enhancing mechanism. Our results show that, under some network structures, the presence of CoCos can increase (and not reduce) financial fragility, because of the occurring of unneeded triggers and consequential suboptimal conversions that damage CoCos investors. We also demonstrate that, in the presence of a moderate financial shock, lightly interconnected financial networks are more robust than highly interconnected networks. This makes them a potentially optimal choice for both CoCos issuers and buyers.

q-fin.GN

Deep Reinforcement Trading with Predictable Returns

Classical portfolio optimization often requires forecasting asset returns and their corresponding variances in spite of the low signal-to-noise ratio provided in the financial markets. Modern deep reinforcement learning (DRL) offers a framework for optimizing sequential trader decisions but lacks theoretical guarantees of convergence. On the other hand, the performances on real financial trading problems are strongly affected by the goodness of the signal used to predict returns. To disentangle the effects coming from return unpredictability from those coming from algorithm un-trainability, we investigate the performance of model-free DRL traders in a market environment with different known mean-reverting factors driving the dynamics. When the framework admits an exact dynamic programming solution, we can assess the limits and capabilities of different value-based algorithms to retrieve meaningful trading signals in a data-driven manner. We consider DRL agents that leverage classical strategies to increase their performances and we show that this approach guarantees flexibility, outperforming the benchmark strategy when the price dynamics is misspecified and some original assumptions on the market environment are violated with the presence of extreme events and volatility clustering.

q-fin.PM

Reinforcement Learning Policy Recommendation for Interbank Network Stability

In this paper, we analyze the effect of a policy recommendation on the performance of an artificial interbank market. Financial institutions stipulate lending agreements following a public recommendation and their individual information. The former is modeled by a reinforcement learning optimal policy that maximizes the system's fitness and gathers information on the economic environment. The policy recommendation directs economic actors to create credit relationships through the optimal choice between a low interest rate or a high liquidity supply. The latter, based on the agents' balance sheet, allows determining the liquidity supply and interest rate that the banks optimally offer their clients within the market. Thanks to the combination between the public and the private signal, financial institutions create or cut their credit connections over time via a preferential attachment evolving procedure able to generate a dynamic network. Our results show that the emergence of a core-periphery interbank network, combined with a certain level of homogeneity in the size of lenders and borrowers, is essential to ensure the system's resilience. Moreover, the optimal policy recommendation obtained through reinforcement learning is crucial in mitigating systemic risk.

econ.GN

Modelling time-varying interactions in complex systems: the Score Driven Kinetic Ising Model

A common issue when analyzing real-world complex systems is that the interactions between the elements often change over time: this makes it difficult to find optimal models that describe this evolution and that can be estimated from data, particularly when the driving mechanisms are not known. Here we offer a new perspective on the development of models for time-varying interactions introducing a generalization of the well-known Kinetic Ising Model (KIM), a minimalistic pairwise constant interactions model which has found applications in multiple scientific disciplines. Keeping arbitrary choices of dynamics to a minimum and seeking information theoretical optimality, the Score-Driven methodology lets us significantly increase the knowledge that can be extracted from data using the simple KIM. In particular, we first identify a parameter whose value at a given time can be directly associated with the local predictability of the dynamics. Then we introduce a method to dynamically learn the value of such parameter from the data, without the need of specifying parametrically its dynamics. Finally, we extend our framework to disentangle different sources (e.g. endogenous vs exogenous) of predictability in real time. We apply our methodology to several complex systems including financial markets, temporal (social) networks, and neuronal populations. Our results show that the Score-Driven KIM produces insightful descriptions of the systems, allowing to predict forecasting accuracy in real time as well as to separate different components of the dynamics. This provides a significant methodological improvement for data analysis in a wide range of disciplines.

q-fin.ST

On the equivalence between the Kinetic Ising Model and discrete autoregressive processes

Binary random variables are the building blocks used to describe a large variety of systems, from magnetic spins to financial time series and neuron activity. In Statistical Physics the Kinetic Ising Model has been introduced to describe the dynamics of the magnetic moments of a spin lattice, while in time series analysis discrete autoregressive processes have been designed to capture the multivariate dependence structure across binary time series. In this article we provide a rigorous proof of the equivalence between the two models in the range of a unique and invertible map unambiguously linking one model parameters set to the other. Our result finds further justification acknowledging that both models provide maximum entropy distributions of binary time series with given means, auto-correlations, and lagged cross-correlations of order one. We further show that the equivalence between the two models permits to exploit the inference methods originally developed for one model in the inference of the other.

cond-mat.stat-mech

Unveiling the relation between herding and liquidity with trader lead-lag networks

We propose a method to infer lead-lag networks of traders from the observation of their trade record as well as to reconstruct their state of supply and demand when they do not trade. The method relies on the Kinetic Ising model to describe how information propagates among traders, assigning a positive or negative "opinion" to all agents about whether the traded asset price will go up or down. This opinion is reflected by their trading behavior, but whenever the trader is not active in a given time window, a missing value will arise. Using a recently developed inference algorithm, we are able to reconstruct a lead-lag network and to estimate the unobserved opinions, giving a clearer picture about the state of supply and demand in the market at all times. We apply our method to a dataset of clients of a major dealer in the Foreign Exchange market at the 5 minutes time scale. We identify leading players in the market and define a herding measure based on the observed and inferred opinions. We show the causal link between herding and liquidity in the inter-dealer market used by dealers to rebalance their inventories.

q-fin.TR

Legendre Equivalences of Spherical Boltzmann Machines

We study either fully visible and restricted Boltzmann machines with sub-Gaussian random weights and spherical or Gaussian priors. We prove the free energies of the spherical and Gaussian models are related by a Legendre transformation. Incidentally our analysis brings also a new purely variational derivation of the free energy of the spherical models.

cond-mat.dis-nn

Inverse problems for structured datasets using parallel TAP equations and RBM

We propose an efficient algorithm to solve inverse problems in the presence of binary clustered datasets. We consider the paradigmatic Hopfield model in a teacher student scenario, where this situation is found in the retrieval phase. This problem has been widely analyzed through various methods such as mean-field approaches or the pseudo-likelihood optimization. Our approach is based on the estimation of the posterior using the Thouless-Anderson-Palmer (TAP) equations in a parallel updating scheme. At the difference with other methods, it allows to retrieve the exact patterns of the teacher and the parallel update makes it possible to apply it for large system sizes. We also observe that the Approximate Message Passing (AMP) equations do not reproduce the expected behavior in the direct problem, questioning the standard practice used to obtain time indexes coming from Belief Propagation (BP). We tackle the same problem using a Restricted Boltzmann Machine (RBM) and discuss the analogies between the two algorithms.

cond-mat.dis-nn

Inference of the Kinetic Ising Model with Heterogeneous Missing Data

We consider the problem of inferring a causality structure from multiple binary time series by using the Kinetic Ising Model in datasets where a fraction of observations is missing. We take our steps from a recent work on Mean Field methods for the inference of the model with hidden spins and develop a pseudo-Expectation-Maximization algorithm that is able to work even in conditions of severe data sparsity. The methodology relies on the Martin-Siggia-Rose path integral method with second order saddle-point solution to make it possible to calculate the log-likelihood in polynomial time, giving as output a maximum likelihood estimate of the couplings matrix and of the missing observations. We also propose a recursive version of the algorithm, where at every iteration some missing values are substituted by their maximum likelihood estimate, showing that the method can be used together with sparsification schemes like LASSO regularization or decimation. We test the performance of the algorithm on synthetic data and find interesting properties when it comes to the dependency on heterogeneity of the observation frequency of spins and when some of the hypotheses that are necessary to the saddle-point approximation are violated, such as the small couplings limit and the assumption of statistical independence between couplings.

physics.data-an