SearcharxivSearch

arXiv subjects

Giuseppe Bruno

Publications and source records attributed to Giuseppe Bruno.

10 recordsLinked to original sources

Quantitative Diffusive Limits for Singular Nonlocal Transport

We study the nonlocal continuity equation \[ \partial_t\mu_b =\operatorname{div}\!\left( \mu_b\nabla\log\bigl((I-b^2\Delta)^{-1}\mu_b\bigr) \right) \] on a closed connected Riemannian manifold. For smooth strictly positive initial data, we prove that as $b \to 0$, its global solution converges to heat flow $\mu(t)$ at the sharp, uniform-in-time rate \[ \sup_{t\ge0}\|\mu_b(t)-\mu(t)\|_{L^1}\le Cb^2. \] The key estimate is the uniform dissipation of a $b$-weighted higher-order resolvent energy, which yields exponential relaxation despite the absence of a Wasserstein gradient-flow structure. On the circle, we also analyze the corresponding deterministic $N$-particle dynamics. A weak--strong modulated energy argument gives \[ \mathbb E\!\left[ \sup_{t\ge0}W_1(\mu_b^N(t),\mu_b(t)) \right] \le C(Nb)^{-1/2} \] for iid initialization. Consequently, the choice $b\asymp N^{-1/5}$ approximates heat flow uniformly in time at rate $N^{-2/5}$.

math.AP

Morrey's problem in $\mathbb{R}^{2 \times 4}$ and $\mathbb{R}^{3 \times 3}_\mathrm{sym}$

We find an explicit rank-one convex non-quasiconvex integrand in $\mathbb{R}^{2\times 4}$: to falsify the quasiconvexity inequality, we exhibit a map $\mathbb{T}^4\to \mathbb{R}^2$ with $12$ non-zero Fourier modes. In fact, this map is obtained from a scalar potential, so we also find a rank one convex integrand in $\mathbb{R}^{4\times 4}_\text{sym}$ which is not quasiconvex. These examples are obtained by transpositions and restrictions of Grabovsky's example of a rank-one convex, non-quasiconvex integrand in $\mathbb{R}^{8 \times 2}.$ We also modify \v{S}ver\'{a}k's example to construct a rank-one convex non-quasiconvex integrand in $\mathbb{R}^{3\times 3}_\text{sym}$.

math.AP

Quasiconvexity for the Dacorogna--Marcellini Energy

We prove that the planar Dacorogna--Marcellini energy $f_\gamma(A)=|A|^4-2\gamma|A|^2\det A$ is quasiconvex exactly when it is rank-one convex, i.e. if and only if $|\gamma|\leq\frac{2}{\sqrt3}$. The proof uses a monotonicity property of the energy functional along the componentwise heat flow. As a corollary of our method, we show that for homogeneous quartic polynomials on $2 \times 2$ matrices invariant by left and right rotation, quasiconvexity is equivalent to rank one convexity.

math.AP

Digital-Based Potentiostat and Mesoporous Microelectrode Co-Design for Non-Enzymatic Glucose Detection at 0.3V-VDD and 1.65nW-Power

This paper presents a proof-of-concept ultra-low voltage and ultra-low power chronoamperometric electrochemical sensor for non-enzymatic glucose readout integrated circuit (IC) in 130nm CMOS detection featuring a reconfigurable Digital-Based (DB) Potentiostat. The signal transfer and noise characteristics of the new digital-based architecture are analytically described in the frequency domain for the first time by an equivalent linearized model that is validated by simulations and experiments. Based on experiments, the proposed DB potentiostat enables the detection of a wide electrochemical current range, spanning from 600pA to 650nA, with R2=0.991 linearity and consumes only 1.65nW (53.5nW) at V dd = 300mV (V dd = 500mV). The proposed DB readout is tested in a proof of-concept platform for non-enzymatic glucose detection with nanostructured microelectrodes, demonstrating successful non enzymatic glucose detection at physiological levels at the lowest reported voltage and power, even in the presence of an interferent (ascorbic acid) and under aerobic conditions, thus revealing a strong potential for emerging Point of Care (PoC) diagnostics applications.

eess.SY

Scaling Limits of Long-Context Transformers

We study the long-context limit of softmax self-attention with a fixed query and a random context of $n$ i.i.d. keys on the sphere, viewing the inverse temperature $\beta_n$ as the scaling parameter that decides whether attention degenerates into uniform averaging or collapses onto the single closest key. We show that the critical scale at which selectivity emerges is determined by the local exponent of the distance-to-query distribution near zero rather than by global features of the context, and scales like $\beta_n^\ast \asymp n^{2/(d-1)}$ for uniform keys on $\mathbb{S}^{d-1}$. Furthermore, we characterize the limiting laws of the ordered attention weights and of the attention output across all regimes of $\beta_n$: a subcritical regime in which the output reduces to a local average around $q$ with explicit deterministic bias and Gaussian fluctuations; a critical regime in which a finite collection of nearest keys retains macroscopic mass without single-key collapse; and a supercritical regime in which all mass concentrates on the closest key. Of notable interest is the subcritical case with identity value matrix where the attention map approximately implements a backward heat equation.

cs.LG

Stochastic Scaling Limits and Synchronization by Noise in Deep Transformer Models

We prove pathwise convergence of the layerwise evolution of tokens in a finite-depth, finite-width transformer model with MultiLayer Perceptron (MLP) blocks to a continuous-time stochastic interacting particle system. We also identify the stochastic partial differential equation describing the evolution of the tokens' distribution in this limit and prove propagation of chaos when the number of such tokens is large. The bounds we establish are quantitative and the limits we consider commute. We further prove that the limiting stochastic model displays synchronization by noise and establish exponential dissipation of the interaction energy on average, provided that the common noise is sufficiently coercive relative to the deterministic self-attention drift. We finally characterize the activation functions satisfying the former condition.

math.PR

A multiscale analysis of mean-field transformers in the moderate interaction regime

In this paper, we study the evolution of tokens through the depth of encoder-only transformer models at inference time by modeling them as a system of particles interacting in a mean-field way and studying the corresponding dynamics. More specifically, we consider this problem in the moderate interaction regime, where the number $N$ of tokens is large and the inverse temperature parameter $\beta$ of the model scales together with $N$. In this regime, the dynamics of the system displays a multiscale behavior: a fast phase, where the token empirical measure collapses on a low-dimensional space, an intermediate phase, where the measure further collapses into clusters, and a slow one, where such clusters sequentially merge into a single one. We provide a rigorous characterization of the limiting dynamics in each of these phases and prove convergence in the above mentioned limit, exemplifying our results with some simulations.

cs.LG

Securities Transaction Settlement Optimization on superconducting quantum devices

We describe a quantum variational algorithm for securities transactions settlement optimization, based on a novel mathematical formalization of the problem that includes the most relevant constraints considered in the pan-European securities settlement platform TARGET2-Securities. The proposed algorithm is designed for Noisy Intermediate-Scale Quantum devices, specifically targeting IBM's superconducting qubit machines. We adopt non-linear activation functions to encode inequality constraints in the objective function of the problem, and design customized noise mitigation techniques to alleviate the effect of readout errors. We consider batches of up to 40 trades obtained from real transactional data to benchmark our algorithm on quantum hardware against classical and quantum-inspired solvers.

quant-ph

Emergence of meta-stable clustering in mean-field transformer models

We model the evolution of tokens within a deep stack of Transformer layers as a continuous-time flow on the unit sphere, governed by a mean-field interacting particle system, building on the framework introduced in (Geshkovski et al., 2023). Studying the corresponding mean-field Partial Differential Equation (PDE), which can be interpreted as a Wasserstein gradient flow, in this paper we provide a mathematical investigation of the long-term behavior of this system, with a particular focus on the emergence and persistence of meta-stable phases and clustering phenomena, key elements in applications like next-token prediction. More specifically, we perform a perturbative analysis of the mean-field PDE around the iid uniform initialization and prove that, in the limit of large number of tokens, the model remains close to a meta-stable manifold of solutions with a given structure (e.g., periodicity). Further, the structure characterizing the meta-stable manifold is explicitly identified, as a function of the inverse temperature parameter of the model, by the index maximizing a certain rescaling of Gegenbauer polynomials.

cs.LG

The Challenges of the Nonlinear Regime for Physics-Informed Neural Networks

The Neural Tangent Kernel (NTK) viewpoint is widely employed to analyze the training dynamics of overparameterized Physics-Informed Neural Networks (PINNs). However, unlike the case of linear Partial Differential Equations (PDEs), we show how the NTK perspective falls short in the nonlinear scenario. Specifically, we establish that the NTK yields a random matrix at initialization that is not constant during training, contrary to conventional belief. Another significant difference from the linear regime is that, even in the idealistic infinite-width limit, the Hessian does not vanish and hence it cannot be disregarded during training. This motivates the adoption of second-order optimization methods. We explore the convergence guarantees of such methods in both linear and nonlinear cases, addressing challenges such as spectral bias and slow convergence. Every theoretical result is supported by numerical examples with both linear and nonlinear PDEs, and we highlight the benefits of second-order methods in benchmark test cases.

cs.LG