SearcharxivSearch

arXiv subjects

Benoît Collins

Publications and source records attributed to Benoît Collins.

At least 19 recordsLinked to original sources

Free Random Projection for In-Context Reinforcement Learning

Hierarchical inductive biases are hypothesized to promote generalizable policies in reinforcement learning, as demonstrated by explicit hyperbolic latent representations and architectures. Therefore, a more flexible approach is to have these biases emerge naturally from the algorithm. We introduce Free Random Projection, an input mapping grounded in free probability theory that constructs random orthogonal matrices where hierarchical structure arises inherently. The free random projection integrates seamlessly into existing in-context reinforcement learning frameworks by encoding hierarchical organization within the input space without requiring explicit architectural modifications. Empirical results on multi-environment benchmarks show that free random projection consistently outperforms the standard random projection, leading to improvements in generalization. Furthermore, analyses within linearly solvable Markov decision processes and investigations of the spectrum of kernel random matrices reveal the theoretical underpinnings of free random projection's enhanced performance, highlighting its capacity for effective adaptation in hierarchically structured state spaces.

cs.LG

Gaussian Equivalence for Self-Attention: Asymptotic Spectral Analysis of Attention Matrix

Self-attention layers have become fundamental building blocks of modern deep neural networks, yet their theoretical understanding remains limited, particularly from the perspective of random matrix theory. In this work, we provide a rigorous analysis of the singular value spectrum of the attention matrix and establish the first Gaussian equivalence result for attention. In a natural regime where the inverse temperature remains of constant order, we show that the singular value distribution of the attention matrix is asymptotically characterized by a tractable linear model. We further demonstrate that the distribution of squared singular values deviates from the Marchenko-Pastur law, which has been believed in previous work. Our proof relies on two key ingredients: precise control of fluctuations in the normalization term and a refined linearization that leverages favorable Taylor expansions of the exponential. This analysis also identifies a threshold for linearization and elucidates why attention, despite not being an entrywise operation, admits a rigorous Gaussian equivalence in this regime.

stat.ML

Operator Norm Bounds for Multi-leg Matrix Tensors and Applications to Random Matrix Theory

We investigate the extremal values of partial traces of matrix tensors under operator norm constraints. To evaluate these multi-linear quantities, we develop a comprehensive graphical formalism that encodes multi-leg partial traces, partial permutations, and their moments using colored directed graphs. With this graphical framework, we establish optimal, sharp bounds for the partial trace $(\mathrm{Tr}_{σ_1} \otimes \ldots \otimes \mathrm{Tr}_{σ_k})(A_1, \ldots, A_m)$ over matrices bounded by $\|A_i\| \le 1$. Specifically, we prove that this maximum evaluates exactly to $N^{M(σ_1,\ldots,σ_k)}$, where $N$ is the dimension and $M$ represents the maximal number of directed cycles in the associated graph across all possible internal vertex pairings. We further derive explicit operator norm estimates for matrices generated by partial traces of partial permutations. Finally, we apply these combinatorial bounds to multi-matrix random matrix theory. By examining models involving Ginibre ensembles, we extend concepts of asymptotic freeness to matrix coefficient algebras, establishing operator norm estimates that rigorously separate the asymptotic behavior of non-crossing and crossing pairings.

math.OA

The complexity of semidefinite programs for testing $k$-block-positivity

We extend \cite{chen2025srkbp} by analyzing the complexity of the $k$-block-positivity testing algorithm that stems from the optimization problem in Definition \ref{definition:SDP-k-block-positivity}. In this paper, we investigate a symmetry reduction scheme based on rectangular shaped Young diagrams. Connecting the complexity to the dimensions of irreducible representations of $\U(d)$, we derive an explicit formula for the complexity, which also clarifies why the semidefinite program hierarchy collapses in the $k=d$ case.

quant-ph

Smooth Multi-Trace Statistics of Classical Ensembles: Large $N$ Expansions, Cumulants, and Matrix Integrals

We consider expectations of the form $E [tr h_1(X_1^N)... tr h_r(X_r^N)]$, where $X_i^N$ are self-adjoint polynomials in various independent classical random matrices and $h_i$ are smooth test function and obtain a large $N$ expansion of these quantities, building on the framework of polynomial approximation and Bernstein-type inequalities recently developed by Chen, Garza-Vargas, Tropp, and van Handel. As applications of the above, we prove the higher-order asymptotic vanishing of cumulants for smooth linear statistics, establish a Central Limit Theorem, and demonstrate the existence of formal asymptotic expansions for the free energy and observables of matrix integrals with smooth potentials.

math.PR

Weingarten Calculus with Virtual Isometries

In this paper, we develop a novel approach to the Weingarten calculus by employing the notion of virtual isometries. Traditionally, Weingarten calculus provides explicit formulas for integrating polynomial functions over compact matrix groups with respect to the Haar measure, yet it faces limitations when evaluating high-degree integrals due to the non-invertibility of the associated matrices. We revisit these classical computations from a new perspective: by constructing Haar-distributed matrices as products of sequences of complex reflections, we derive new recursive structures for the Weingarten functions across different dimensions. This framework leads to two main results: (1) an explicit Weingarten calculus for complex reflections, yielding systematic moment computations for associated rank-one matrices, and (2) a novel convolution formula that connects Weingarten functions in dimension $n$ to those in dimension $n-1$, through the introduction of ascension functions in the symmetric group algebra. Our approach not only provides a unified treatment for unitary groups, but also sheds light on the algebraic and probabilistic aspects of high-degree integral computations. We present several examples and applications.

math.PR

Operator-valued Khintchine inequality for $ε$-free semicircles

We exhibit several bounds for operator norms of the sum of $ε$-free semicircular random variables introduced in the paper of Speicher and Wysoczański. In particular, using the first and second largest eigenvalues of the adjacency matrix $ε$, we show analogs of the operator-valued Khintchine-type inequality obtained by Haagerup and Pisier.

math.FA

Holdout cross-validation for large non-Gaussian covariance matrix estimation using Weingarten calculus

Cross-validation is one of the most widely used methods for model selection and evaluation; its efficiency for large covariance matrix estimation appears robust in practice, but little is known about the theoretical behavior of its error. In this paper, we derive the expected Frobenius error of the holdout method, a particular cross-validation procedure that involves a single train and test split, for a generic rotationally invariant multiplicative noise model, therefore extending previous results to non-Gaussian data distributions. Our approach involves using the Weingarten calculus and the Ledoit-Péché formula to derive the oracle eigenvalues in the high-dimensional limit. When the population covariance matrix follows an inverse Wishart distribution, we approximate the expected holdout error, first with a linear shrinkage, then with a quadratic shrinkage to approximate the oracle eigenvalues. Under the linear approximation, we find that the optimal train-test split ratio is proportional to the square root of the matrix dimension. Then we compute Monte Carlo simulations of the holdout error for different distributions of the norm of the noise, such as the Gaussian, Student, and Laplace distributions and observe that the quadratic approximation yields a substantial improvement, especially around the optimal train-test split ratio. We also observe that a higher fourth-order moment of the Euclidean norm of the noise vector sharpens the holdout error curve near the optimal split and lowers the ideal train-test ratio, making the choice of the train-test ratio more important when performing the holdout method.

q-fin.ST

Variations on quantum de Finetti theorems and operator valued Martin boundaries: a Choquet-Deny approach

We revisit the quantum de Finetti theorem. We state and prove a couple of variants thereof. In parallel, we introduce an operator version of the Martin boundary on quantum groups and prove generalizations of Biane's theoresm. Our proof of the de Finetti theorem is new in the sense that it is based on an analogy with the theory of operator valued Martin boundary that we introduce.

math.OA

Symmetry reduction for testing $k$-block-positivity via extendibility

We study the problem of testing $k$-block-positivity via symmetric $N$-extendibility by taking the tensor product with a $k$-dimensional maximally entangled state. We exploit the unitary symmetry of the maximally entangled state to reduce the size of the corresponding semidefinite programs (SDP). For example, for $k=2$, the SDP is reduced from one block of size $2^{N+1} d^{N+1}$ to $\lfloor \frac{N+1}{2} \rfloor$ blocks of size $\approx O( (N-1)^{-1} 2^{N+1} d^{N+1} )$.

quant-ph

Weingarten calculus for centered random permutation matrices

We introduce and study the Weingarten calculus for centered random permutation matrices in the symmetric group S_N. After presenting a formulation of the Weingarten calculus on the symmetric group, we derive a formula in the centered case, as well as a sign-respecting formula. Our investigations uncover the fact that a building block of this Weingarten calculus is Kummer's confluent hypergeometric function. It allows us to derive multiple algebraic properties of the Weingarten function and uniform estimate. These results shed a conceptual light on phenomena that take place regarding the algebraic and asymptotic behavior of moments of random permutations in the resolution of Bordenave and Bordenave-Collins of strong convergence. We obtain multiple new non-trivial estimates for moments of coefficients in centered moments.

math.PR

Fluctuations of eigenvalues of a polynomial on Haar unitary and finite rank matrices

This paper calculates the fluctuations of eigenvalues of polynomials on large Haar unitaries cut by finite rank deterministic matrices. When the eigenvalues are all simple, we can give a complete algorithm for computing the fluctuations. When multiple eigenvalues are involved, we present several examples suggesting that a general algorithm would be much more complex.

math.PR

Strong convergence for tensor GUE random matrices

Haagerup and Thorbjørnsen proved that iid GUEs converge strongly to free semicircular elements as the dimension grows to infinity. Motivated by considerations from quantum physics -- in particular, understanding nearest neighbor interactions in quantum spin systems -- we consider iid GUE acting on multipartite state spaces, with a mixing component on some sites and identity on the remaining sites. We show that under proper assumptions on the dimension of the sites, strong asymptotic freeness still holds. Our proof relies on an interpolation technology recently introduced by Bandeira, Boedihardjo and van Handel.

math.PR

Eigenvalues of random lifts and polynomials of random permutation matrices

Consider a finite sequence of independent random permutations, chosen uniformly either among all permutations or among all matchings on n points. We show that, in probability, as n goes to infinity, these permutations viewed as operators on the (n-1) dimensional vector space orthogonal to the vector with all coordinates equal to 1, are asymptotically strongly free. Our proof relies on the development of a matrix version of the non-backtracking operator theory and a refined trace method. As a byproduct, we show that the non-trivial eigenvalues of random n-lifts of a fixed based graphs approximately achieve the Alon-Boppana bound with high probability in the large n limit. This result generalizes Friedman's Theorem stating that with high probability, the Schreier graph generated by a finite number of independent random permutations is close to Ramanujan. Finally, we extend our results to tensor products of random permutation matrices. This extension is especially relevant in the context of quantum expanders.

math.PR

Projections of Orbital Measures and Quantum Marginal Problems

This paper studies projections of uniform random elements of (co)adjoint orbits of compact Lie groups. Such projections generalize several widely studied ensembles in random matrix theory, including the randomized Horn's problem, the randomized Schur's problem, and the orbital corners process. In this general setting, we prove integral formulae for the probability densities, establish some properties of the densities, and discuss connections to multiplicity problems in representation theory as well as to known results in the symplectic geometry literature. As applications, we show a number of results on marginal problems in quantum information theory and also prove an integral formula for restriction multiplicities.

math-ph

Convergence for noncommutative rational functions evaluated in random matrices

One of the main applications of free probability is to show that for appropriately chosen independent copies of $d$ random matrix models, any noncommutative polynomial in these $d$ variables has a spectral distribution that converges asymptotically and can be described with the help of free probability. This paper aims to show that this can be extended to noncommutative rational functions, answering an open question by Roland Speicher. This paper also provides a noncommutative probability approach to approximating the free field. At the algebraic level, its construction relies on the approximation by generic matrices. On the other hand, it admits many embeddings in the algebra of operators affiliated with a $II_1$ factor. A consequence of our result is that, as soon as the generators admit a random matrix model, the approximation of any self-adjoint noncommutative rational function by generic matrices can be upgraded at the level of convergence in distribution.

math.OA

Asymptotic Freeness of Unitary Matrices in Tensor Product Spaces for Invariant States

In this paper, we pursue our study of asymptotic properties of families of random matrices that have a tensor structure. In previous work, the first- and second-named authors provided conditions under which tensor products of unitary random matrices are asymptotically free with respect to the normalized trace. Here, we extend this result by proving that asymptotic freeness of tensor products of Haar unitary matrices holds with respect to a significantly larger class of states. Our result relies on invariance under the symmetric group, and therefore on traffic probability. As a byproduct, we explore two additional generalisations: (i) we state results of freeness in a context of general sequences of representations of the unitary group -- the fundamental representation being a particular case that corresponds to the classical asymptotic freeness result for Haar unitary matrices, and (ii) we consider actions of the symmetric group and the free group simultaneously and obtain a result of asymptotic freeness in this context as well.

math.PR

Matrix models for cyclic monotone and monotone independences

Cyclic monotone independence is an algebraic notion of noncommutative independence, introduced in the study of multi-matrix random matrix models with small rank. Its algebraic form turns out to be surprisingly close to monotone independence, which is why it was named cyclic monotone independence. This paper conceptualizes this notion by showing that the same random matrix model is also a model for the monotone convergence with an appropriately chosen state. This observation provides a unified nonrandom matrix model for both types of monotone independences.

math.OA