Searcharxiv⌕ Search

arXiv subjects

Pierpaolo Vivo

Publications and source records attributed to Pierpaolo Vivo.

At least 19 recordsLinked to original sources

Rank-One Signal Recovery in Sparse Wishart Noise

We study the high-dimensional recovery of a signal vector $\mathbf{x}$ in the presence of sparse Wishart-like noise. We define an $N \times N$ matrix $A = J+(θ/N)\mathbf{xx}^{\top}$, where $\mathbf{xx}^{\top}$ is the rank-one deformation of the random noise matrix $J$. We consider a Wishart-like matrix $J={X}^{\top} X$, where $X$ is a sparse $M \times N$ random matrix with entries $X_{ij} = c_{ij}W_{ij}$, with $c_{ij}$ regulating the density of non-zero elements, and $W_{ij}$ the bond weights. Using the replica method, we compute analytically the top eigenpair statistics of $A$, and their dependence on the signal strength $θ$, the rectangularity ratio $α=\sqrt{M/N}$, and the average connectivity of the noise. The spectral observables are expressed in terms of a system of Recursive Distributional Equations, which are efficiently solved via a Population Dynamics algorithm. They allow us to compute the average largest eigenvalue $\langleλ_1\rangle_{A}$, the average top eigenvector component density, and the average overlap between the top eigenvector of $A$ and $\mathbf{x}$. We identify a critical threshold $θ_{\mathrm{crit}}$--depending on the average connectivity of the noise--that marks a BBP-like phase transition: below this value, $\langleλ_1\rangle_{A}$ is unaffected by the signal, and the overlap vanishes. Thus, the signal is not recoverable from the top eigenvector of $A$. For $θ>θ_{\mathrm{crit}}$, the signal-related outlier eigenvalue becomes $\langleλ_1\rangle_{A}$ and the overlap is nonzero, allowing for recovery of the signal. The results are in excellent agreement with numerical diagonalisation. We show that in the dense limit, the recovery threshold and eigen-statistics converge to the results predicted by the classical BBP transition for additive rank-one deformations of dense Wishart matrices.

cond-mat.dis-nn↗

Data as Commodity: a Game-Theoretic Principle for Information Pricing

Data is the central commodity of the digital economy. Unlike physical goods, data exhibits properties that defy the standard theory of supply and demand: it is non-rival (the same dataset can be sold to multiple buyers without degradation), it is replicable at near-zero cost, and it is traded under heterogeneous licensing rules that restrict lawful use. Determining a new pricing principle to attach a fair price tag to datasets is therefore a difficult but central problem. We propose a game-theoretic framework in which the value of a data string emerges from strategic competition among $N$ players betting on a stochastic process with asymmetric information about past outcomes. A better-informed player may either exploit her advantage or sell part of her dataset to less informed competitors. By analytically deriving the Nash equilibrium, we identify the price range for a mutually beneficial trade. The model reveals market dynamics that depart from textbook intuition: informed players may compete or jointly exploit the least informed; data can be shared even at zero price without reducing the seller`s utility; rivalry among well-informed players can benefit uninformed ones; and trades infeasible in small markets can be viable in larger ones. These findings establish a theoretical foundation for the pricing of intangible goods in interacting digital markets, which are in need of robust valuation principles.

physics.soc-ph↗

Maximal Minimal Spacing for Random Points

From $N+1$ random points on a line we wish to select $M+1$ points so as to maximize the minimal spacing between them. We consider an initial configuration with independent and identically distributed spacings. Equivalently, the points are arrival times of a generic renewal process. For general spacing distributions, and for all $M\leq N$, we derive exact distributional identities for the maximal minimal spacing and obtain its asymptotic behavior. The problem admits a reformulation in terms of a threshold-resetting random walk. The walk advances by successive random increments and is reset to the origin upon exceeding a fixed threshold. The probability that the optimal spacing exceeds a given value coincides with the probability that the walk completes at least $M$ reset cycles within $N$ steps. This yields an exact representation in terms of first-passage functionals of the walk. The same mapping suggests a numerical scheme for the max-min spacing problem in the regime of large $N$ and $M$, whose accuracy is tested against the exact results obtained here.

math-ph↗

Large deviations for linear regressions

Linear regression is one of the simplest and most widely used tools to learn patterns from data: it fits a set of coefficients so that a linear combination of predictors best matches observed responses. The quality of the fit is measured by the residual sum of squares, the total squared mismatch between predictions and data, whose minimum defines the training loss. We consider Gaussian design and noise, with teacher coefficients independently drawn from a general distribution $p(β)$, and a general class of separable regularizers, including Ridge and Lasso. Using the zero-temperature replica method, we compute analytically the large-deviation statistics of the minimum training loss for large numbers $P$ of predictors and $N$ of observations, with $r=P/N$ fixed. The rate function we compute governs rare sample-to-sample fluctuations of the optimal loss. Extensive numerical simulations are in excellent agreement with our theory and clearly show a pronounced deviation from the Gaussian regime of typical fluctuations in the tails.

cond-mat.stat-mech↗

Largest eigenvalue and top eigenvector statistics of large Euclidean random matrices

Euclidean random matrices arise in a wide range of physical systems where interactions are determined by spatial configurations, including disordered media and cooperative phenomena in atomic ensembles. Unlike classical random matrix ensembles, their entries are strongly correlated through the geometry of the underlying random points, making their analytical treatment challenging. While global spectral properties such as the spectral density are relatively well understood, much less is known about extremal eigenvalues and the associated eigenvectors, despite their central role in applications. Here we address the problem of characterising the largest eigenvalue and the corresponding top eigenvector of large Euclidean random matrices with a generic symmetric kernel. For vectors in any dimension $d\geq 1$ drawn independently from a common distribution, we show that both quantities can be computed within a unified replica-based framework. In the large-$N$ limit, the replica saddle-point equations reduce to an eigenvalue problem for a Fredholm integral operator determined by the kernel and by the underlying distribution. We then specialise this general formulation to the quadratic distance kernel, for which the Fredholm problem reduces to a finite set of $d+2$ self-consistent equations. In this case, we obtain an explicit expression for the average largest eigenvalue, fully determined by low-order moments of the underlying distribution, and an analytical characterisation of the density of the top eigenvector's components. We further perform extensive numerical simulations that confirm these predictions. More broadly, our work provides a general framework to access extremal spectral properties of Euclidean random matrices.

cond-mat.stat-mech↗

Sparse corruption in low-rank matrix inference: the PCA benchmark

Principal Component Analysis (PCA) is a standard tool for extracting a low-rank signal from noisy observations. It is known that applying PCA to a rank-one signal corrupted by a dense, homogeneous noise, in the large matrix size limit, the celebrated BBP transition occurs, where the emergence of an outlying eigenvalue and the alignment of the corresponding eigenvector occur at the same critical signal strength. Here we study the case of sparse noise corruption. The noise matrix is modelled as the adjacency matrix of a weighted undirected graph with finite average connectivity. Using the replica method, we analytically compute the typical top eigenvalue, the top eigenvector component density, and the squared overlap with the signal, through recursive distributional equations solved by population dynamics. We identify two signal-strength transitions as functions of graph connectivity: $θ_{\rm crit}$, marking signal recovery by the top eigenvector and generalising the BBP transition, and $θ_{\rm b}$, where the signal-related eigenvalue detaches from the bulk. For noise with nonzero mean, these transitions need not coincide because of a structural sparse-graph outlier, leading to a discontinuous transition in the squared overlap with the top eigenvector. The same top-eigenpair formalism also predicts the overlap of the signal with the eigenvector associated with the second largest eigenvalue when the signal eigenvalue is an outlier but remains below the structural outlier, where the transition is continuous. We specialise the equations to Poissonian and Random Regular degree distributions, recover dense-noise results in the large-connectivity limit, and validate the theory by numerical diagonalisation of large matrices.

stat.ML↗

From Judgments to Issues: Structured Extraction of Legal Reasoning with Citation-Hallucination Control

We present an automated pipeline that decomposes Italian tax-court judgments into individual legal issues and extracts, for each issue, a structured XML representation grounded in the IRAC framework and the legal syllogism. The pipeline targets a corpus of approximately $330{,}000$ first- and second-instance decisions of the Italian tax courts and is built around a capable yet cost-efficient general-purpose model (DeepSeek V3), a choice driven by the need to process several hundred thousand documents at a sustainable cost. To address the well-documented unreliability of large language models on legal citations, we couple the extraction step with an automatic hallucination-detection filter that compares the references produced by the model with those identified in the judgment text by a dedicated parser (Linkoln), normalised to standard identifiers (URN-NIR, ECLI, CELEX). We validate the pipeline on $50$ judgments annotated by two PhDs in tax law, computing inter-annotator agreement and LLM-vs-expert agreement on both issue extraction and legal citations, together with a stand-alone evaluation of the hallucination filter. To the best of our knowledge, this is the first issue-level, expert-validated structured extraction pipeline with hallucination control for Italian tax-court decisions, and it provides a concrete starting point for downstream applications such as issue-level retrieval, citation-network analysis, and the construction of large-scale datasets of legal reasoning.

cs.CL↗

Spectral criteria for generalization in unsupervised Hebbian nets

We consider an unsupervised Hebbian network where the pairwise interactions among neurons are built on noisy realizations of hidden ground-truth vectors. Unlike classical Hopfield models, designed as memory devices, this class of networks can be employed to extract latent structure and generalize beyond the "training" set. By combining random matrix theory and replica methods, we derive the asymptotic spectrum of the corresponding interaction matrix and show that the onset of generalization is controlled by a sharp spectral transition. Depending on the quality and the size of the accessible dataset, the spectrum displays either two separated bulks, encoding informative and noisy directions, or a merged single-bulk phase where such distinction is lost. We show that, when coupled with regularization, the emergence of such a spectral split predicts the network's capability to reconstruct the ground-truth vectors from corrupted samples.

cond-mat.dis-nn↗

Queue & AI: When Faster Tasks Slow Down the Workflow

Quantifying the workplace productivity effects of Generative Artificial Intelligence is now central to economics, management, and public policy. The deployment of AI tools in customer service, writing, software development, and consulting operations has been reported to generate large per-task productivity gains, typically measured as tasks completed per worker-hour or reductions in mean handle time. We argue that such mean-based metrics can misrepresent AI's effects in workflows where tasks accumulate and compete for scarce human attention. AI assistance can generate a deceptive productivity signature: average completion times fall because AI tools typically supply a fast first draft, yet workflow-level performance deteriorates when a subset of AI errors escapes review and returns as costly downstream rework. We call this divergence between mean task speed and system-level delay the variance wedge. Depending on the operational parameters, the most time-efficient way to complete a workflow may undergo a transition between two task-processing regimes, a fully AI-assisted and a fully manual one. We formalize the mechanism as a queueing model and derive two main implications analytically. First, under congestion, reviewers rationally raise the risk threshold for checking AI outputs, reducing scrutiny precisely when it would matter the most. Second, AI assistance can stabilize an overloaded workflow only when (i) the fraction of tasks handled by AI exceeds a critical threshold, and (ii) the human attention required for review and expected rework is lower than the attention for manual completion, a requirement substantially more stringent than faster draft generation. These results suggest that AI deployment should be evaluated not only by average task speed, but by its overall effects on congestion, rework, and the robustness of human oversight under load.

cs.CY↗

The Most Dispersed Subset of Random Points in $\mathbb{R}^d$

Consider a population of $N$ individuals, each having $d\geq 1$ different traits, and an additive measure, called dispersion, which rewards large pairwise separations between traits. The goal is to select $M\leq N$ individuals such that their traits are as dispersed as possible. We compute analytically the full statistics (including large deviation tails) of the maximally achievable dispersion among sub-populations of size $M$ when the traits are independent and identically distributed. Two complementary approaches are developed, one based on a mean-field theory for order statistics, and the other on the replica method from the field of disordered systems. In all dimensions $d$, and for rotationally symmetric distributions, the optimal subset for large populations consists of all points lying outside a $d$-dimensional ball whose radius is determined self-consistently. For a single trait ($d=1$), the statistics of the maximal dispersion can be tackled for finite $N,M$ as well. The formulae we obtained are corroborated by numerical simulations on small instances and by heuristic algorithms that find near-optimal solutions.

cond-mat.stat-mech↗

A calibrated model of debt recycling with interest costs and tax shields: viability under different fiscal regimes and jurisdictions

Debt recycling is a leveraged equity management strategy in which homeowners use accumulated home equity to finance investments, applying the resulting returns to accelerate mortgage repayment. We propose a novel framework to model equity and mortgage dynamics in presence of mortgage interest rates, borrowing costs on equity-backed credit lines, and tax shields arising from interest deductibility. The model is calibrated on three jurisdictions -- Australia, Germany, and Switzerland -- representing diverse interest rate environments and fiscal regimes. Results demonstrate that introducing positive interest rates without tax shields contracts success regions and lengthens repayment times, while tax shields partially reverse these effects by reducing effective borrowing costs and adding equity boosts from mortgage interest deductibility. Country-specific outcomes vary systematically, and rental properties consistently outperform owner-occupied housing due to mortgage interest deductibility provisions.

q-fin.RM↗

Mapping Microscopic and Systemic Risks in TradFi and DeFi: a literature review

This work explores the formation and propagation of systemic risks across traditional finance (TradFi) and decentralized finance (DeFi), offering a comparative framework that bridges these two increasingly interconnected ecosystems. We propose a conceptual model for systemic risk formation in TradFi, grounded in well-established mechanisms such as leverage cycles, liquidity crises, and interconnected institutional exposures. Extending this analysis to DeFi, we identify unique structural and technological characteristics - such as composability, smart contract vulnerabilities, and algorithm-driven mechanisms - that shape the emergence and transmission of risks within decentralized systems. Through a conceptual mapping, we highlight risks with similar foundations (e.g., trading vulnerabilities, liquidity shocks), while emphasizing how these risks manifest and propagate differently due to the contrasting architectures of TradFi and DeFi. Furthermore, we introduce the concept of crosstagion, a bidirectional process where instability in DeFi can spill over into TradFi, and vice versa. We illustrate how disruptions such as liquidity crises, regulatory actions, or political developments can cascade across these systems, leveraging their growing interdependence. By analyzing this mutual dynamics, we highlight the importance of understanding systemic risks not only within TradFi and DeFi individually, but also at their intersection. Our findings contribute to the evolving discourse on risk management in a hybrid financial ecosystem, offering insights for policymakers, regulators, and financial stakeholders navigating this complex landscape.

q-fin.RM↗

Top eigenpair statistics of diluted Wishart matrices

Using the replica method, we compute the statistics of the top eigenpair of diluted covariance matrices of the form $\mathbf{J} = \mathbf{X}^T \mathbf{X}$, where $\mathbf{X}$ is a $N\times M$ sparse data matrix, in the limit of large $N,M$ with fixed ratio and a bounded number of nonzero entries. We allow for random non-zero weights, provided they lead to an isolated largest eigenvalue. By formulating the problem as the optimisation of a quadratic Hamiltonian constrained to the $N$-sphere at low temperatures, we derive a set of recursive distributional equations for auxiliary probability density functions, which can be efficiently solved using a population dynamics algorithm. The average largest eigenvalue is identified with a Lagrange parameter that governs the convergence of the algorithm, and the resulting stable populations are then used to evaluate the density of the top eigenvector's components. We find excellent agreement between our analytical results and numerical results obtained from direct diagonalisation.

cond-mat.stat-mech↗

DebtStreamness: An Ecological Approach to Credit Flows in Inter-Firm Networks

Understanding how credit flows through inter-firm networks is critical for assessing financial stability and systemic risk. In this study, we introduce DebtStreamness, a novel metric inspired by trophic levels in ecological food webs, to quantify the position of firms within credit chains. By viewing credit as the ``primary energy source'' of the economy, we measure how far credit travels through inter-firm relationships before reaching its final borrowers. Applying this framework to Uruguay's inter-firm credit network, using survey data from the Central Bank, we find that credit chains are generally short, with a tiered structure in which some firms act as intermediaries, lending to others further along the chain. We also find that local network motifs such as loops can substantially increase a firm's DebtStreamness, even when its direct borrowing from banks remains the same. Comparing our results with standard economic classifications based on input-output linkages, we find that DebtStreamness captures distinct financial structures not visible through production data. We further validate our approach using two maximum-entropy network reconstruction methods, demonstrating the robustness of DebtStreamness in capturing systemic credit structures. These results suggest that DebtStreamness offers a complementary ecological perspective on systemic credit risk and highlights the role of hidden financial intermediation in firm networks.

econ.GN↗

Correlation between upstreamness and downstreamness in random global value chains

This paper is concerned with upstreamness and downstreamness of industries and countries. Upstreamness and downstreamness measure respectively the average distance of an industrial sector from final consumption and from primary inputs. Recently, Antràs and Chor reported a puzzling and counter-intuitive finding in data from the period 1995-2011, namely that (at country level) upstreamness appears to be positively correlated with downstreamness, with a correlation slope close to $+1$. We first analyze a simple model of random Input/Output tables, and we show that, under minimal and realistic structural assumptions, there is a natural positive correlation emerging between upstreamness and downstreamness of the same industrial sector/country, with correlation slope equal to $+1$. This effect is robust against changes in the randomness of the entries of the I/O table and different aggregation protocols. Secondly, we perform experiments by randomly reshuffling the entries of the empirical I/O table where these puzzling correlations are detected, in such a way that the global structural constraints are preserved. Again, we find that the upstreamness and downstreamness of the same industrial sector/country are positively correlated with slope close to $+1$. Our results strongly suggest that (i) extra care is needed when interpreting these measures as simple representations of each sector's positioning along the value chain, as the ``curse of the input-output identities'' and labor effects effectively force the value chain to acquire additional links from primary factors of production, and (ii) the empirically observed puzzling correlation may rather be a necessary consequence of the few structural constraints (positive entries, and sub-stochasticity) that Input/Output tables and their surrogates must meet.

stat.AP↗

Financial instability transition under heterogeneous investments and portfolio diversification

We analyze the stability of financial investment networks, where financial institutions hold overlapping portfolios of assets. We consider the effect of portfolio diversification and heterogeneous investments using a random matrix dynamical model driven by portfolio rebalancing. While heterogeneity generally correlates with heightened volatility, increasing diversification may have a stabilizing or destabilizing effect depending on the connectivity level of the network. The stability/instability transition is dictated by the largest eigenvalue of the random matrix governing the time evolution of the endogenous components of the returns, for which different approximation schemes are proposed and tested against numerical diagonalization.

q-fin.RM↗

Phase transitions in debt recycling

Debt recycling is an aggressive equity extraction strategy that potentially permits faster repayment of a mortgage. While equity progressively builds up as the mortgage is repaid monthly, mortgage holders may obtain another loan they could use to invest on a risky asset. The wealth produced by a successful investment is then used to repay the mortgage faster. The strategy is riskier than a standard repayment plan since fluctuations in the house market and investment's volatility may also lead to a fast default, as both the mortgage and the liquidity loan are secured against the same good. The general conditions of the mortgage holder and the outside market under which debt recycling may be recommended or discouraged have not been fully investigated. In this paper, to evaluate the effectiveness of traditional monthly mortgage repayment versus debt recycling strategies, we build a dynamical model of debt recycling and study the time evolution of equity and mortgage balance as a function of loan-to-value ratio, house market performance, and return of the risky investment. We find that the model has a rich behavior as a function of its main parameters, showing strongly and weakly successful phases - where the mortgage is eventually repaid faster and slower than the standard monthly repayment strategy, respectively - a default phase where the equity locked in the house vanishes before the mortgage is repaid, signalling a failure of the debt recycling strategy, and a permanent re-mortgaging phase - where further investment funds from the lender are continuously secured, but the mortgage is never fully repaid. The strategy's effectiveness is found to be highly sensitive to the initial mortgage-to-equity ratio, the monthly amount of scheduled repayments, and the economic parameters at the outset. The analytical results are corroborated with numerical simulations with excellent agreement.

q-fin.RM↗

Statistics of the non-zero eigenvalues and singular values of low-rank random matrices with non-negative entries

We compute analytically the probability distribution and moments of the sum and product of the non-zero eigenvalues and singular values of random matrices with (i) non-negative entries, (ii) fixed rank, and (iii) prescribed sums of the entries in each row. Applications of such matrices are discussed in the context of Markov chains, economics and social networks to name a few. All results are valid at finite matrix size and are given in terms of the statistics of vectors of general Dirichlet random variables. Analytical results are corroborated by numerical simulations throughout with excellent agreement.

cond-mat.stat-mech↗