Searcharxiv⌕ Search

arXiv subjects

Silvia Bartolucci

Publications and source records attributed to Silvia Bartolucci.

At least 19 recordsLinked to original sources

Large deviations for linear regressions

Linear regression is one of the simplest and most widely used tools to learn patterns from data: it fits a set of coefficients so that a linear combination of predictors best matches observed responses. The quality of the fit is measured by the residual sum of squares, the total squared mismatch between predictions and data, whose minimum defines the training loss. We consider Gaussian design and noise, with teacher coefficients independently drawn from a general distribution $p(β)$, and a general class of separable regularizers, including Ridge and Lasso. Using the zero-temperature replica method, we compute analytically the large-deviation statistics of the minimum training loss for large numbers $P$ of predictors and $N$ of observations, with $r=P/N$ fixed. The rate function we compute governs rare sample-to-sample fluctuations of the optimal loss. Extensive numerical simulations are in excellent agreement with our theory and clearly show a pronounced deviation from the Gaussian regime of typical fluctuations in the tails.

cond-mat.stat-mech↗

Queue & AI: When Faster Tasks Slow Down the Workflow

Quantifying the workplace productivity effects of Generative Artificial Intelligence is now central to economics, management, and public policy. The deployment of AI tools in customer service, writing, software development, and consulting operations has been reported to generate large per-task productivity gains, typically measured as tasks completed per worker-hour or reductions in mean handle time. We argue that such mean-based metrics can misrepresent AI's effects in workflows where tasks accumulate and compete for scarce human attention. AI assistance can generate a deceptive productivity signature: average completion times fall because AI tools typically supply a fast first draft, yet workflow-level performance deteriorates when a subset of AI errors escapes review and returns as costly downstream rework. We call this divergence between mean task speed and system-level delay the variance wedge. Depending on the operational parameters, the most time-efficient way to complete a workflow may undergo a transition between two task-processing regimes, a fully AI-assisted and a fully manual one. We formalize the mechanism as a queueing model and derive two main implications analytically. First, under congestion, reviewers rationally raise the risk threshold for checking AI outputs, reducing scrutiny precisely when it would matter the most. Second, AI assistance can stabilize an overloaded workflow only when (i) the fraction of tasks handled by AI exceeds a critical threshold, and (ii) the human attention required for review and expected rework is lower than the attention for manual completion, a requirement substantially more stringent than faster draft generation. These results suggest that AI deployment should be evaluated not only by average task speed, but by its overall effects on congestion, rework, and the robustness of human oversight under load.

cs.CY↗

Efficiency for Experts, Visibility for Newcomers: A Case Study of Label-Code Alignment in Kubernetes

Labels on platforms such as GitHub support triage and coordination, yet little is known about how well they align with code modifications or how such alignment affects collaboration across contributor experience levels. We present a case study of the Kubernetes project, introducing label-diff congruence - the alignment between pull request labels and modified files - and examining its prevalence, stability, behavioral validation, and relationship to collaboration outcomes across contributor tiers. We analyse 18,020 pull requests (2014--2025) with area labels and complete file diffs, validate alignment through analysis of over one million review comments and label corrections, and test associations with time-to-merge and discussion characteristics using quantile regression and negative binomial models stratified by contributor experience. Congruence is prevalent (46.6\% perfect alignment), stable over years, and routinely maintained (9.2\% of PRs corrected during review). It does not predict merge speed but shapes discussion: among core developers (81\% of the sample), higher congruence predicts quieter reviews (18\% fewer participants), whereas among one-time contributors it predicts more engagement (28\% more participants). Label-diff congruence influences how collaboration unfolds during review, supporting efficiency for experienced developers and visibility for newcomers. For projects with similar labeling conventions, monitoring alignment can help detect coordination friction and provide guidance when labels and code diverge.

cs.SE↗

A calibrated model of debt recycling with interest costs and tax shields: viability under different fiscal regimes and jurisdictions

Debt recycling is a leveraged equity management strategy in which homeowners use accumulated home equity to finance investments, applying the resulting returns to accelerate mortgage repayment. We propose a novel framework to model equity and mortgage dynamics in presence of mortgage interest rates, borrowing costs on equity-backed credit lines, and tax shields arising from interest deductibility. The model is calibrated on three jurisdictions -- Australia, Germany, and Switzerland -- representing diverse interest rate environments and fiscal regimes. Results demonstrate that introducing positive interest rates without tax shields contracts success regions and lengthens repayment times, while tax shields partially reverse these effects by reducing effective borrowing costs and adding equity boosts from mortgage interest deductibility. Country-specific outcomes vary systematically, and rental properties consistently outperform owner-occupied housing due to mortgage interest deductibility provisions.

q-fin.RM↗

Mapping Microscopic and Systemic Risks in TradFi and DeFi: a literature review

This work explores the formation and propagation of systemic risks across traditional finance (TradFi) and decentralized finance (DeFi), offering a comparative framework that bridges these two increasingly interconnected ecosystems. We propose a conceptual model for systemic risk formation in TradFi, grounded in well-established mechanisms such as leverage cycles, liquidity crises, and interconnected institutional exposures. Extending this analysis to DeFi, we identify unique structural and technological characteristics - such as composability, smart contract vulnerabilities, and algorithm-driven mechanisms - that shape the emergence and transmission of risks within decentralized systems. Through a conceptual mapping, we highlight risks with similar foundations (e.g., trading vulnerabilities, liquidity shocks), while emphasizing how these risks manifest and propagate differently due to the contrasting architectures of TradFi and DeFi. Furthermore, we introduce the concept of crosstagion, a bidirectional process where instability in DeFi can spill over into TradFi, and vice versa. We illustrate how disruptions such as liquidity crises, regulatory actions, or political developments can cascade across these systems, leveraging their growing interdependence. By analyzing this mutual dynamics, we highlight the importance of understanding systemic risks not only within TradFi and DeFi individually, but also at their intersection. Our findings contribute to the evolving discourse on risk management in a hybrid financial ecosystem, offering insights for policymakers, regulators, and financial stakeholders navigating this complex landscape.

q-fin.RM↗

Top eigenpair statistics of diluted Wishart matrices

Using the replica method, we compute the statistics of the top eigenpair of diluted covariance matrices of the form $\mathbf{J} = \mathbf{X}^T \mathbf{X}$, where $\mathbf{X}$ is a $N\times M$ sparse data matrix, in the limit of large $N,M$ with fixed ratio and a bounded number of nonzero entries. We allow for random non-zero weights, provided they lead to an isolated largest eigenvalue. By formulating the problem as the optimisation of a quadratic Hamiltonian constrained to the $N$-sphere at low temperatures, we derive a set of recursive distributional equations for auxiliary probability density functions, which can be efficiently solved using a population dynamics algorithm. The average largest eigenvalue is identified with a Lagrange parameter that governs the convergence of the algorithm, and the resulting stable populations are then used to evaluate the density of the top eigenvector's components. We find excellent agreement between our analytical results and numerical results obtained from direct diagonalisation.

cond-mat.stat-mech↗

Cryptocurrencies in the Balance Sheet: Insights from (Micro)Strategy -- Bitcoin Interactions

This paper investigates the evolving link between cryptocurrency and equity markets in the context of the recent wave of corporate Bitcoin (BTC) treasury strategies. We assemble a dataset of 39 publicly listed firms holding BTC, from their first acquisition through April 2025. Using daily logarithmic returns, we first document significant positive co-movements via Pearson correlations and single factor model regressions, discovering an average BTC beta of 0.62, and isolating 12 companies, including Strategy (formerly MicroStrategy, MSTR), exhibiting a beta exceeding 1. We then classify firms into three groups reflecting their exposure to BTC, liquidity, and return co-movements. We use transfer entropy (TE) to capture the direction of information flow over time. Transfer entropy analysis consistently identifies BTC as the dominant information driver, with brief, announcement-driven feedback from stocks to BTC during major financial events. Our results highlight the critical need for dynamic hedging ratios that adapt to shifting information flows. These findings provide important insights for investors and managers regarding risk management and portfolio diversification in a period of growing integration of digital assets into corporate treasuries.

q-fin.GN↗

DebtStreamness: An Ecological Approach to Credit Flows in Inter-Firm Networks

Understanding how credit flows through inter-firm networks is critical for assessing financial stability and systemic risk. In this study, we introduce DebtStreamness, a novel metric inspired by trophic levels in ecological food webs, to quantify the position of firms within credit chains. By viewing credit as the ``primary energy source'' of the economy, we measure how far credit travels through inter-firm relationships before reaching its final borrowers. Applying this framework to Uruguay's inter-firm credit network, using survey data from the Central Bank, we find that credit chains are generally short, with a tiered structure in which some firms act as intermediaries, lending to others further along the chain. We also find that local network motifs such as loops can substantially increase a firm's DebtStreamness, even when its direct borrowing from banks remains the same. Comparing our results with standard economic classifications based on input-output linkages, we find that DebtStreamness captures distinct financial structures not visible through production data. We further validate our approach using two maximum-entropy network reconstruction methods, demonstrating the robustness of DebtStreamness in capturing systemic credit structures. These results suggest that DebtStreamness offers a complementary ecological perspective on systemic credit risk and highlights the role of hidden financial intermediation in firm networks.

econ.GN↗

Introducing Repository Stability

Drawing from engineering systems and control theory, we introduce a framework to understand repository stability, which is a repository activity capacity to return to equilibrium following disturbances - such as a sudden influx of bug reports, key contributor departures, or a spike in feature requests. The framework quantifies stability through four indicators: commit patterns, issue resolution, pull request processing, and community engagement, measuring development consistency, problem-solving efficiency, integration effectiveness, and sustainable participation, respectively. These indicators are synthesized into a Composite Stability Index (CSI) that provides a normalized measure of repository health proxied by its stability. Finally, the framework introduces several important theoretical properties that validate its usefulness as a measure of repository health and stability. At a conceptual phase and open to debate, our work establishes mathematical criteria for evaluating repository stability and proposes new ways to understand sustainable development practices. The framework bridges control theory concepts with modern collaborative software development, providing a foundation for future empirical validation.

cs.SE↗

Correlation between upstreamness and downstreamness in random global value chains

This paper is concerned with upstreamness and downstreamness of industries and countries. Upstreamness and downstreamness measure respectively the average distance of an industrial sector from final consumption and from primary inputs. Recently, Antràs and Chor reported a puzzling and counter-intuitive finding in data from the period 1995-2011, namely that (at country level) upstreamness appears to be positively correlated with downstreamness, with a correlation slope close to $+1$. We first analyze a simple model of random Input/Output tables, and we show that, under minimal and realistic structural assumptions, there is a natural positive correlation emerging between upstreamness and downstreamness of the same industrial sector/country, with correlation slope equal to $+1$. This effect is robust against changes in the randomness of the entries of the I/O table and different aggregation protocols. Secondly, we perform experiments by randomly reshuffling the entries of the empirical I/O table where these puzzling correlations are detected, in such a way that the global structural constraints are preserved. Again, we find that the upstreamness and downstreamness of the same industrial sector/country are positively correlated with slope close to $+1$. Our results strongly suggest that (i) extra care is needed when interpreting these measures as simple representations of each sector's positioning along the value chain, as the ``curse of the input-output identities'' and labor effects effectively force the value chain to acquire additional links from primary factors of production, and (ii) the empirically observed puzzling correlation may rather be a necessary consequence of the few structural constraints (positive entries, and sub-stochasticity) that Input/Output tables and their surrogates must meet.

stat.AP↗

Mining a Decade of Event Impacts on Contributor Dynamics in Ethereum: A Longitudinal Study

We analyze developer activity across 10 major Ethereum repositories (totaling 129884 commits, 40550 issues) spanning 10 years to examine how events such as technical upgrades, market events, and community decisions impact development. Through statistical, survival, and network analyses, we find that technical events prompt increased activity before the event, followed by reduced commit rates afterwards, whereas market events lead to more reactive development. Core infrastructure repositories like Go-Ethereum exhibit faster issue resolution compared to developer tools, and technical events enhance core team collaboration. Our findings show how different types of events shape development dynamics, offering insights for project managers and developers in maintaining development momentum through major transitions. This work contributes to understanding the resilience of development communities and their adaptation to ecosystem changes.

cs.SE↗

Financial instability transition under heterogeneous investments and portfolio diversification

We analyze the stability of financial investment networks, where financial institutions hold overlapping portfolios of assets. We consider the effect of portfolio diversification and heterogeneous investments using a random matrix dynamical model driven by portfolio rebalancing. While heterogeneity generally correlates with heightened volatility, increasing diversification may have a stabilizing or destabilizing effect depending on the connectivity level of the network. The stability/instability transition is dictated by the largest eigenvalue of the random matrix governing the time evolution of the endogenous components of the returns, for which different approximation schemes are proposed and tested against numerical diagonalization.

q-fin.RM↗

Phase transitions in debt recycling

Debt recycling is an aggressive equity extraction strategy that potentially permits faster repayment of a mortgage. While equity progressively builds up as the mortgage is repaid monthly, mortgage holders may obtain another loan they could use to invest on a risky asset. The wealth produced by a successful investment is then used to repay the mortgage faster. The strategy is riskier than a standard repayment plan since fluctuations in the house market and investment's volatility may also lead to a fast default, as both the mortgage and the liquidity loan are secured against the same good. The general conditions of the mortgage holder and the outside market under which debt recycling may be recommended or discouraged have not been fully investigated. In this paper, to evaluate the effectiveness of traditional monthly mortgage repayment versus debt recycling strategies, we build a dynamical model of debt recycling and study the time evolution of equity and mortgage balance as a function of loan-to-value ratio, house market performance, and return of the risky investment. We find that the model has a rich behavior as a function of its main parameters, showing strongly and weakly successful phases - where the mortgage is eventually repaid faster and slower than the standard monthly repayment strategy, respectively - a default phase where the equity locked in the house vanishes before the mortgage is repaid, signalling a failure of the debt recycling strategy, and a permanent re-mortgaging phase - where further investment funds from the lender are continuously secured, but the mortgage is never fully repaid. The strategy's effectiveness is found to be highly sensitive to the initial mortgage-to-equity ratio, the monthly amount of scheduled repayments, and the economic parameters at the outset. The analytical results are corroborated with numerical simulations with excellent agreement.

q-fin.RM↗

Deep Limit Order Book Forecasting

We exploit cutting-edge deep learning methodologies to explore the predictability of high-frequency Limit Order Book mid-price changes for a heterogeneous set of stocks traded on the NASDAQ exchange. In so doing, we release `LOBFrame', an open-source code base to efficiently process large-scale Limit Order Book data and quantitatively assess state-of-the-art deep learning models' forecasting capabilities. Our results are twofold. We demonstrate that the stocks' microstructural characteristics influence the efficacy of deep learning methods and that their high forecasting power does not necessarily correspond to actionable trading signals. We argue that traditional machine learning metrics fail to adequately assess the quality of forecasts in the Limit Order Book context. As an alternative, we propose an innovative operational framework that evaluates predictions' practicality by focusing on the probability of accurately forecasting complete transactions. This work offers academics and practitioners an avenue to make informed and robust decisions on the application of deep learning techniques, their scope and limitations, effectively exploiting emergent statistical properties of the Limit Order Book.

q-fin.TR↗

HLOB -- Information Persistence and Structure in Limit Order Books

We introduce a novel large-scale deep learning model for Limit Order Book mid-price changes forecasting, and we name it `HLOB'. This architecture (i) exploits the information encoded by an Information Filtering Network, namely the Triangulated Maximally Filtered Graph, to unveil deeper and non-trivial dependency structures among volume levels; and (ii) guarantees deterministic design choices to handle the complexity of the underlying system by drawing inspiration from the groundbreaking class of Homological Convolutional Neural Networks. We test our model against 9 state-of-the-art deep learning alternatives on 3 real-world Limit Order Book datasets, each including 15 stocks traded on the NASDAQ exchange, and we systematically characterize the scenarios where HLOB outperforms state-of-the-art architectures. Our approach sheds new light on the spatial distribution of information in Limit Order Books and on its degradation over increasing prediction horizons, narrowing the gap between microstructural modeling and deep learning-based forecasting in high-frequency financial markets.

q-fin.TR↗

Distribution of centrality measures on undirected random networks via cavity method

The Katz centrality of a node in a complex network is a measure of the node's importance as far as the flow of information across the network is concerned. For ensembles of locally tree-like and undirected random graphs, this observable is a random variable. Its full probability distribution is of interest but difficult to handle analytically because of its "global" character and its definition in terms of a matrix inverse. Leveraging a fast Gaussian Belief Propagation-cavity algorithm to solve linear systems on a tree-like structure, we show that (i) the Katz centrality of a single instance can be computed recursively in a very fast way, and (ii) the probability $P(K)$ that a random node in the ensemble of undirected random graphs has centrality $K$ satisfies a set of recursive distributional equations, which can be analytically characterized and efficiently solved using a population dynamics algorithm. We test our solution on ensembles of Erdős-Rényi and scale-free networks in the locally tree-like regime, with excellent agreement. The distributions display a crossover between multimodality and unimodality as the mean degree increases, where distinct peaks correspond to the contribution to the centrality coming from nodes of different degrees. We also provide an approximate formula based on a rank-$1$ projection that works well if the network is not too sparse, and we argue that an extension of our method could be efficiently extended to tackle analytical distributions of other centrality measures such as PageRank for directed networks in a transparent and user-friendly way.

physics.soc-ph↗

Upstreamness and downstreamness in input-output analysis from local and aggregate information

Ranking sectors and countries within global value chains is of paramount importance to estimate risks and forecast growth in large economies. However, this task is often non-trivial due to the lack of complete and accurate information on the flows of money and goods between sectors and countries, which are encoded in Input-Output (I-O) tables. In this work, we show that an accurate estimation of the role played by sectors and countries in supply chain networks can be achieved without full knowledge of the I-O tables, but only relying on local and aggregate information, e.g., the total intermediate demand per sector. Our method, based on a rank-$1$ approximation to the I-O table, shows consistently good performance in reconstructing rankings (i.e., upstreamness and downstreamness measures for countries and sectors) when tested on empirical data from the World Input-Output Database. Moreover, we connect the accuracy of our approximate framework with the spectral properties of the I-O tables, which ordinarily exhibit relatively large spectral gaps. Our approach provides a fast and analytical tractable framework to rank constituents of a complex economy without the need of matrix inversions and the knowledge of finer intersectorial details.

physics.soc-ph↗

DApps Ecosystems: Mapping the Network Structure of Smart Contract Interactions

In recent years, decentralized applications (dApps) built on blockchain platforms such as Ethereum and coded in languages such as Solidity, have gained attention for their potential to disrupt traditional centralized systems. Despite their rapid adoption, limited research has been conducted to understand the underlying code structure of these applications. In particular, each dApp is composed of multiple smart contracts, each containing a number of functions that can be called to trigger a specific event, e.g., a token transfer. In this paper, we reconstruct and analyse the network of contracts and functions calls within the dApp, which is helpful to unveil vulnerabilities that can be exploited by malicious attackers. We show how decentralization is architecturally implemented, identifying common development patterns and anomalies that could influence the system's robustness and efficiency. We find a consistent network structure characterized by modular, self-sufficient contracts and a complex web of function interactions, indicating common coding practices across the blockchain community. Critically, a small number of key functions within each dApp play a pivotal role in maintaining network connectivity, making them potential targets for cyber attacks and highlighting the need for robust security measures.

cs.CY↗